
A deep, well-organized survey of self-supervised representation learning methods (CPC, MoCo, SimCLR, BYOL) that gives engineers the conceptual to build and choose / approaches for AI-native products.
“Given a task and enough labels, supervised learning can solve it really well. Good performance usually requires a decent amount of labels, but collecting manual labels is expensive (i.e. ImageNet) and hard to be scaled up.”
Lilian Weng
“The self-supervised task , also known as pretext task , guides us to a supervised loss function. However, we usually don’t care about the final performance of this invented task.”
Lilian Weng
“In order to identify the same image with different rotations, the model has to learn to recognize high level object parts, such as heads, noses, and eyes, and the relative positions of these parts, rather than local patterns.”
Lilian Weng
“Other than trivial signals like boundary patterns or textures continuing, another interesting and a bit surprising trivial solution was found, called “chromatic aberration” . It is triggered by different focal lengths of lights at different wavelengths passing through the lens.”
Lilian Weng
“A video contains a sequence of semantically related frames. Nearby frames are close in time and more correlated than frames further away. The order of frames describes certain rules of reasonings and physical logics; such as that object motion should be smooth and gravity is pointing down.”
Lilian Weng
Checking sign-in…
Loading comments…