Vibeleaderboard
Index / article

Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

qwenlm.github.io
Visit qwenlm.github.io
Category
Other
Type
ARTICLE
Added
Jul 21, 2026

About

CLIP1 is a phenomenal playmaker in vision and multimodal representation learning. It plays not only as a foundation model but also a bridge between vision and language. It has triggered a series of research in different fields, especially text-to-image generation. However, we find that there is a necessity for a language-specific CLIP for applications, especially cross-modal retrieval, and there is no opensourced Chinese CLIP with good performance. We therefore launched this project to promote t

Why it made the leaderboard

If you're building cross-modal retrieval or text-to-image pipelines for Chinese content, this gives you a purpose-trained Chinese CLIP rather than forcing English-centric embeddings onto Chinese text and image data. It targets the multilingual gap directly instead of relying on translation workarounds.

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.