Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese
qwenlm.github.io- Category
- Other
- Type
- ARTICLE
- Added
- Jul 21, 2026
About
CLIP1 is a phenomenal playmaker in vision and multimodal representation learning. It plays not only as a foundation model but also a bridge between vision and language. It has triggered a series of research in different fields, especially text-to-image generation. However, we find that there is a necessity for a language-specific CLIP for applications, especially cross-modal retrieval, and there is no opensourced Chinese CLIP with good performance. We therefore launched this project to promote t
Why it made the leaderboard
If you're building cross-modal retrieval or text-to-image pipelines for Chinese content, this gives you a purpose-trained Chinese CLIP rather than forcing English-centric embeddings onto Chinese text and image data. It targets the multilingual gap directly instead of relying on translation workarounds.
Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.