- Category
- AI Tools
- Pricing
- Open Source
- Type
- TOOL
- Builder
- openai
- GitHub
- 34.1k stars
- Added
- May 22, 2026
About
OpenAI's vision-language model that predicts the most relevant text snippet for a given image, trained on 400M image-text pairs from the web.
Why it made the leaderboard
OpenAI's vision-language model trained on 400M image-text pairs — predicts the most relevant text for a given image, enabling image-text matching without task-specific training.
Tags
clipvisionmultimodalopenaiembeddings
Tech Stack
Python
Media
Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.
