📢 Announcing voyage-multimodal-3, our first multimodal embedding model! It vectorizes interleaved text & images, capturing key visual features from screenshots of PDFs, slides, tables, figures, etc. +19.63% accuracy gain on 3 multimodal retrieval tasks (20 datasets)! 🧵🧵

Most multimodal embedding models (e.g. @cohere multimodal v3) mimic @OpenAI CLIP and use separate networks for images & text. In contrast, voyage-multimodal-3, inspired by the modern vision-language model architecture, processes both text and visuals within the same transformer.

Thanks to the unified architecture, voyage-multimodal-3 excels in mixed-modality retrieval. Testing on varying ratios of text and screenshots with identical content showed that, unlike other models, its performance remains steady as the screenshot ratio increases.

No need anymore for screen parsing models, layout analysis, or other complex text extraction pipelines. Simply take a screenshot of your document and/or prepend extra text (e.g., meta info), and vectorize the interleaved data! Start with our notebook: https://t.co/QIbj9zO3wM
a document screenshot directly removes layout analysis and text extraction from the pipeline, and the unified encoder avoids the accuracy drop dual-encoder models show on mixed text and image corpora.
Checking sign-in…
Loading comments…