Vibeleaderboard
Index / tool

Youtu-Parsing-Omni

huggingface.co
Visit huggingface.co
Category
Developer Tools
Rank
No. 3053Tools index
Type
TOOL
Date

About

A multimodal model from Tencent on Hugging Face for document parsing and OCR, tagged for image, audio and video input. It loads through Transformers with custom code and can be served with vLLM or SGLang through an OpenAI-compatible API.

Tags

ocrdocument-parsingmultimodaltencenttransformersvision-language

Media

Youtu-Parsing-Omni

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.