Youtu-Parsing-Omni
huggingface.co- Category
- Developer Tools
- Rank
- No. 3053Tools index
- Type
- TOOL
- Date
About
A multimodal model from Tencent on Hugging Face for document parsing and OCR, tagged for image, audio and video input. It loads through Transformers with custom code and can be served with vLLM or SGLang through an OpenAI-compatible API.
Tags
ocrdocument-parsingmultimodaltencenttransformersvision-language
Media
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.