VisionLaya
huggingface.co- Category
- AI Tools
- Rank
- No. 2323Tools index
- Type
- TOOL
- Date
About
VisionLaya adds image understanding to the Laya model by swapping its ModernBERT text encoder for SmolVLM-256M-Instruct, so it can answer choice, score, and yes/no-probability questions about an image plus optional text in a single forward pass with no text generation. It keeps Laya's existing predict(state, questions) API, proper-scoring-rule training, and temperature calibration unchanged.
Tags
vision-language-modelcalibrationsmolvlmhuggingfacestructured-output
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.