Vibeleaderboard
Index / tool
Visit huggingface.co
Category
AI Tools
Rank
No. 2323Tools index
Type
TOOL
Date

About

VisionLaya adds image understanding to the Laya model by swapping its ModernBERT text encoder for SmolVLM-256M-Instruct, so it can answer choice, score, and yes/no-probability questions about an image plus optional text in a single forward pass with no text generation. It keeps Laya's existing predict(state, questions) API, proper-scoring-rule training, and temperature calibration unchanged.

Tags

vision-language-modelcalibrationsmolvlmhuggingfacestructured-output

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.