Find the right skill, CLI, harness, or service for the job.
Pairing LFM2.5-VL-3B with this small 280M-parameter draft model cuts decode latency substantially on both H100 and Apple silicon, a practical lever for anyone serving vision-language models without touching output quality.
Checking sign-in…
Loading comments…