Check out our demos using LFM2.5-VL-3B, our latest lightweight, vision-language model that reads screens, documents, and the physical world. First up: LFM2.5-VL-3B running fully on-device in the browser with WebGPU to understand a document page. The model parses the entire layout in one pass and returns regions and labels that the interface renders as an overlay. The demo highlights OCR and layout understanding for visually structured content such as forms, reports, receipts, and other documents. 🧵
Here, LFM2.5-VL-3B runs fully on-device in the browser with WebGPU to ground objects in an image. Given a natural-language request, the model finds the referenced objects in the scene and returns their locations. The interface then outlines each result with a bounding box, connecting visual understanding with precise spatial coordinates. (2/4)
We combine visual understanding with tool calling, running fully on-device in the browser with WebGPU. Given an image of a dish and a request for its recipe, the model understands the visual input and decides to call an appropriate search tool. The demo illustrates how LFM2.5-VL-3B can turn information from an image into a structured function call and an actionable workflow. (3/4)
Finally, here’s our model understanding and navigating a digital interface. Given the goal of opening a specified documentation page, the model analyzes the current screen and identifies where to click next through visual grounding. Repeating this process across successive screens creates a multi-step navigation flow that reaches the requested page.
A 3B vision model handling OCR, and screen navigation on-device means computer-use and document workflows can run client-side, with no image leaving the browser.
Checking sign-in…
Loading comments…