Vibeleaderboard
Index / tool

LocateAnything

research.nvidia.com
Visit research.nvidia.com
Category
AI Tools
Rank
No. 1310Tools index

Previous survey · No. 1315 ·

Pricing
Open Source
Type
TOOL
Use case
Models: Train & Run · Data Processing
Interfaces
Web
Builder
NVlabs
Date

About

LocateAnything is a fast, high-quality vision-language grounding framework that uses Parallel Box Decoding (PBD) to predict bounding boxes as atomic units in a single forward pass. It supports diverse localization tasks including document understanding, GUI grounding, dense object detection, and OCR localization. By decoding geometric elements in parallel rather than sequentially, it achieves up to 2.5× faster throughput while improving localization accuracy.

What it does

The page describes LocateAnything as a vision-language model that finds and boxes objects, text, and UI elements in an image from a text query. It predicts each box in one step instead of token by token, which the page says makes it faster and more accurate on tasks like document parsing, GUI finding, and OCR.

Stated on the product site

Inference modes offered
The model offers a fast parallel decoding mode, a slower sequential decoding mode, and a hybrid mode that switches between them
Model architecture components
The model is built from a named vision encoder and a named language decoder joined by a projector layer
Training dataset availability
The training dataset used to build the model is stated to be publicly released on Hugging Face
Supported task types
The page groups four task types under one model: document understanding, GUI grounding, dense object detection, and OCR localization

Not stated on the site

  • The page does not show pricing or licensing terms for using the model.
  • The page does not show what platforms, operating systems, or hardware (beyond the GPU used in its own benchmarks) the model is meant to run on.
  • The page does not show whether there is a hosted app, demo, or API for trying the tool directly, as opposed to downloading model weights or code.

Written from the product site at research.nvidia.com.

What it can do

  • Locate objects in an image using natural language queries

    Image and a text description of a target object → Bounding box coordinates around the described object

  • Detect and localize UI elements in GUI screenshots

    Screenshot of a graphical user interface and a text query → Bounding boxes identifying the specified GUI components

  • Extract and localize text regions in documents

    Document image and OCR localization query → Bounding boxes around detected text regions or specific text content

  • Perform dense object detection across an entire image

    Image with multiple objects → Bounding boxes for all detected object instances with labels

  • Ground language references to regions in complex documents

    Document image and a natural language reference to a specific region or element → Bounding box pinpointing the referenced document region

  • Run high-throughput batch visual grounding

    Multiple images paired with language queries → Bounding box predictions for all queries processed in parallel at up to 2.5× faster speed

Tags

vision-languageobject detectiongroundingbounding boxvlmparallel decodingcomputer visionmultimodal

Media

LocateAnything

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.