Vibeleaderboard
Index / tool
Visit research.nvidia.com
Category
AI Tools
Rank
No. 645Tools index
Pricing
Open Source
Platform
web
Type
TOOL
Added
Jul 2, 2026

About

LocateAnything is a fast, high-quality vision-language grounding framework that uses Parallel Box Decoding (PBD) to predict bounding boxes as atomic units in a single forward pass. It supports diverse localization tasks including document understanding, GUI grounding, dense object detection, and OCR localization. By decoding geometric elements in parallel rather than sequentially, it achieves up to 2.5× faster throughput while improving localization accuracy.

What it can do

  • Locate objects in an image using natural language queries

    Image and a text description of a target objectBounding box coordinates around the described object

  • Detect and localize UI elements in GUI screenshots

    Screenshot of a graphical user interface and a text queryBounding boxes identifying the specified GUI components

  • Extract and localize text regions in documents

    Document image and OCR localization queryBounding boxes around detected text regions or specific text content

  • Perform dense object detection across an entire image

    Image with multiple objectsBounding boxes for all detected object instances with labels

  • Ground language references to regions in complex documents

    Document image and a natural language reference to a specific region or elementBounding box pinpointing the referenced document region

  • Run high-throughput batch visual grounding

    Multiple images paired with language queriesBounding box predictions for all queries processed in parallel at up to 2.5× faster speed

Why it made the leaderboard

NVIDIA's grounding framework that decodes bounding boxes in parallel — each box an atomic unit in a single forward pass — instead of dribbling out coordinate tokens sequentially. Trained on 138M queries and 785M boxes, it covers GUI grounding, document layout, OCR localization, and dense detection in one model.

Tags

vision-languageobject detectiongroundingbounding boxvlmparallel decodingcomputer visionmultimodal

Media

LocateAnything

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.