grounding — Tying a model's answers to checkable sources — retrieved documents, live data, tool results — instead of letting it answer from memory alone.
hallucination — When a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Gemini 3.5 Flash's native Computer Use API posted the best mean reward (0.267) among tested frontier models on a hard CAD benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition →, though it still solved only 5 of 25 tasks and hallucinated screen state after repeated screenshots.