ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence
Source
arxiv.org
Author
Sanjay Mishra, Divya Chukkapalli, Ganesh R. Naik
Date
Why it matters
NL2SQL accuracy above 89% on academic benchmarks drops to 57-80% on tiered enterprise Oracle schemas, and the silent-divergence metric catches queries that execute cleanly while returning wrong results.
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.