
NL2SQL accuracy above 89% on academic benchmarks drops to 57-80% on tiered enterprise Oracle schemas, and the silent-divergence metric catches queries that execute cleanly while returning wrong results.
articleBC-Bench: Evaluating Agentic Engineering in a Domain-Specific Language for ERPHaoran Sun, Klaus Marius Hansen
articleHRAG – Hybrid RAG on €116/month of Hetzner, officially benchmarkedvictor_edka
articleAre the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon StatementsXinke Tong, Xuanming Zhang, Tianyi Tang, An Yang, Jiatu Hu, Guojie Lin, Zhenzhen Shi, Lingfeng Zeng, Boyu Yang, Bing Zhao, Hu Wei, Lin Qu, Dayiheng LiuChecking sign-in…
Loading comments…