
Concrete evidence that parameter count is the wrong selection axis for classification-style routing: tuned 3B models beat larger base models, and the benchmarks most people cite no longer separate anything.
articleOptimismBench: Forecasting Bias and the Alignment Effect in Language Model JudgmentSeonglae Cho, Adriano Koshiyama
articleWhen Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey ResponsesZihan Chen, Di Zhu, Lei Nico Zheng
articleTraining Compute-Optimal Large Language ModelsJordan Hoffmann et al.Checking sign-in…
Loading comments…