
Model-size decisions get justified with FLOPs; if that number does not track energy on real hardware, the sizing argument needs a different basis.
articleDiagnostic Foundation for Evaluating LLMs' Research Integrity as Co-ScientistsYash Tripathi, Silu Sharma, Sai Sidhanth Manoharan Jayanthi, Shivank Garg, Lin Li
articleBetter, Faster, Stronger: Programmatic Skill Learning Best Reduces Agent CostZixi Huang, Xiheng Wang, Andrew Wang, William Jurayj, Bernal Jim\'enez Guti\'errez, Daniel Khashabi, Nicholas Andrews
articleHow well LLM-based test generation techniques perform with newer LLM versions?Michael Konstantinou, Renzo Degiovanni, Mike PapadakisSign in to comment.
Loading comments…