
Compressing structured records into compact before analysis cuts token use to roughly 1-2% of the original while preserving accuracy and producing traceable rationales.
“TEFM achieves token efficiency by compressing lengthy structured observations into compact Behavioral Code tokens, dramatically reducing token consumption with minimal information loss.”
“Comprehensive experiments across various domain datasets and model backbones (Qwen3, Gemma-2, Phi-4) show that TEFM achieves competitive classification accuracy with dramatic token reduction (approximately 1% token retention in clinical and 2% in security domains) while producing faithful rationales.”
articleAccelerating LLM Inference via Vector Index Based Output EmbeddingsMartin Loretz, Sepp Hochreiter
articleCode Health in LLM-Based Test Generation: Effectiveness and Token EfficiencyFreya Wirdemann, Markus Borg, Nadim Hagatulah, Adam Tornhill
articleA Removal Based Approach to Improve LLM Faithfulness at Test-TimeQinglan Luo, S M A Nahian, John Guttag, S. Mazdak Abulnaga, Katie MattonChecking sign-in…
Loading comments…