About
ICML 2025 paper accelerating large language models by compressing each segment into a single separator token to speed up inference.
Why it made the leaderboard
ICML 2025 technique that speeds up LLM inference by compressing each text segment into a single separator token — a concrete way to cut attention cost if inference latency is your bottleneck.
Tags
llminferencespeedupresearchcompression
Tech Stack
CC++CudaMakefilePythonShell
Media

Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.
