New Research & Open Source: We are releasing CL-bench, a systematic benchmark to advance Context Learning. To deliver real-world utility, LMs must transition from static parameter memorization to dynamic reasoning within complex contexts. Addressing this challenge, Tencent HY and Fudan University have jointly released a new paper, "CL-bench: A Benchmark for Context Learning". This is the first research output from Vinces Yao (@ShunyuYao12)’s team since he joined Tencent. This release also signals the official launch of the Tencent HY Research, a blog dedicated to sharing our frontier insights and innovations. 🌐 Project Page: https://t.co/B3x5nZgnGE 📖 Blog:


Why CL-bench? Most benchmarks test what a model remembers, but real-world high-value tasks like coding in massive private repositories or real-time financial analysis demand what a model can learn on the fly. We define this as "Context Learning". This is the ability to leverage new knowledge beyond pre-training. While humans learn from context naturally, even frontier models show significant gaps. CL-bench shifts the focus toward how models reason and resolve tasks using information provided in-context.
To bridge the gap between benchmarks and reality, CL-bench features: -500 complex contexts & 1,899 expert-curated tasks. -31,000+ rigorous validation rules. -Pure Contextual Reasoning: Tasks require models to apply new knowledge NOT found in pre-training but provided exclusively within the context. Our evaluations of ten frontier models find that models solve only 17.2% of tasks on average, revealing that LMs have yet to achieve effective context learning, which poses a critical bottleneck for tackling real-world, complex context-dependent tasks.
Our goal is to move LMs beyond leaderboard chasing toward true utility. By open-sourcing CL-bench, we aim to enable the community to make LMs more intelligent and advancing their deployment in real-world scenarios.
This quantifies the gap that appears when an must work inside a private repo or fresh document set rather than recall . A 17.2% average is a hard ceiling to design around.
Checking sign-in…
Loading comments…