If frontier models reason in latent loops rather than visible , monitoring and any safety tooling built on readable traces stop seeing the actual computation.
articleHow Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought TrajectoriesHui Wei, Junda Wu, Sheldon Yu, Sizhe Zhou, Yizhu Jiao, Ming Zhong, Bowen Jin, Tong Yu, Shijia Pan, Jiawei Han, Julian McAuley
articleLanguage Models Don T Always Say What They Think Unfaithful Explanations In Chain Of Thought Prompting 2023 05 07CohereChecking sign-in…
Loading comments…