
Teams doing RLVR post-training outside English have been extrapolating from English-only results. This says the native-language reasoning penalty is small, which changes how those training runs get configured.
articleModular Cognitive Architecture Emerges in Large Language ModelsPengrui Han, Jacob Andreas, Evelina Fedorenko, Andrea Gregor de Varda
articleDon't Want Your LLM to Recommend Nuclear Strike? Try Asking It in JapaneseRian Touchent (ALMAnaCH)Sign in to comment.
Loading comments…