On the Robustness of LLMs' Internal Representation of Code Correctness
Source
arxiv.org
Author
Francisco Ribeiro, Sohaila Abdulsattar, Renata Gonzalez, Mahmoud Kassem, Sarah Nadi
Date
Why it matters
Before building verification on internal correctness signals, know the effect looks sensitive to how it is extracted rather than being a stable property of the model — the earlier result does not transfer as a drop-in.