Vibeleaderboard
← All Intel
Intel / article

How Should Diffusion Language Models Edit Code?

Source
Xijia Tao, Ziru Liu, Shansan Gong, Jiacheng Ye, Kecheng Chen, Zirui Wu, Lin Zheng, Xinyu Fu, Rui Liu, Lingpeng Kong
Author
Xijia Tao, Ziru Liu, Shansan Gong, Jiacheng Ye, Kecheng Chen, Zirui Wu, Lin Zheng, Xinyu Fu, Rui Liu, Lingpeng Kong
Date
Key takeaways · AI-distilled
  • The study compares four editing interfaces for masked diffusion language models on CanItEdit: whole-file rewriting, search-and-replace, locate-then-infill, and -level editing.
  • Diffusion models produce coordinated changes well when the correct edit locations are supplied, but lose much of that ability when they must predict the locations themselves, a gap the authors call the composition gap.
  • Giving the model the intact original code helps it fill multiple edit regions, but does not solve the problem of choosing which regions to edit.
  • A sentence-level Wiki editing probe shows the same gap between supplied and predicted locations, suggesting the problem is not specific to code.
Terms in this piece · Glossary
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
Why it matters

If you consider diffusion models for code editing, the bottleneck is choosing edit regions, not generating the changes. Regions must cover every required change yet stay tight, because widening them to ensure coverage sharply reduces success.

Recommended reads
Comments

Checking sign-in…

Loading comments…