
Stacking many instructions in a single prompt causes non-linear, reproducible failure (e.g., JSON output conflicting with other constraints); a training-free prompt-compiler rewrite recovers up to 11 points of instruction-following for weaker models, offering a cheap mitigation for production prompt pipelines.
“Instruction-following degrades non-linearly: the follow rate falls from ~96% to as low as 20%, driven by a structured and reproducible set of pairwise conflicts.”
“A single "output JSON" constraint, for example, is jointly unsatisfiable with nine others.”
“It recovers up to +11 points of follow rate for weaker models, which are also the models most often deployed at scale, while leaving stronger models, which already internalise the same structure, essentially unchanged.”
articleCompound Prompt Constraints in LLM Code Generation: A Factorial Study of Format, Persona, and UrgencyShrenik Jadhav, Nickalsa LaPlaca, Caleb Stone, Ashok Raja, Omar Ochoa, Vidhyashree Nagaraju
articleSteering Instruction Hierarchies at Inference TimeSiqi Zeng, Sewoong Lee, Han Zhao, Julia Hockenmaier
articleHow well LLM-based test generation techniques perform with newer LLM versions?Michael Konstantinou, Renzo Degiovanni, Mike PapadakisChecking sign-in…
Loading comments…