Code LLMs are central to software engineering, but their stochasticity poses real-world risks. Code-MUE measures uncertainty through execution-based semantic interaction graphs, revealing most models can't predict their own errors. If your code pipeline leans on a model that can't say when it's wrong, you don't actually know what it'll do.
“arXiv:2607.12273v2 Announce Type: replace-cross Abstract: As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-world risks, where even minor errors can lead …”
Read the source →STATUS
ACTIVE
CATEGORY
Models
EVIDENCE
Not yet assessed
ENTITY
Code-MUE, Code LLMs, semantic interaction graphs
DECISION
Automated · no editorial override
LAST OBSERVED
Aug 3, 2026