The Curve Was Stable. The Judgment Was Not.
Published in Computational Sciences
The easiest AI failures to measure are the ones for which we already know the correct answer. A hallucinated date can be checked. A calculation can be recomputed. A factual claim can be compared with a source.
Normative conflict offers no such luxury.
Consider an AI system asked whether privacy should be protected when disclosure might prevent harm, or whether a painful truth should be told when emotional protection also carries moral weight. The difficulty is not simply choosing between two options. A competent response has to keep both claims visible long enough to explain what is gained, what is lost, and why one course of action is nevertheless recommended.
I became interested in the point at which that capacity disappears.
From failure to transition
The failure that first interested me was not dramatic. Some responses became smoother. They moved away from the particulars of the dilemma, relied on generic reassurance or safety language, or avoided naming the conflict altogether.
That suggested a different alignment question: instead of asking only whether a model produced an acceptable answer, could I estimate the probability that it would stop sustaining a coherent interpretation as normative pressure increased?
I used 60 normatively ambiguous scenarios spanning five conflict levels, more than 9,000 generated responses, three transformer-based text-only LLM families, five temperature settings, two sampling variants, and four levels of persona perturbation. I estimated interpretative-collapse rates as effective sampling variance and per-token entropy changed.
Adapted from Fig. 1 of Kim (2026), CC BY 4.0.
Across the tested model families, the fitted Kramers-like slopes were statistically indistinguishable within a prespecified equivalence margin. Adjusted R² values ranged from 0.93 to 0.94 across the main JPPI conditions. In entropy space, collapse probability increased monotonically across the observed range.
The claim is deliberately narrow. I am not proposing a physical thermodynamics of language models. The formulation is phenomenological. Effective variance and entropy are two parameterisations of the same calibrated relationship, not independent discoveries. Stability across three tested transformer families does not establish a universal alignment constant.
The curve was cleaner than the construct
The most uncomfortable result was not the collapse curve. It was the human judgment underneath it.
For 1,500 unique responses, three trained raters produced 4,500 binary collapse judgments. Pairwise Cohen’s kappa values were only about 0.04 to 0.05, despite raw agreement of approximately 0.57. Agreement on the auxiliary affective-degradation ratings was also low.
When the reliability statistics came back, they forced a second measurement problem into the foreground. The model’s behaviour was uncertain, but so was the human criterion by which collapse was being judged. I did not treat that disagreement as either simple annotation noise or proof of moral plurality. I reported it, retained a conservative majority-and-adjudication rule, and released the reliability summary with the dataset.
A clean model fit does not make a contested construct clean.
Disagreement may contain genuine normative plurality, but also ambiguity, threshold differences, finite-rater noise, or ordinary error. As Plank argued in work on human label variation, disagreement is informative, but it does not interpret itself.
This separates two problems that are often collapsed together. Model uncertainty and evaluative uncertainty are not the same thing. Farquhar and colleagues showed how semantic entropy can help detect factual confabulation. Normative conflict adds another layer: sometimes there is no factual answer waiting to be recovered, and the evaluator must also decide what counts as an adequate representation of competing values.
What collapse looked like, and what would challenge the relationship
Responses that remained coherent typically kept the conflict visible. They acknowledged competing norms, articulated trade-offs, and stayed anchored to the specific human situation. Collapsed responses more often oscillated between incompatible commitments, retreated into generic safety language, or offered reassurance without engaging the conflict itself.
Failure did not always sound like failure.
That matters for value-sensitive AI. A system used in counselling, education, healthcare, law, or intimate advice may need not only to avoid factual error, but also to preserve tensions that should not yet be resolved.
The paper also specifies what would weaken the relationship. Architecture-specific slope divergence beyond the equivalence margin would challenge cross-family stability. Stable non-monotonic collapse under increasing conflict would challenge the rate-form interpretation. Failure to reproduce the entropy-collapse relationship independently would suggest pipeline specificity. If prompt-level hierarchical modelling with scenario-level random effects removed the pattern, the current slopes would need to be reinterpreted as prompt heterogeneity.
These are not footnotes to the claim. They define its boundaries.
A different question for alignment
Much of alignment evaluation is organised around convergence: did the model arrive at the desired answer, preference, policy, or behaviour?
Normatively difficult interaction suggests another question:
Can a system preserve a conflict that should not yet be collapsed?
That shifts attention from the final recommendation to the reasoning that precedes it. It asks whether competing values remain represented, whether scenario-specific costs survive compression, and whether uncertainty produces careful qualification or merely generic closure.
The present study does not settle that problem. It gives us a way to begin measuring one part of it.
The relevant boundary may not lie between a model that is right and a model that is wrong. It may lie between a system that can still carry a difficult problem and one that has quietly made the difficulty disappear.
The article, dataset, and analysis code are openly available:
Ryan SangBaek Kim (2026), “The affective thermodynamic relationship: an empirical information-theoretic scaling relationship for normative-conflict collapse in large language models,” Communications AI & Computing, 1, Article 14.
DOI: 10.1038/s44488-026-00006-y
Further reading
Farquhar, S. et al. (2024). Detecting hallucinations in large language models using semantic entropy. Nature, 630, 625–630. DOI: 10.1038/s41586-024-07421-0.
Plank, B. (2022). The “problem” of human label variation: on ground truth in data, modeling and evaluation. EMNLP 2022, 10671–10682. DOI: 10.18653/v1/2022.emnlp-main.731.
Please sign in or register for FREE
If you are a registered user on Research Communities by Springer Nature, please sign in