A great deal of modern computational research is built around robustness.
We ask whether a result survives another model specification, another random seed, another estimator, another prompt, another threshold, another architecture.
These are good habits.
But while working on my latest paper, I found myself returning to a different question.
What if the result survives the model, but not the measurement?
That question emerged from an analysis of 351,734 naturalistic relationship narratives. The paper is now published in Acta Psychologica, but the result that stayed with me was not the largest coefficient, the most elaborate model, or the clustering solution.
It was a much simpler observation.
The inferential statement changed when the affective measure changed.
That experience has made me think more seriously about a form of robustness that affective science and emotion AI do not always test explicitly:
robustness to plausible measurement choices.
The first result was almost flat
The original question was straightforward.
Are more structurally complex narratives also more affectively intense?
The corpus contained 351,734 relationship narratives. Narrative structure was represented through Level of Complexity, or LoC. Affect was initially summarised using the absolute magnitude of net valence.
The correlation was:
r = -0.061
With a sample this large, even a very small association can be statistically significant. But statistical significance was not the interesting issue. The corresponding explained variance was only 0.37%.
Substantively, the association was extremely weak.
That result could have supported a simple conclusion: narrative complexity and affective magnitude are almost unrelated.
But the phrase “affective magnitude” was doing more work than it first appeared to be doing.
Net valence has a particular mathematical property.
Positive and negative sentence-level signals can cancel each other.
A narrative containing strong positive and strong negative language may therefore end up close to zero after aggregation.
That is not a computational error.
It is what the operation is designed to do.
The question was whether that operation answered the question we thought we were asking.
What survives cancellation?
To examine that issue, I compared net-valence magnitude with two summaries that retain more of the magnitude of sentence-level affect before opposite directions cancel.
Mean absolute valence, or MAV, produced:
r = 0.186
Root-mean-square valence, or RMS, produced:
r = 0.220
The explained variance was 3.47% for MAV and 4.85% for RMS.
These remain small effects.
I want to emphasize that because the story of this paper is not that a weak effect suddenly became strong.
It did not.
What changed was the kind of statement the data supported.
Under net valence, the most defensible interpretation was that narrative complexity was almost unrelated to affective magnitude.
Under MAV and RMS, the most defensible interpretation became that narrative complexity was weakly, but consistently, associated with cancellation-resistant affective magnitude.
The corpus was the same.
The narrative-complexity measure was the same.
The sentence-level affective signals were the same.
The difference was the operation used to summarise those signals.
The arithmetic is trivial. The methodological consequence is not.
There is no mathematical surprise here.
If +8 and -8 are averaged, they can become zero.
If their magnitudes are retained before aggregation, they do not disappear in the same way.
This is elementary arithmetic.
That is precisely why I think the example matters.
Many important measurement decisions are mathematically obvious at the moment we make them, but become invisible once the resulting variable enters a dataset.
After transformation, the variable acquires a clean name.
“Emotion.”
“Valence.”
“Engagement.”
“Stress.”
“Ground truth.”
We then model the variable as though the transformation that created it were no longer part of the scientific claim.
But every summary operation has a characteristic loss.
A mean reduces variation.
A majority vote removes minority judgments.
A discrete class compresses within-category structure.
A temporal average removes trajectory.
Net valence allows opposing directions to cancel.
None of these operations is inherently wrong. Reduction is unavoidable in science.
The problem begins when the reduction becomes invisible.
The same corpus had already taught me a different lesson
This was not the first paper I had written from these 351,734 narratives.
An earlier PLOS ONE study used the same ANAD corpus to examine narrative-affect discrepancy as a structured expressive space.
That analysis asked where narrative structure and expressed affect diverged.
The Acta Psychologica paper asks a different question.
How much of that apparent relationship depends on the way affect itself is operationalised?
The progression has been useful for me because the same empirical substrate has exposed different layers of the measurement problem.
ANAD provides the derived-feature infrastructure.
The PLOS ONE study mapped the geometry of discrepancy.
The Acta Psychologica study asks how the interpretation of that geometry changes when the affective measure changes.
What began as a study of discrepancy has gradually become, for me, a study of the conditions under which discrepancy becomes visible at all.
Affective science already knows that targets are constructed
This problem has a much longer intellectual history than my dataset.
Lisa Feldman Barrett and colleagues showed how difficult it is to infer discrete emotional states directly from facial movements. Abigail Jacobs and Hanna Wallach formalised the broader problem of construct validity in computational systems, showing that theoretical constructs and their operationalisations can diverge. Saif Mohammad's Ethics Sheet for Automatic Emotion Recognition and Sentiment Analysis made many of the assumptions surrounding data, labels, methods, and evaluation in emotion recognition explicit.
More recently, D'Amelio and colleagues examined electrodermal-activity-based emotion recognition and found that arousal prediction generally outperformed valence prediction. Their review also identified a mismatch between the dimensional emotion frameworks commonly invoked in the literature and the continued dominance of classification approaches in machine learning.
These literatures converge on an uncomfortable fact.
Before we evaluate prediction, we have already made decisions about what is being predicted.
That decision is not merely a preprocessing choice.
It is part of the scientific theory.
Emotion AI makes the problem harder
The problem becomes especially consequential in emotion AI because several distinct objects can all be labelled “emotion.”
A target can represent a person's self-report.
It can represent what an observer perceived.
It can represent the category an experimental stimulus was intended to elicit.
It can represent a majority vote among annotators.
It can represent a thresholded position in valence-arousal space.
It can represent a physiological activation pattern.
All of these may be legitimate objects of prediction.
But they are not interchangeable.
A model that predicts observer labels with 95% accuracy may be excellent at reproducing observer judgments.
That does not establish 95% access to another person's subjective experience.
A model trained on self-report may accurately predict self-report.
That does not transform self-report into an infallible window onto emotional experience.
A model trained on majority labels may accurately reproduce consensus.
That does not prove that minority interpretations were erroneous.
Model performance inherits the epistemic status of the target.
No improvement in architecture can remove that dependency.
Robustness should not stop at the model
This is where I think there is room for a more explicit research programme.
We already expect serious computational work to examine sensitivity to modelling decisions.
In affective research, we should increasingly ask the same question about measurement decisions.
For affective inference, robustness should include sensitivity to plausible measurement choices, not only sensitivity to model choices.
The relevant question is not whether every alternative measure gives exactly the same coefficient.
That would be unrealistic.
The stronger question is whether the substantive inference survives reasonable alternative operationalisations.
Does the sign remain?
Does the ranking remain?
Does the category assignment remain?
Does the conclusion move from “no meaningful relationship” to “weak relationship”?
Does an observation that appears affectively neutral under one representation retain substantial affective magnitude under another?
Does a model appear better simply because the target representation is easier to predict?
These are measurement-robustness questions.
They are especially important when the construct is not directly observable.
What this paper does not tell us
The Acta Psychologica study does not identify a person's “true emotion.”
The affective measures are text-derived proxies.
VADER-derived valence captures surface-level linguistic affect. MAV and RMS reduce cancellation, but they do not directly measure mixed emotional experience.
The data are English-language relationship narratives from a single online context.
The analyses are correlational.
The five expressive profiles reported in the paper are descriptive textual configurations, not psychological types or diagnoses.
These boundaries matter because the paper's methodological implication depends on keeping the distinction between representation and experience visible.
If that distinction disappears, a proxy becomes a state.
A score becomes emotion.
And a model's success at predicting the score begins to look like access to the person.
The next study I would like to see
The current result raises questions that this dataset cannot answer.
I am increasingly interested in designs where the same underlying episode can be represented in several ways at once.
A participant's first-person report.
An observer's judgment.
Language-derived affect.
Physiological measures.
Repeated measurements across time.
Potentially, model-generated interpretations as well.
That kind of design would allow us to ask a more useful question than “Which measure is correct?”
We could ask which substantive conclusions are stable across representations, which depend on a particular measurement choice, and where disagreements between measures contain scientifically meaningful information.
This is also where I think collaboration could be particularly valuable.
I would be interested in working with researchers who have access to:
- multilingual narrative or conversational corpora with first-person affective reports;
- longitudinal or experience-sampling data where affect changes across time;
- multimodal datasets linking language with physiology, behaviour, voice, or facial movement;
- emotion-recognition benchmarks in which the same observations can be re-analysed under alternative target constructions.
The goal would not be to declare one representation the ground truth.
It would be to map the conditions under which an inference survives measurement.
That is a different standard of robustness.
And I suspect it will become increasingly important as emotion AI moves from research environments into systems that act on people.
When measurement becomes intervention
There is a final step in this argument.
An affective score inside a paper is a representation.
An affective score that changes what happens next is also an intervention.
If a system infers that a student is frustrated and changes the lesson, the score has entered a decision loop.
If a workplace system infers disengagement and alerts a manager, the score has acquired institutional consequences.
If an AI companion interprets vulnerability and changes its language, the inference is shaping the environment from which the person's next emotional response will emerge.
At that point, measurement and interpretive authority begin to meet.
The more consequential the output becomes, the less acceptable it is for the construction of its target to remain implicit.
A high-performing model does not acquire epistemic authority simply because it performs well.
The legitimacy of the inference still depends on what was measured, how it was measured, whose judgment established the reference, what information was lost, and what the output is allowed to do.
What I took away from the paper
When I began this analysis, I expected the paper to be about whether narrative complexity and affective intensity were coupled.
The answer turned out to be modest.
They are weakly coupled under some representations and almost uncoupled under another.
But the smallness of the effect redirected the question.
The important issue was no longer simply whether narrative structure and affect were related.
It became:
How much of the relationship we observe is a property of the phenomenon, and how much is a property of the measurement that made the phenomenon visible?
That question now seems more durable to me than any individual coefficient.
We routinely ask whether a result is robust to another model.
For affective science and emotion AI, I think we should ask the parallel question just as seriously.
Is the inference robust to another defensible measure?
Research discussed
Ryan SangBaek Kim (2026). Narrative Complexity Is Weakly Coupled With Affective Energy, Not Net Valence, in Relationship Narratives. Acta Psychologica, 270, 107689.
DOI: https://doi.org/10.1016/j.actpsy.2026.107689
Earlier analysis of the same ANAD corpus:
Ryan SangBaek Kim (2026). Narrative-Affect Discrepancy as a Regulated Degree of Freedom in 351,734 Relationship Narratives. PLOS ONE, 21(5), e0348715.
DOI: https://doi.org/10.1371/journal.pone.0348715
Open derived-feature research resource:
ANEST Narrative-Affect Representations (ANAD v1), Version 1.4.0
DOI: https://doi.org/10.5281/zenodo.22135369
Selected references
Barrett, L. F., Adolphs, R., Marsella, S., Martinez, A. M., & Pollak, S. D. (2019). Emotional Expressions Reconsidered: Challenges to Inferring Emotion From Human Facial Movements. Psychological Science in the Public Interest, 20(1), 1-68.
https://doi.org/10.1177/1529100619832930
Jacobs, A. Z., & Wallach, H. (2021). Measurement and Fairness. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 375-385.
https://doi.org/10.1145/3442188.3445901
Mohammad, S. M. (2022). Ethics Sheet for Automatic Emotion Recognition and Sentiment Analysis. Computational Linguistics, 48(2), 239-278.
https://doi.org/10.1162/coli_a_00433
D'Amelio, T. A., et al. (2025). Emotion Recognition Systems With Electrodermal Activity: From Affective Science to Affective Computing. Neurocomputing, 651, 130831.
https://doi.org/10.1016/j.neucom.2025.130831
For potential research collaboration on measurement robustness, multimodal affective inference, or cross-linguistic replication:
Ryan SangBaek Kim
Ryan Research Institute (RRI), Paris
ryan@ryanresearch.org