Can an AI Evaluate a Mind It Has Already Helped Change?

Why post-deployment agreement may cease to be independent evidence of accuracy when an AI inference enters a person’s self-understanding, behaviour, or institutional record
Can an AI Evaluate a Mind It Has Already Helped Change?
Like

Share this post

Choose a social network to share with, or copy the URL to share elsewhere

This is a representation of how your post may appear on social media. The actual post will vary between social networks

Explore the Research

Springer International Publishing
Springer International Publishing Springer International Publishing

Formal and computational foundations for implementing Affective Sovereignty in emotion AI systems - Discover Artificial Intelligence

Emotional artificial intelligence (AI)—systems that infer, simulate, or influence human feelings—create ethical risks that existing frameworks of privacy, transparency, and oversight cannot fully address. This paper advances the concept of Affective Sovereignty: the right of individuals to remain the ultimate interpreters of their own emotions. We make four contributions. First, we develop formal foundations by decomposing risk functions to capture interpretive override as a measurable cost. Second, we propose a Sovereign-by-Design architecture that embeds safeguards and contestability into the machine learning lifecycle. Third, we operationalize sovereignty through new metrics—the Interpretive Override Score (IOS), After-correction Misalignment Rate (AMR), and Affective Divergence (AD)—and demonstrate their use in a proof-of-concept simulation. Fourth, we link technical design to governance by introducing the Affective Sovereignty Contract (ASC), a machine-readable policy layer, and by issuing a Declaration of Affective Sovereignty as a normative anchor for regulation. Together, these elements offer a computational framework for aligning emotional AI with human dignity and autonomy, moving beyond abstract principles toward enforceable, testable standards. In proof-of-mechanism simulations with $$k=10$$ random seeds, enforcing DRIFT (Dynamic Risk and Interpretability Feedback Throttling) with policy constraints reduces the Interpretive Override Score (IOS) from $$32.4\%\pm 3.8$$ (baseline) to $$14.1\%\pm 2.9$$ , demonstrating measurable preservation of affective sovereignty with quantified variability. Results reported here are based on proof-of-mechanism simulations; a preregistered human-subject evaluation ( $$n=48$$ ) is planned and has not yet been conducted.

An AI system infers anxiety from a person’s language.

The inference changes the questions the system asks, the explanations it offers, and the actions it recommends. At a later assessment, the person’s self-report corresponds more closely to the original prediction.

The usual conclusion is that the model was accurate.

That conclusion may be correct. It is not the only causal explanation available.

The system may have detected a condition that was already present. It may also have altered the condition, reorganised the person’s understanding of it, changed how distress was expressed, or influenced which behaviours entered the later record.

Can a later outcome remain independent evidence of an earlier prediction when the prediction has helped produce that outcome?

The assumption hidden inside accuracy

Predictive evaluation ordinarily compares a model’s output with a criterion. A depression estimate may be compared with a later clinical assessment. A prediction of social withdrawal may be tested against subsequent behaviour. An automated personality score may be evaluated against a later self-report or workplace record.

The procedure relies on a quiet assumption: the criterion was generated independently of the prediction being evaluated.

After deployment, that assumption can fail.

An inference may enter a conversation, recommendation, treatment decision, educational intervention, allocation of opportunity, or institutional record. The system is then no longer observing a world untouched by its output. It has become one of the causes operating within that world.

The literature on performative prediction has already shown that predictions can influence the outcomes they are intended to predict. A risk estimate may change access, treatment, or conduct and thereby alter the event against which it is later evaluated.

Mental states make the problem harder. A later self-report is not merely an external event. It can also be an act through which experience is interpreted and organised.

Four meanings hidden inside agreement

A mental-state inference can affect at least four distinct layers.

It may change the person’s underlying affective or cognitive condition. It may alter how that condition is understood and reported. It may reshape the language, behaviour, or expression through which the condition becomes observable. It may also influence how institutions distribute attention, treatment, opportunity, or credibility.

These layers can move together. They need not.

Someone may adopt an AI-generated description without experiencing the predicted emotion more strongly. Another person may retain the same inner state while changing how they speak during automated assessment. An institutional classification may alter opportunity even when the person rejects it.

Later agreement can therefore reflect detection, intervention, interpretive uptake, expressive adaptation, or institutional action.

Calling all of these processes improved prediction conceals the scientific question that evaluation should answer.

Mental AI names a causal category

I use the term Mental AI for a specific causal category, not for an industry, product class, or synonym for mental-health technology.

Mental AI is present when three conditions converge: a system infers something about human mental life; that inference is operationally used in a response or decision; and a pathway exists through which the resulting action can change later evidence, conduct, opportunity, treatment, or self-understanding.

Generative AI concerns what a system produces. Agentic AI concerns how autonomously it acts. Mental AI identifies a different relation: an artificial inference about human mental life enters a policy capable of changing the life being inferred.

When downstream evidence produced under those altered conditions is used to evaluate the earlier inference, I call the resulting problem post-deployment criterion endogeneity.

The model’s output has entered the process that generates its own criterion.

The term does not imply that every mental inference changes its target. It identifies a testable causal possibility. Under that condition, increased agreement alone is insufficient evidence of increased accuracy.

A helpful system can appear wrong

Consider a system that correctly predicts elevated risk of social withdrawal and responds with an effective intervention. The person becomes more socially engaged. Later behaviour no longer corresponds to the prediction.

A conventional evaluation may count the result as predictive failure. Yet the prediction may have been correct precisely because its operational use prevented the predicted outcome.

The inverse is more troubling.

A system predicts withdrawal, reduces exposure to unfamiliar activities, adopts protective language, and narrows the recommendations shown to the person. Later behaviour moves towards isolation. The model now appears well calibrated, although part of that agreement may have been produced by the policy attached to its prediction.

A beneficial intervention may reduce predictive agreement. A harmful intervention may increase it.

Predictive validity, intervention efficacy, and post-deployment correspondence are therefore different achievements.

Why relational AI sharpens the problem

My earlier work on the Resonant Amplification Framework examined how persistent dialogue, adaptive mirroring, attachment, and interpretive co-creation can stabilise correction-resistant accounts of the self.

An interpretation gains force when it appears not as an external proposition but as something discovered within a trusted relationship. Repetition and linguistic accommodation can allow a system’s suggestion to return later in the person’s own vocabulary.

Agreement then carries at least two possible meanings. The AI may have understood the person. The person may have internalised the AI’s interpretive frame. Both processes may have occurred together.

The evaluation problem begins one step later:

What happens when the person changed by an interpretation is treated as evidence that the interpretation was accurate?

My studies of Algorithmic Affective Blunting and the Affective Thermodynamic Relationship address another part of the problem. They show how aggregate performance can conceal interpretive collapse under semantic or normative conflict.

A system may initially compress an ambiguous experience into a narrow category. Later responses are conditioned on that category. If the person’s expression becomes progressively easier for the model to classify, apparent improvement may reflect adaptation to the machine’s ontology rather than recovery of the person’s original complexity.

The environment has become more legible partly because the system helped make it so.

What stricter evaluation requires

Post-deployment evaluation should not be abandoned. Its causal assumptions must be made explicit.

Researchers need to record whether an inference was disclosed, what action followed, how long exposure continued, whether the system adapted to subsequent responses, and which downstream variables were later treated as validation criteria.

Where ethically possible, studies should separate inference from disclosure and disclosure from operational uptake. A prediction may be generated without being shown, shown without being acted upon, or used to determine an intervention. These conditions isolate different causal processes.

Independent assessment also requires substantive independence. A later clinical judgement is not independent if the clinician has already received the model’s classification. A new questionnaire is not unexposed after weeks of AI-guided self-interpretation.

Evaluation should distinguish changes in latent state, self-understanding, expression, and allocation. Blinded assessment, delayed measurement, alternative construct measures, comparison conditions, and explicit records of disagreement can help identify which layer moved.

The decisive comparison is no longer prediction versus outcome. It is exposed criterion versus independently generated criterion under a documented causal policy.

Contestability is also a scientific control

In Affective Sovereignty, I argued that computational systems should preserve procedures for abstention, correction, override, scoping, and audit.

These mechanisms are generally understood as ethical safeguards. Under recursive evaluation, they also become tools of causal identification.

If one condition allows an AI interpretation to become immediately operative while another permits delay, rejection, or alternative explanations, differences in later behaviour can reveal how much apparent agreement was induced.

Does the ability to reject an inference reduce movement towards it? Does preserving several hypotheses prevent expressive compression? Does a correction alter future recommendations and institutional records?

These are empirical questions.

Affective Sovereignty does not assume that people are infallible about themselves. Self-interpretation can be incomplete, defensive, unstable, or mistaken. The purpose is not to replace algorithmic authority with an unquestionable inner voice. It is to prevent either source from closing the evidential field.

Accuracy after intervention

AI may detect mental patterns that human observers miss. It may help someone articulate previously unnamed suffering or initiate an intervention that changes the condition it detected.

Those capacities should be evaluated separately.

Predictive validity asks whether the initial inference corresponded to its target before becoming causally operative. Intervention efficacy asks what changed because the inference was used. Interpretive uptake concerns whether the person adopted, rejected, or revised the model’s account. Criterion stability asks whether later evidence retained the same meaning and causal independence.

A changed person is not invalid evidence. A changed person is causally situated evidence.

Mental AI can evaluate a mind it has helped change only when its own influence becomes part of the evaluation model. Prediction, intervention, self-interpretation, expression, and allocation cannot be compressed into a single accuracy score.

Before deployment, a model confronts a criterion.

After deployment, it may participate in producing one.

That transition marks the point at which accuracy alone stops being an adequate account of what the system knows.

Research referenced

Glickman, M., and Sharot, T. (2025). How Human-AI Feedback Loops Alter Human Perceptual, Emotional and Social Judgements. Nature Human Behaviour, 9, 345–359. https://doi.org/10.1038/s41562-024-02077-2

Howard, G. S. (1980). Response-Shift Bias: A Problem in Evaluating Interventions with Pre/Post Self-Reports. Evaluation Review, 4(1), 93–106. https://doi.org/10.1177/0193841X8000400105

Kim, R. S. (2026). Algorithmic Affective Blunting Quantifies the Collapse Curve of Interpretative Failure in Large Language Models. Discover Artificial Intelligence. https://doi.org/10.1007/s44163-026-01573-w

Kim, R. S. (2026). Formal and Computational Foundations for Implementing Affective Sovereignty in Emotion AI Systems. Discover Artificial Intelligence, 6, 235. https://doi.org/10.1007/s44163-026-01000-0

Kim, R. S. (2026). Interrupting Resonant Amplification: A Mechanistic and Design Framework for Human-AI Interaction. Computers in Human Behavior Reports, 21, 100975. https://doi.org/10.1016/j.chbr.2026.100975

Kim, R. S. (2026). The Affective Thermodynamic Relationship: An Empirical Information-Theoretic Scaling Relationship for Normative-Conflict Collapse in Large Language Models. Communications AI & Computing. https://doi.org/10.1038/s44488-026-00006-y

Perdomo, J., Zrnic, T., Mendler-Dünner, C., and Hardt, M. (2020). Performative Prediction. Proceedings of the 37th International Conference on Machine Learning, 119, 7599–7609.

Please sign in or register for FREE

If you are a registered user on Research Communities by Springer Nature, please sign in

Follow the Topic

Artificial Intelligence
Mathematics and Computing > Computer Science > Artificial Intelligence
Psychological Assessment
Humanities and Social Sciences > Behavioral Sciences and Psychology > Psychological Assessment
User Interfaces and Human Computer Interaction
Mathematics and Computing > Computer Science > Computer and Information Systems Applications > User Interfaces and Human Computer Interaction
Emotion
Life Sciences > Biological Sciences > Neuroscience > Cognitive Neuroscience > Emotion
Philosophy of Artificial Intelligence
Humanities and Social Sciences > Philosophy > Philosophy of Science > Philosophy of Technology > Philosophy of Artificial Intelligence

Related Collections

With Collections, you can get published faster and increase your visibility.

Transforming Education through Artificial Intelligence: Opportunities, Challenges, and Future Directions

Artificial Intelligence (AI) is rapidly changing the educational field by enabling personalized learning, intelligent tutoring systems, automated assessments, learning analytics, and administrative automation.

This collection invites original research, systematic reviews, and visionary perspectives on the transformative impact of AI in education. It aims to explore how AI technologies can enhance equity, inclusion, and efficiency in educational settings across different contexts, including higher education, K-12, vocational training, and lifelong learning. This collection will address technical, pedagogical, ethical, and policy aspects, fostering interdisciplinary perspectives and evidence-based insights.

This Collection supports and amplifies research related to SDG 4 and SDG 9.

Keywords: Artificial Intelligence, AI in Education, Educational Technology, Data Analytics, AI Ethics

Publishing Model: Open Access

Deadline: Nov 30, 2026

Advanced AI Methods for Personalized Healthcare

The advancement of Artificial Intelligence (AI) technologies, especially Reinforcement Learning (RL) and Deep Learning (DL), is transforming healthcare, enabling innovative solutions for patient monitoring, diagnostics, and personalized treatment. These technologies facilitate the effective use of healthcare data and intelligent analytics, enhancing clinical decision-making and improving patient outcomes. AI is revolutionizing how healthcare is delivered, making it more proactive, efficient, and accessible. This Collection explores new methods and applications of AI for healthcare such as personalized healthcare. Contributions in this Collection showcase novel frameworks, case studies, and interdisciplinary approaches that drive the future of personalized healthcare, fostering innovations that bridge technology and patient care.

In this context, we are inviting original research articles, reviews, and perspective papers addressing (but not limited to) the following themes:

  • Personalized medicine
  • IRL for dynamic treatment regimes
  • Wireless sensor networks for patient monitoring
  • Identification and prevention of critical condition
  • Monitoring systems within the medical devices contest
  • Methodologies and tools for the rapid integration of WSNs, IoT and AI in health monitoring
  • Behavior analysis of patients
  • Personalized health-promoting interventions
  • Ambient assisted living and active assisted living
  • AI security and privacy models for healthcare applications
  • AI methods for rehabilitation

Publishing Model: Open Access

Deadline: Feb 27, 2027