Beyond the Demo: Rewiring Colorectal Cancer Care with Multi-Agent Systems

A model in a demo is not an agent on a ward. After five posts on building the data foundation, here is the harder question: what actually changes when multi-agent systems run in colorectal cancer care — and what stubbornly refuses to change.

Published in Computational Sciences and Surgery

Share this post

Choose a social network to share with, or copy the URL to share elsewhere

This is a representation of how your post may appear on social media. The actual post will vary between social networks

A Model in a Demo Is Not an Agent on a Ward

Across the first five posts of this series, I argued for a deliberate order: a research-ready cohort (Post 1), a multi-agent architecture (Post 2), a governance layer (Post 3), a standard data set (Post 4), and finally data elementization (Post 5). That arc built the instrument. This post is about something different and harder: what changes at the bedside once the instrument is actually running.

Here is the uncomfortable observation that motivates this piece. Most institutions that "adopted AI" changed nothing. A model was purchased, a pilot was announced, a press release was issued — and clinical behavior on Monday morning looked exactly as it did before. The failure was almost never algorithmic. It was that a tool had been added, but no workflow had been rewired.

Our group has now moved from architecture to deployment. What we run is not a single model but a working multi-agent platform for colorectal cancer care, composed of registered software modules: a clinical decision-support platform for rectal cancer, a postoperative intelligent follow-up system, and an explainable-AI engine (software copyright registration 2025SR0278495). Two specialized sub-engines developed with our group have been recognized nationally — a neoadjuvant-response prediction engine (first prize, 3rd National Simulation Innovation Application Competition) and an imaging-service platform built on heterogeneous data fusion. Together with DACCA, a 5,000+ case cohort refreshed daily, the agents reason over the same governed standard described in Posts 3–4.

Deployment Is Workflow Rewiring, Not Tool Stacking

The distinction I want to draw is simple but decisive:

A tool is something you add to your workflow. A deployed agent changes what the workflow is.

Adding a tool leaves the clinician doing the same cognitive labor, only now with an extra screen. Rewiring means some part of the clinical work is now performed differently — because an agent reliably does something that previously consumed a surgeon's scarce attention. That is the only definition of "landed" I find defensible.

Three places in colorectal cancer care have rewired for us, and they are worth examining honestly — including what did not change.

Three Places Where Agents Now Do Real Work

1. Assisting decisions: the sphincter-preservation judgment

Few judgments in colorectal surgery are as consequential, or as uncomfortable, as deciding whether a very low rectal cancer can be resected with sphincter preservation — or whether the patient must leave the operating room with a permanent stoma. The decision integrates tumor height, sphincter anatomy, response to neoadjuvant therapy, functional baseline, and the patient's own values. It is exactly the kind of judgment that is easy to make badly when it is made hurriedly.

Our decision-support agent does not make this call. What it does is assemble the multi-modal preoperative picture — imaging, endoscopy, pathology, and molecular markers — against the governed standard, and produce a recommendation with the rationale attached: which features drove it, which guideline clause it rests on, and where the evidence is thin. The surgeon then decides, with the reasoning visible rather than implicit.

Critically, we did not rush to claim benefit. The work now underway is a two-stage validation of this agent-assisted decision pathway — retrospective first, then prospective — precisely because a surgical decision aid that changes whether someone keeps their sphincter deserves more than a retrospective AUC. In our hands, sphincter preservation for ultra-low rectal cancer has reached 85.3%, against an industry average around 60%; the agent's job is to make that judgment reproducible rather than dependent on who is on call.

2. Automating follow-up: from episodic visits to continuous care

In Post 2, I described a Follow-up Agent that turns a transient clinic encounter into a continuous, personalized one. That is no longer a design sketch. Low anterior resection syndrome (LARS) — the bowel-dysfunction constellation that follows sphincter-preserving surgery — is the clearest test case, because it is common, under-discussed, and deeply quality-of-life-altering.

We built LARS management as a closed loop: predict → assess → individualized intervention. The follow-up agent carries the longitudinal risk map forward between visits, flags deteriorating trajectories early, and triggers the intervention tier appropriate to that specific patient. Where a conventional follow-up is episodic and generic — blood test, imaging, see you in three months — the agent-assisted visit arrives already informed.

With 73.2% of our patients achieving good-to-excellent LARS outcomes under this loop, I would argue the number matters less than its shape: outcomes improved not because a model was clever, but because care became continuous instead of intermittent.

3. Research archiving: turning care into cohort

The third rewiring is the least visible and, in my view, the most structurally important. The Research Archive Agent extracts variables from the clinical record and maintains the research cohort automatically, rather than requiring a separate abstraction step performed by a human reading charts after hours.

This closes a loop I opened in Post 1, where a pre-standardization audit found up to 30% source disagreement on key postoperative variables. If the caring and the recording are the same act, that disagreement largely disappears. The practical consequence is cultural as much as technical: clinical work produces research data as a by-product, not as unpaid overtime. When participation in research costs the clinician nothing extra, cohorts stop being a burden and become infrastructure.

The Entry Ticket: Explainability

There is one property that governs whether any of this is admissible in surgery, and it is not accuracy.

A black box can be tolerated in recommendation engines. It cannot be tolerated where a recommendation leads to an operation. Every agent output in our platform is required to be reconstructable to its inputs — traceable to specific report spans, specific structured fields, specific guideline clauses — the same traceability discipline described in Post 3. Where evidence is insufficient, the system's correct output is "this cannot be determined from the available data." We registered an explainable-AI system as part of the platform precisely because explainability is not a feature we can add later; it is the condition of entry.

What Stubbornly Refuses to Change

It would be easy to end this post in technological triumph. That would be dishonest.

After deployment, here is what has not changed. The decision is still made by a surgeon. The responsibility for it is still non-transferable. The difficult conversation with a patient about a possible stoma is still had by a human being, sitting down, looking them in the eye. And the two numbers I quoted — 85.3% sphincter preservation, 73.2% good-to-excellent LARS outcomes — are not abstractions. They are whether a person can live without a bag, and whether they can leave the house without planning their day around a toilet.

What agents changed is not who is accountable. It is how much of a surgeon's scarce attention is spent reconstructing the past rather than judging the present. That is the only trade I consider worth making.

What's Next

The two-stage validation now underway will report whether agent-assisted decision-making holds prospectively, which is the question that actually matters. Beyond that lies multi-center deployment and extending the same loop across the full care continuum. I hope to report results rather than architecture.


Selected Readings & Resources

  1. Wang XD, et al. Value-Based Colorectal Cancer Standard Data Set [以价值医疗为导向的结直肠癌标准数据集]. ISBN 978-7-5727-0816-9. — the governed standard the deployed agents reason over.
  1. Wang XD. Two-stage validation of a multi-agent assisted decision system for sphincter-preserving surgery in ultra-low rectal cancer (West China Clinical Research Fund, ongoing). — the prospective validation now running.
  1. Registered platform components: explainable-AI engine (software copyright 2025SR0278495); LARS trajectory prediction model patent (ZL202311684024.9).

Join the Conversation

I'd like to learn from others deploying clinical agents:

  • Where has deployment actually changed clinical behavior in your setting — and where did it quietly fail to?
  • For decision aids near an irreversible outcome (a stoma, a resection): what level of validation do you require before routine use?
  • Does your institution let clinical documentation double as research capture? If not, what blocks it?

Comment below or reach out directly. We should be comparing notes on the unglamorous part — the rewiring, not the demo.


The views expressed in this post are the author's own and do not represent the position of any affiliated institution.