Legion Health: The AI Decides Whether a Psychiatric Refill Reaches a Clinician
Evidence-Based Responsibility Reconstruction℠ Case Study
Legion Health is not building a clinical tool for physicians. It is building a clinic whose operating architecture assumes that work performed through software should eventually be executable by an LLM.
Legion describes itself as putting its technology identity ahead of its care-delivery identity. Its internal rule is that “anything a human can do, an LLM has to be able to read and eventually do as well.”
Utah has now authorized one consequential step along that path. Legion’s AI may determine whether a patient qualifies for a psychiatric medication renewal, authorize the renewal, and transmit it to a pharmacy without a clinician reviewing the individual decision first.
This case applies Evidence-Based Responsibility Reconstruction℠ to the signed Utah agreement, Legion’s proposal, the state’s public description of the pilot, Legion’s Terms of Service, and a partnered company interview describing its AI architecture.
The responsibility problem is not hidden:
The AI does not merely decide whether medication should be renewed. It decides whether the patient’s case reaches a clinician at all.
“Clinician in the Loop” Does Not Mean Case-Level Review
Human Review Declines as the Pilot Advances
| Phase | Human involvement | Advancement condition |
|---|---|---|
| Phase A | A Utah-licensed clinician reviews the first 250 requests before completion. | Reported concordance must exceed 98%. |
| Phase B | The AI may send renewals before the next 1,000 requests receive retrospective review. | Reported concordance must exceed 99%. |
| Phase C | Approximately 5%–10% of cases receive monthly sampling, with additional incident audits. | Ongoing reporting and oversight. |
During Phase A, a clinician can stop an individual renewal before it reaches the pharmacy.
During Phase B, the review occurs after the AI has acted. During Phase C, most individual decisions are not scheduled for clinician review at all.
The phrase “clinician in the loop” therefore describes the operating system as a whole. It does not establish that a clinician participated in a particular patient’s renewal.
A Physician’s Name Is Not Evidence of Physician Judgment
For an automated renewal, the named physician may not have:
- Seen the patient’s answers.
- Reviewed the refill packet.
- Evaluated the safety flags.
- Approved the individual renewal.
- Had an opportunity to intervene before transmission.
The name makes the prescription attributable to a licensed professional. It does not prove that the professional formed or adopted the judgment reflected in the prescription.
Legion’s design exposes the difference between having a human name on the outcome and having evidence that a human made the decision.
The AI Decides When Human Judgment Begins
Safety Screens Control Access to a Clinician
Legion’s proposal establishes hard stops for conditions including:
- Suicidality or self-harm risk.
- Mania or hypomania.
- Pregnancy-related changes.
- New or worsening symptoms.
- Loss of medication efficacy.
- Severe adverse effects.
- Contraindications or prescription discrepancies.
If the system detects one of those conditions, encounters ambiguity, or cannot verify the patient or prescription, the request enters a clinician queue. Patients and pharmacists may also request escalation.
If the system detects no qualifying condition, the renewal proceeds.
| When the system detects risk or ambiguity | When the system does not detect risk or ambiguity |
|---|---|
| The case is sent for clinician review. | The AI may authorize and transmit the renewal. |
| A clinician can examine the chat log and medical record. | A clinician may never examine the individual case. |
| The escalation creates a review record. | The absence of an escalation may be the only recorded safety result. |
The human safeguard begins only after the automated system decides that human judgment is necessary.
A Deterministic Rule Does Not Make Detection Deterministic
The agreement says each hard stop is tested through deterministic unit tests and synthetic red-team scripts. Those tests can establish that a known trigger produces the expected action.
They do not establish that the system will correctly recognize every real patient statement that should activate the trigger.
A rule can reliably escalate a positive suicidality flag while the language system still fails to recognize that a patient’s words indicate suicidality. If the flag is never generated, the deterministic safeguard never begins.
Naming a missed escalation would identify the technical failure. It would not establish who designed the threshold, who validated it, who determined its performance was sufficient, or who accepted the risk of letting it control access to a clinician.
The Concordance Percentage Does Not Measure Every Disagreement Equally
Legion Uses a Safety-Weighted Calculation
The agreement classifies clinician comparisons as follows:
| AI decision | Clinician decision | Classification | Effect on advancement target |
|---|---|---|---|
| Refill | Refill | Concordant | Included as agreement |
| Refill | Escalate | Discordant | Counts against the target |
| Escalate | Escalate | Concordant | Included as agreement |
| Escalate | Refill | Discordant | Does not count against the target |
Excluding unnecessary escalations from the advancement penalty creates a conservative safety preference. That is a defensible design choice.
It also means that “98% concordance” or “99% concordance” is not ordinary agreement accuracy. One category of disagreement is omitted from the performance threshold.
The Headline Percentage Cannot Stand Alone
Questions the Percentage Cannot Answer
A decision-maker needs the full matrix:
- How often did the AI renew when the clinician would have escalated?
- How often did the AI escalate when the clinician would have renewed?
- Which disagreements counted against advancement?
- Were clinicians reviewing independently, or did they see the AI’s decision and rationale first?
- Did the same disagreement pattern concentrate around a medication, symptom, model, or prompt version?
The percentage describes a scoring rule. It does not by itself establish that the AI and clinicians agreed in 98% or 99% of all reviewed cases.
Responsibility Is Split Across the Clinic, the Platform, and the State
The Agreement Combines Entities That the Terms Separate
Legion’s Terms of Service distinguish between two organizations:
| Entity | Publicly described role |
|---|---|
| Legion Health Inc. | Operates the platform and provides administrative and management services. It says it does not provide medical care or control clinical judgment. |
| Legion Health PA | Provides medical services through licensed clinicians. |
The Utah agreement takes a stronger approach by naming Legion Health Inc. and Legion Health PA jointly as the “Participant.” Both are therefore parties to the state agreement. Jonathan Kole, M.D., signed for Legion Health PA, and Yash Patel signed for Legion Health Inc.
That prevents the pilot from being presented to Utah as the responsibility of only one side of the corporate structure. It still does not establish which entity performed each act after a specific failure.
The Current Public Terms Describe a Different Prescription Sequence
Legion’s Terms of Service, last updated in October 2025, say a prescription requires a provider consultation, a provider determination that the medication is appropriate, and a prescription written by that provider. The Utah agreement signed in March 2026 allows an AI-authorized renewal under a named prescriber who may not interact with the patient.
The terms predate the pilot and may be revised before patients participate. As currently published, however, they do not explain the individual authority Utah has authorized the AI to exercise. Legion Health Terms of Service
Utah Remains in the Responsibility Chain
Utah granted regulatory relief that makes the automated renewal structure possible. The agreement says Utah does not endorse Legion, requires Legion to protect the state from claims, and preserves legally available remedies for patients and third parties.
Those provisions allocate legal and financial positions after harm. They do not erase Utah’s authorizing act.
Utah decided that a physician could supply professional authority through an approved AI protocol without reviewing every prescription. Responsibility Reconstruction must retain that decision in the record.
Legion Creates Evidence—and Plans to Purge Some of It
The Refill Packet Creates Traceability
Every AI decision is supposed to produce a structured refill packet containing:
- Patient inputs.
- Rule checks.
- Safety flags.
- A decision rationale.
Legion must report volumes, dispositions, escalation reasons, clinician agreement, complaints, adverse outcomes, safety signals, and selected case excerpts to Utah. This is a stronger evidence design than a system that produces only a final prescription.
The packet may help establish what the system recorded and why it reported making the decision.
It is not necessarily the complete evidence needed to reconstruct the decision.
The Audit Record Has an Expiration Problem
What Legion Plans to Preserve
The proposal says clinical records will follow Legion’s standard medical-record retention policy. Pilot-specific audit artifacts will be retained through the pilot and Utah’s closeout, then moved to reduced retention and purged.
The agreement does not state a precise public retention period for those audit artifacts.
What Utah Receives
Utah’s monthly reports are treated as protected records. The state ordinarily receives de-identified data and redacted excerpts rather than complete chat logs. Legion therefore remains central to holding, selecting, explaining, and preserving much of the evidence.
| Evidence the agreement creates | What issues remain uncertain after harm |
|---|---|
| Structured refill packet | Whether it preserves the original inputs or only the system’s processed account of them |
| Model decision and rationale | Whether the exact model and version can be identified later |
| Safety flags and rule checks | Whether missed signals can be reconstructed from the complete interaction |
| Clinician comparison | Whether the clinician formed an independent judgment |
| Monthly state report | Whether Utah received full records or selected excerpts |
| Pilot audit artifacts | Whether they still exist when a later injury or pattern becomes visible |
Traceability exists only as long as the necessary record survives.
What the Public Record Establishes—and What Remains Unresolved
| What the public record establishes | What issues remain unresolved |
|---|---|
| Utah authorized Legion’s AI to determine eligibility and complete qualifying psychiatric renewals. | Which individual approved the AI’s decision thresholds and determined they were sufficient. |
| A named prescriber may supply authority without reviewing the patient’s individual case. | What the named physician knew, reviewed, or could have stopped. |
| Human review declines from pre-issuance review to retrospective review and sampling. | Who authorized each transition and what complete evidence supported it. |
| Safety flags determine whether the patient reaches a clinician. | The system’s false-negative rate for suicidality, mania, pregnancy changes, adverse effects, and other risks. |
| Legion’s broader architecture uses frontier models, prompts, workflows, APIs, and vendor systems. | Which model and version power the Utah workflow and how substitutions or vendor updates are governed. |
| The concordance calculation excludes one category of disagreement from the advancement penalty. | Whether decision-makers receive the full matrix rather than only the headline percentage. |
| Legion Health Inc. and Legion Health PA are jointly bound by the agreement. | Which entity designed, operated, monitored, changed, and preserved each component. |
| Every AI decision generates a refill packet. | Whether the surviving packet and underlying records are sufficient for an independent reconstruction. |
| Utah receives monthly reports and may request further information. | What evidence Utah possesses independently of Legion’s reporting and classification. |
| Legally available patient remedies remain intact. | Whether an injured patient can obtain the technical and institutional evidence needed to prove the responsibility chain. |
| Pilot audit artifacts are eventually moved to reduced retention and purged. | Who determines when responsibility evidence is no longer worth preserving. |
These are the questions that matter after a patient is harmed. “AI error,” “prompt failure,” “physician oversight,” or “missed escalation” would describe possible mechanisms. None would establish the responsibility chain.
Responsibility Reconstruction Finding
Legion designed a clinic around progressively transferring software-mediated work from humans to LLMs.
Utah authorized the AI to determine whether a psychiatric medication renewal can proceed and whether the patient’s circumstances require clinician involvement.
A physician’s name remains attached to the prescription. That name does not establish that the physician reviewed or adopted the individual decision.
The safety process contains deterministic rules, but those rules depend on the system first recognizing the condition that should activate them.
The performance thresholds use a concordance calculation that does not count every disagreement equally.
The broader Legion architecture can use models from multiple outside providers, while the public agreement does not identify the model and version governing the refill workflow.
Legion creates a structured operating record, but some pilot-specific evidence is scheduled for reduced retention and eventual deletion.
Utah’s pilot therefore preserves a familiar human name on the prescription while permitting the clinical act beneath that name to be performed by a changing AI system.
If a patient is harmed, the presence of a physician’s name will make the prescription attributable. It will not prove that a physician made the decision. Responsibility Reconstruction must determine who designed the decision, who authorized the system, who validated its controls, who could have intervened, whether anyone did, and what evidence survives to prove each connection.
Naming the error is not owning responsibility.