Services Case Studies Webinar Blog News & Press Events About Request a Briefing
Home  /  Case Studies  /  Legion Health
Evidence-Based Responsibility Reconstruction℠ for AI-Mediated Conduct

Legion Health: The AI Decides Whether a Psychiatric Refill Reaches a Clinician

AI use Psychiatric medication renewals
Human checkpoint AI-triggered escalation and phased review
Evidence source Refill packets, review samples, and provider reports
Reconstruction finding Attribution does not establish physician judgment

Evidence-Based Responsibility Reconstruction℠ Case Study

Legion Health is not building a clinical tool for physicians. It is building a clinic whose operating architecture assumes that work performed through software should eventually be executable by an LLM.

Legion describes itself as putting its technology identity ahead of its care-delivery identity. Its internal rule is that “anything a human can do, an LLM has to be able to read and eventually do as well.”

Utah has now authorized one consequential step along that path. Legion’s AI may determine whether a patient qualifies for a psychiatric medication renewal, authorize the renewal, and transmit it to a pharmacy without a clinician reviewing the individual decision first.

This case applies Evidence-Based Responsibility Reconstruction℠ to the signed Utah agreement, Legion’s proposal, the state’s public description of the pilot, Legion’s Terms of Service, and a partnered company interview describing its AI architecture.

The responsibility problem is not hidden:

The AI does not merely decide whether medication should be renewed. It decides whether the patient’s case reaches a clinician at all.

What Utah Authorized

The AI Can Complete the Renewal

The signed agreement authorizes Legion’s “Mental Health Refill Autopilot” to renew a limited group of existing, non-controlled psychiatric medications.

The system may:

  • Verify the patient and prescription.
  • Collect information about efficacy, side effects, allergies, clinical changes, and symptoms.
  • Screen for conditions requiring escalation.
  • Determine whether the request is eligible.
  • Authorize the renewal.
  • Transmit it to the patient’s pharmacy.

This is not an AI recommendation awaiting a physician’s decision. When the request falls within scope and the system detects no risk or ambiguity, the AI completes the renewal.

The Scope Is Narrow but the Decision Is Consequential

The pilot excludes controlled substances, benzodiazepines, antipsychotics, lithium, valproate, clozapine, new prescriptions, dose changes, medication switches, and renewals requiring new laboratory work or an ECG.

It includes commonly used medications for depression, anxiety, panic disorder, PTSD, OCD, insomnia, and related conditions. Automated renewals are limited to ten between provider reviews or six months, whichever occurs first.

Those boundaries reduce the number of eligible cases. They do not change who makes the decision inside those boundaries.

Patients and Pharmacists Must Be Told AI Is Acting

Before the AI engages with a Utah patient about a renewal, Legion must disclose that the patient is interacting with generative AI, explain how the patient’s information and correspondence will be used or shared, identify the temporary Utah program, and obtain an acknowledgement. Patients may request human review.

Pharmacists must be told that the renewal was generated by AI and receive a direct way to contact Legion and Utah.

Those are meaningful disclosures. They make the AI’s participation visible. They do not establish which model made the decision, who approved its operating rules, or who is responsible for the design. A patient’s acknowledgement authorizes participation; it does not transfer responsibility for the system.

Utah Created a Named-Prescriber Structure

Schedule A allows an affiliated professional entity or provider to rely on Legion’s AI “solely by being the named prescriber” without interacting directly with the patient or another healthcare provider.

Utah’s public explanation says each refill is approved by a licensed physician, either directly or “vicariously through the AI system’s protocol.”

That word—vicariously—carries the responsibility structure.

The physician’s professional authority remains attached to the prescription even when the physician did not review the individual renewal.

“Clinician in the Loop” Does Not Mean Case-Level Review

Human Review Declines as the Pilot Advances

PhaseHuman involvementAdvancement condition
Phase A A Utah-licensed clinician reviews the first 250 requests before completion. Reported concordance must exceed 98%.
Phase B The AI may send renewals before the next 1,000 requests receive retrospective review. Reported concordance must exceed 99%.
Phase C Approximately 5%–10% of cases receive monthly sampling, with additional incident audits. Ongoing reporting and oversight.

During Phase A, a clinician can stop an individual renewal before it reaches the pharmacy.

During Phase B, the review occurs after the AI has acted. During Phase C, most individual decisions are not scheduled for clinician review at all.

The phrase “clinician in the loop” therefore describes the operating system as a whole. It does not establish that a clinician participated in a particular patient’s renewal.

A Physician’s Name Is Not Evidence of Physician Judgment

For an automated renewal, the named physician may not have:

  • Seen the patient’s answers.
  • Reviewed the refill packet.
  • Evaluated the safety flags.
  • Approved the individual renewal.
  • Had an opportunity to intervene before transmission.

The name makes the prescription attributable to a licensed professional. It does not prove that the professional formed or adopted the judgment reflected in the prescription.

Legion’s design exposes the difference between having a human name on the outcome and having evidence that a human made the decision.

The AI Decides When Human Judgment Begins

Safety Screens Control Access to a Clinician

Legion’s proposal establishes hard stops for conditions including:

  • Suicidality or self-harm risk.
  • Mania or hypomania.
  • Pregnancy-related changes.
  • New or worsening symptoms.
  • Loss of medication efficacy.
  • Severe adverse effects.
  • Contraindications or prescription discrepancies.

If the system detects one of those conditions, encounters ambiguity, or cannot verify the patient or prescription, the request enters a clinician queue. Patients and pharmacists may also request escalation.

If the system detects no qualifying condition, the renewal proceeds.

When the system detects risk or ambiguityWhen the system does not detect risk or ambiguity
The case is sent for clinician review. The AI may authorize and transmit the renewal.
A clinician can examine the chat log and medical record. A clinician may never examine the individual case.
The escalation creates a review record. The absence of an escalation may be the only recorded safety result.

The human safeguard begins only after the automated system decides that human judgment is necessary.

A Deterministic Rule Does Not Make Detection Deterministic

The agreement says each hard stop is tested through deterministic unit tests and synthetic red-team scripts. Those tests can establish that a known trigger produces the expected action.

They do not establish that the system will correctly recognize every real patient statement that should activate the trigger.

A rule can reliably escalate a positive suicidality flag while the language system still fails to recognize that a patient’s words indicate suicidality. If the flag is never generated, the deterministic safeguard never begins.

Naming a missed escalation would identify the technical failure. It would not establish who designed the threshold, who validated it, who determined its performance was sufficient, or who accepted the risk of letting it control access to a clinician.

Utah Authorized a System, Not a Stable Model

Legion’s Broader Architecture Uses Multiple Frontier Models

In a November 2025 interview, Legion said its broader AI stack uses frontier models from OpenAI, Google, and Anthropic and switches among them as needed. The company said most of its performance comes from prompts and workflow design rather than a custom-trained model.

Legion also described an API-first architecture connecting LLMs to scheduling, billing, clinical records, patient support, and clinical workflows. The interview was presented in partnership with Legion, so it is evidence of the company’s stated design and strategy—not independent validation of its performance. Legion Health interview

The agreement does not identify which model or model version powers the Utah refill workflow. It does not establish whether the workflow uses one provider consistently or can move among providers.

The Regulated System Has Multiple Moving Parts

The operating system is not merely “the AI.” It includes:

model + model version + prompts + deterministic rails + workflow logic + APIs + clinical data + escalation thresholds + vendor systems

What the Agreement Allows to Change

The agreement allows deterministic-rail updates and prompt or template revisions informed by internal quality review. Formal amendments to the approved proposal require written approval from the parties.

What the Agreement Does Not Answer

The public record does not clearly establish:

  • Whether replacing the underlying model counts as an amendment.
  • Whether every replacement model must repeat Phase A.
  • Who decides that two models are sufficiently equivalent.
  • How silent updates by a model vendor are identified.
  • How model-specific errors are separated from prompt, workflow, or data errors.
  • Whether Utah receives notice before or after a material system change.

If the model changes while the named workflow remains the same, the regulated label may remain stable while the behavior beneath it changes.

The Concordance Percentage Does Not Measure Every Disagreement Equally

Legion Uses a Safety-Weighted Calculation

The agreement classifies clinician comparisons as follows:

AI decisionClinician decisionClassificationEffect on advancement target
Refill Refill Concordant Included as agreement
Refill Escalate Discordant Counts against the target
Escalate Escalate Concordant Included as agreement
Escalate Refill Discordant Does not count against the target

Excluding unnecessary escalations from the advancement penalty creates a conservative safety preference. That is a defensible design choice.

It also means that “98% concordance” or “99% concordance” is not ordinary agreement accuracy. One category of disagreement is omitted from the performance threshold.

The Headline Percentage Cannot Stand Alone

Questions the Percentage Cannot Answer

A decision-maker needs the full matrix:

  • How often did the AI renew when the clinician would have escalated?
  • How often did the AI escalate when the clinician would have renewed?
  • Which disagreements counted against advancement?
  • Were clinicians reviewing independently, or did they see the AI’s decision and rationale first?
  • Did the same disagreement pattern concentrate around a medication, symptom, model, or prompt version?

The percentage describes a scoring rule. It does not by itself establish that the AI and clinicians agreed in 98% or 99% of all reviewed cases.

Responsibility Is Split Across the Clinic, the Platform, and the State

The Agreement Combines Entities That the Terms Separate

Legion’s Terms of Service distinguish between two organizations:

EntityPublicly described role
Legion Health Inc. Operates the platform and provides administrative and management services. It says it does not provide medical care or control clinical judgment.
Legion Health PA Provides medical services through licensed clinicians.

The Utah agreement takes a stronger approach by naming Legion Health Inc. and Legion Health PA jointly as the “Participant.” Both are therefore parties to the state agreement. Jonathan Kole, M.D., signed for Legion Health PA, and Yash Patel signed for Legion Health Inc.

That prevents the pilot from being presented to Utah as the responsibility of only one side of the corporate structure. It still does not establish which entity performed each act after a specific failure.

The Current Public Terms Describe a Different Prescription Sequence

Legion’s Terms of Service, last updated in October 2025, say a prescription requires a provider consultation, a provider determination that the medication is appropriate, and a prescription written by that provider. The Utah agreement signed in March 2026 allows an AI-authorized renewal under a named prescriber who may not interact with the patient.

The terms predate the pilot and may be revised before patients participate. As currently published, however, they do not explain the individual authority Utah has authorized the AI to exercise. Legion Health Terms of Service

Utah Remains in the Responsibility Chain

Utah granted regulatory relief that makes the automated renewal structure possible. The agreement says Utah does not endorse Legion, requires Legion to protect the state from claims, and preserves legally available remedies for patients and third parties.

Those provisions allocate legal and financial positions after harm. They do not erase Utah’s authorizing act.

Utah decided that a physician could supply professional authority through an approved AI protocol without reviewing every prescription. Responsibility Reconstruction must retain that decision in the record.

Legion Creates Evidence—and Plans to Purge Some of It

The Refill Packet Creates Traceability

Every AI decision is supposed to produce a structured refill packet containing:

  • Patient inputs.
  • Rule checks.
  • Safety flags.
  • A decision rationale.

Legion must report volumes, dispositions, escalation reasons, clinician agreement, complaints, adverse outcomes, safety signals, and selected case excerpts to Utah. This is a stronger evidence design than a system that produces only a final prescription.

The packet may help establish what the system recorded and why it reported making the decision.

It is not necessarily the complete evidence needed to reconstruct the decision.

The Audit Record Has an Expiration Problem

What Legion Plans to Preserve

The proposal says clinical records will follow Legion’s standard medical-record retention policy. Pilot-specific audit artifacts will be retained through the pilot and Utah’s closeout, then moved to reduced retention and purged.

The agreement does not state a precise public retention period for those audit artifacts.

What Utah Receives

Utah’s monthly reports are treated as protected records. The state ordinarily receives de-identified data and redacted excerpts rather than complete chat logs. Legion therefore remains central to holding, selecting, explaining, and preserving much of the evidence.

Evidence the agreement createsWhat issues remain uncertain after harm
Structured refill packet Whether it preserves the original inputs or only the system’s processed account of them
Model decision and rationale Whether the exact model and version can be identified later
Safety flags and rule checks Whether missed signals can be reconstructed from the complete interaction
Clinician comparison Whether the clinician formed an independent judgment
Monthly state report Whether Utah received full records or selected excerpts
Pilot audit artifacts Whether they still exist when a later injury or pattern becomes visible

Traceability exists only as long as the necessary record survives.

What the Public Record Establishes—and What Remains Unresolved

What the public record establishesWhat issues remain unresolved
Utah authorized Legion’s AI to determine eligibility and complete qualifying psychiatric renewals. Which individual approved the AI’s decision thresholds and determined they were sufficient.
A named prescriber may supply authority without reviewing the patient’s individual case. What the named physician knew, reviewed, or could have stopped.
Human review declines from pre-issuance review to retrospective review and sampling. Who authorized each transition and what complete evidence supported it.
Safety flags determine whether the patient reaches a clinician. The system’s false-negative rate for suicidality, mania, pregnancy changes, adverse effects, and other risks.
Legion’s broader architecture uses frontier models, prompts, workflows, APIs, and vendor systems. Which model and version power the Utah workflow and how substitutions or vendor updates are governed.
The concordance calculation excludes one category of disagreement from the advancement penalty. Whether decision-makers receive the full matrix rather than only the headline percentage.
Legion Health Inc. and Legion Health PA are jointly bound by the agreement. Which entity designed, operated, monitored, changed, and preserved each component.
Every AI decision generates a refill packet. Whether the surviving packet and underlying records are sufficient for an independent reconstruction.
Utah receives monthly reports and may request further information. What evidence Utah possesses independently of Legion’s reporting and classification.
Legally available patient remedies remain intact. Whether an injured patient can obtain the technical and institutional evidence needed to prove the responsibility chain.
Pilot audit artifacts are eventually moved to reduced retention and purged. Who determines when responsibility evidence is no longer worth preserving.

These are the questions that matter after a patient is harmed. “AI error,” “prompt failure,” “physician oversight,” or “missed escalation” would describe possible mechanisms. None would establish the responsibility chain.

Responsibility Reconstruction Finding

Legion designed a clinic around progressively transferring software-mediated work from humans to LLMs.

Utah authorized the AI to determine whether a psychiatric medication renewal can proceed and whether the patient’s circumstances require clinician involvement.

A physician’s name remains attached to the prescription. That name does not establish that the physician reviewed or adopted the individual decision.

The safety process contains deterministic rules, but those rules depend on the system first recognizing the condition that should activate them.

The performance thresholds use a concordance calculation that does not count every disagreement equally.

The broader Legion architecture can use models from multiple outside providers, while the public agreement does not identify the model and version governing the refill workflow.

Legion creates a structured operating record, but some pilot-specific evidence is scheduled for reduced retention and eventual deletion.

Utah’s pilot therefore preserves a familiar human name on the prescription while permitting the clinical act beneath that name to be performed by a changing AI system.

If a patient is harmed, the presence of a physician’s name will make the prescription attributable. It will not prove that a physician made the decision. Responsibility Reconstruction must determine who designed the decision, who authorized the system, who validated its controls, who could have intervened, whether anyone did, and what evidence survives to prove each connection.

Naming the error is not owning responsibility.