A regional health plan was drowning in prior-authorization backlog and could not defend its denials to state regulators. We modernized authorization with explainable AI that routes clear-approve cases automatically and hands genuine gray-area cases to nurses with the policy citation attached. Median turnaround fell from 6.4 days to 1.9, the auto-approve rate reached 54 percent with zero clinical overrides missed, and every denial carried an audit-grade decision trail that held up in a state market-conduct review. Faster for members, defensible for the plan. This is explainable AI where the stakes are real.
A backlog the plan could not defend
A regional health plan covering about 480,000 members processed roughly 12,000 prior-authorization requests a month across imaging, specialty drugs, and elective procedures. The process was manual, and it showed. Median turnaround was 6.4 days, the backlog routinely exceeded 3,000 open cases, and the appeals rate on denials had climbed to 19 percent. The deeper problem surfaced when the state department of insurance opened a market-conduct examination and asked the plan to reconstruct the reasoning behind a sample of denials. For many cases, the plan could not. The decision lived in a reviewer's judgment, unrecorded, and a denial the plan cannot explain is a denial the plan cannot defend.
The executive team had circled automation for several quarters and stalled each time on the same fear: an AI that denies care it should have approved is a clinical and legal catastrophe, not an efficiency gain. The Stratenity engagement was scoped around a decision the chief medical officer and the general counsel could both sign, which meant the design had to be faster for members and more defensible for the plan at the same time, not one at the expense of the other.
The starting evidence was uneven in a familiar way. The claims and authorization systems held clean transactional data, but the reasoning behind individual determinations lived in reviewers' heads and in scattered notes. The first two weeks assembled a shared evidence base: the published medical policies, the actual criteria reviewers applied in practice, and a sample of denials the examiners had flagged. Reading that sample together made the gap concrete. The plan was not making bad decisions; it was making unrecorded ones, and an unrecorded decision is indistinguishable from an arbitrary one to a regulator.
Automate the clear cases, cite the reasoning on every one
We split the authorization population by how clear the policy answer was, not by how much we wanted to automate. Where the clinical criteria were met unambiguously and the request matched published medical policy, the system could auto-approve, because approving care that meets criteria carries no member harm. Where criteria were partially met, ambiguous, or the request fell outside policy, the case routed to a licensed nurse reviewer with the relevant policy section and the matching and missing criteria already assembled. The model never issued a denial on its own. Denials were a human decision, and every decision, automated approval or nurse determination, wrote a structured trail: the policy version, the criteria evaluated, the evidence cited, and the reviewer identity.
| Case pattern | Share of volume | System action | Human control |
|---|---|---|---|
| Criteria fully met, in policy | 54% | Auto-approve with trail | Sampled audit |
| Criteria partially met | 21% | Route with gap analysis | Nurse determination |
| Outside published policy | 14% | Route with policy citation | Nurse or medical director |
| Missing clinical data | 8% | Auto-request records | Reviewer on return |
| Any potential denial | All denials | Assemble reasoning only | Human issues decision |
A worked case shows the split. A request for an advanced MRI arrived with documentation of six weeks of conservative therapy, the exact threshold the medical policy required. The model matched every criterion, cited the policy section and the chart notes proving the therapy timeline, and auto-approved in under two minutes with the full trail recorded. A second request, for the same scan but with only three weeks documented, did not auto-approve. It routed to a nurse with the four-week gap flagged and the policy citation attached. The nurse requested the missing records rather than denying, and the member got a decision in a day instead of a week. Neither outcome was a black box. The second case matters more than the first, because the plan's old fear was that automation would deny the three-week request outright to clear the queue. The design made that outcome structurally impossible: the model could route and could request records, but it could not deny, so the only paths available were approval or human review.
We built the operating cadence before scaling past imaging. A weekly review sampled auto-approvals for accuracy, tracked the nurse override rate on routed cases, and watched the appeals trend as a lagging safety signal. A single medical director owned the auto-approve criteria with named authority to tighten a rule the moment a sampled case looked wrong, so governance was a standing function rather than a project deliverable that expired when the engagement closed.
Faster for members, defensible for the plan
- Median authorization turnaround fell from 6.4 days to 1.9 days within the first quarter of operation.
- The auto-approve rate reached 54 percent of volume, with no case in the sampled audit found to have been approved outside criteria.
- The open backlog fell below 600 cases, from a standing level above 3,000, as clear approvals stopped queuing behind gray-area reviews.
- Every denial carried an audit-grade decision trail; when the market-conduct examination requested a fresh sample, the plan reconstructed the full reasoning for each in minutes.
- The appeals rate on denials fell from 19 percent to 11 percent, because members received the specific policy criterion and missing evidence rather than an opaque denial.
What the engagement taught
- Segment by clarity of the answer, not by ambition to automate. Auto-approving clear cases is safe; automating denials is not, and the line between them is the design.
- Never let the model deny. Assembling the reasoning for a human is a support tool; issuing the adverse decision is a governance failure waiting for a regulator.
- The decision trail is the product, not a byproduct. A market-conduct examination turns an unrecorded judgment into an indefensible one overnight.
- Explainability lowered appeals as much as it satisfied regulators, because a member who sees the missing criterion can supply it rather than contest the whole decision.
- Auto-requesting missing records converted a class of would-be denials into approvals, which was both better care and fewer appeals.
How to run this in your plan
- Classify your authorization volume by how unambiguously each case meets published policy, and automate only the clear-approve segment.
- Route every ambiguous or out-of-policy case to a licensed reviewer with the policy citation and criteria gap already assembled.
- Forbid the model from issuing denials; have it prepare the reasoning and let a human make the adverse determination.
- Write a structured decision trail on every case, policy version, criteria, evidence, and reviewer, queryable for any regulatory sample.
- Track turnaround, auto-approve accuracy, and appeals rate against a protected baseline, and expand automation only where accuracy holds under audit.