A state workforce agency was placing 41 percent of program entrants into jobs, and only 58 percent of those placements survived six months. Stratenity helped redesign the program so an AI model ranked candidate-to-opening matches while a human caseworker made every placement decision. The model surfaced options; the caseworker owned the call and recorded the reason. Six-month placement retention rose from 58 to 74 percent, time-to-first-placement fell from 47 to 29 days, and every decision carried a written rationale a caseworker could defend to an auditor or an appeals board. Automation did the ranking. People kept the judgment.
A placement engine that placed people who did not stay
The agency ran a state workforce development program serving roughly 12,000 job seekers a year across 22 local offices, funded through a mix of federal formula grants and state appropriation. Its headline metric, the one the legislature watched, was the entered-employment rate. On paper it was climbing. Underneath, two numbers told a harder story. Only 41 percent of program entrants were placed into a job at all, and of those, just 58 percent were still employed at the six-month checkpoint. The program was moving people into openings that did not fit, and the churn was invisible in the top-line number because a placement counted the day it happened, not the day it failed.
The caseworkers were not the problem. Each carried a load of 90 to 110 active cases and matched candidates to openings out of a spreadsheet of listings, ranked mostly by recency and geography. There was no systematic read on which openings a given candidate was actually likely to hold. Senior caseworkers had a feel for it, but that feel lived in their heads and left when they retired. The executive team had circled the same question for three budget cycles: could the agency use an AI model to match candidates to openings without turning placement into a black box that a caseworker could not explain to an appeals board or an auditor. The engagement was scoped to answer that with a working operating model, not a pilot memo.
The model ranks, the caseworker decides, the reason is recorded
The design rule was fixed on day one and never moved. The AI model would rank matches and explain why it ranked them. It would never place anyone. Every placement decision stayed with a named human caseworker, who could accept the model's top option, choose a lower-ranked one, or reject the list entirely, and who had to record a short reason for the choice. That reason became the audit trail. It meant the agency could defend any single placement to a regulator or an appeals board by showing both what the model saw and what the human decided, which is the standard a public program is held to.
The model was trained on five years of the agency's own outcome history, weighting features that predicted six-month retention rather than same-day placement. The team ran it in shadow mode against 3,000 historical cases before any caseworker saw it, to confirm it did not encode the biases already sitting in the data.
The operating model was expressed as a five-stage sequence, and the discipline of the design was locating the model precisely where it added value and locking it out everywhere it created liability. The team wrote the stage table below as the contract between the automation and the humans who operated it, so that a new caseworker or an external reviewer could read in one glance who owned each decision and what the model was and was not permitted to touch.
| Stage | Who decides | What the AI does | Human checkpoint |
|---|---|---|---|
| Intake and profiling | Caseworker | Extracts skills and constraints from intake notes, flags gaps | Caseworker confirms the profile before it enters the model |
| Match ranking | Model proposes, human disposes | Ranks live openings by predicted six-month retention, shows top drivers | Caseworker reviews ranked list with reasons |
| Placement decision | Caseworker only | Nothing; the model is read-only here | Caseworker records the chosen opening and a one-line rationale |
| Fairness review | Program integrity lead | Reports selection rates by demographic group weekly | Integrity lead signs off or pauses the model |
| Retention follow-up | Caseworker | Flags placements trending toward early exit at day 30 and day 90 | Caseworker triggers a retention touch or a re-match |
The last two rows were the ones the executive team debated hardest. The fairness review gave a single accountable person the authority to pause the model if selection rates diverged across groups, and it fired weekly, not quarterly. The retention follow-up turned placement from a one-time event into a tracked relationship, which is where the six-month number actually moved.
Retention up, time down, every call defensible
- Six-month placement retention rose from 58 percent to 74 percent across the first two cohorts, measured against a held-out baseline of offices that adopted later.
- Median time-to-first-placement fell from 47 days to 29 days, because caseworkers stopped hand-searching listings and started from a ranked short list.
- Overall placement rate climbed from 41 percent to 49 percent within two quarters, without loosening any eligibility standard.
- Every placement carried a recorded human rationale, so a sample audit of 200 cases by the state auditor closed with zero unexplained decisions.
- Weekly fairness review caught and corrected one feature that was quietly penalizing candidates with employment gaps, before it affected a full cohort.
What the agency learned it would keep
- Deciding on day one that the model would never place anyone removed the entire governance fight; the debate became about ranking quality, not authority.
- Training the model on six-month retention rather than same-day placement changed which openings it favored, and that single reframing drove most of the gain.
- Requiring a one-line rationale per placement felt like friction and turned out to be the audit defense; caseworkers stopped fearing the model once they, not it, were on record.
- Weekly fairness reporting caught a bias the quarterly review would have missed for a full cohort, so cadence beat depth.
- The retention follow-up, not the matching, was where the durable outcome lived; matching gets someone hired, follow-up keeps them there.
How to run this in another public program
- Fix the rule before you build: name the exact decision the model may never make, and put a human on record for it.
- Train and evaluate against the outcome that actually matters (retention, not the event), and run shadow mode on historical cases before any live use.
- Require a short written rationale at the human decision point; that record is your audit and appeals defense.
- Stand up a weekly fairness review with one accountable owner who can pause the model, not a committee that meets quarterly.
- Instrument the follow-up, not just the placement, and give caseworkers a trigger to re-match when a placement trends toward early exit.