A mid-cap commercial bank had sixteen AI pilots running and nothing in production. Each business line had funded its own experiment, so the same fraud model was being built three times while shared data and model risk governance sat unowned. The engagement stopped starting new pilots and instead sequenced a platform program: shared data foundation first, then a governed model risk pipeline, then business use cases riding on top. Within three quarters the bank had four use cases in production and a repeatable path from idea to deployment, rather than sixteen demos that never met a regulator.
Sixteen pilots, four business lines, zero in production
The bank was a roughly 22 billion dollar commercial lender with a retail arm, a commercial and industrial book, a treasury services unit, and a wealth division. Each of the four had responded to board pressure on AI the same way: by funding its own pilot. Sixteen were live when the engagement began, spread across fraud detection, credit decisioning support, document extraction, and customer service. Not one had reached production. The commercial and retail teams were independently building near-identical transaction fraud models on overlapping data. Nobody owned the shared data plumbing, and model risk management, the function that a bank regulator will always ask about first, had been treated as a step to clear at the end rather than a design constraint at the start.
The cost was not only wasted spend, though the duplicated fraud work alone was around 1.4 million dollars a year. The deeper cost was that the bank had no answer when its regulators asked how models were validated, monitored, and governed. The engagement was scoped to convert a scatter of pilots into a sequenced program with a single readiness posture, and to get a first cohort of use cases from idea to supervised production without tripping a supervisory finding. The board wanted proof that AI could scale inside the bank's risk appetite; the examiners wanted proof that any model touching a lending or fraud decision was inventoried, validated, and monitored. A credible program had to satisfy both audiences at once, which meant governance could not be the thing the bank got to last.
Foundation first, use cases second
We inverted the order the bank had been working in. Instead of pushing sixteen use cases forward in parallel, we froze new pilots and sequenced the program so that the shared foundations were built before anything scaled on top of them. That freeze was the hardest conversation of the engagement, because every business line believed its pilot was the exception that deserved to keep moving. The readiness assessment gave the argument a common language: rather than debating whose pilot mattered most, the leadership team looked at where the bank as a whole was weakest and sequenced against that. The assessment scored the bank across six dimensions, and the table below shows the baseline, the target, and the first move for each.
| Readiness dimension | Baseline score | Target | First move |
|---|---|---|---|
| Data foundation | 2.1 / 5 | 3.5 | Single governed feature store for shared fraud and credit data |
| Model risk governance | 1.8 / 5 | 4.0 | Validation and monitoring pipeline aligned to SR 11-7 before deployment |
| Talent and operating model | 2.4 / 5 | 3.5 | Central platform team plus embedded owners per business line |
| Use case portfolio | 2.9 / 5 | 3.8 | Consolidate sixteen pilots to four sequenced use cases |
| Infrastructure and MLOps | 2.2 / 5 | 3.5 | Standard deployment and rollback path shared across lines |
| Regulatory and controls | 2.0 / 5 | 4.2 | Model inventory and audit trail queryable by examiners |
The three highest-leverage moves were the shared feature store, the governed model risk pipeline, and the consolidation of sixteen pilots into four. The credit and fraud teams stopped building their own data pipelines and drew from one governed source, which meant a feature engineered once could be reused and its lineage traced in a single place. Model risk validation became a gate every use case passed through, with monitoring thresholds, challenger models, and rollback triggers defined before deployment rather than reconstructed under pressure after an examiner asked. That reordering is what turned demos into a deployable program. It also changed the operating model: a small central platform team owned the shared feature store and the model risk pipeline, while each business line kept an embedded owner accountable for its own use cases. Neither extreme would have worked. Pure central ownership starved the use cases of business context, and pure local ownership was exactly how the bank had ended up building the same fraud model three times in the first place.
What the program produced
- Four use cases reached supervised production within three quarters: transaction fraud scoring, commercial document extraction, credit memo drafting support, and service triage.
- The duplicated fraud modeling work collapsed into one shared model, retiring around 1.4 million dollars a year of parallel build cost.
- Model risk governance moved from 1.8 to 3.9 on the readiness scale, and the model inventory became queryable by examiners rather than reconstructed for each exam.
- Fraud scoring on the shared model lifted caught-fraud value by 18 percent while cutting false-positive review volume by roughly a third.
- Time from approved use case to supervised production fell from an open-ended average that had never once completed to a predictable eleven-week path through the newly built shared pipeline, so the twelve shelved pilots could be re-evaluated against the new foundation rather than restarted from scratch.
What transferred beyond this engagement
- In a regulated institution, model risk governance is a design constraint, not a closing step. Build the validation and monitoring path before the use case, or the use case never ships.
- Parallel pilots across business lines duplicate the expensive parts and starve the shared parts. Consolidate to a sequenced portfolio before you scale.
- The data foundation is the real bottleneck. Use cases built on ungoverned, duplicated data cannot be defended to a regulator or trusted by a risk committee.
- A central platform team plus embedded business owners beats either extreme. Pure central ownership starves context; pure local ownership rebuilds the same plumbing four times.
- Readiness is measurable. Scoring each dimension against a target turns a vague AI ambition into a sequenced program a board and an examiner can both follow.
How to run this pattern yourself
- Inventory every live and shelved AI pilot across business lines and map where they duplicate data, models, or infrastructure.
- Score readiness across data, model risk, talent, portfolio, MLOps, and regulatory controls against explicit target levels.
- Freeze new pilots and sequence the program so the shared feature store and governed model risk pipeline are built before use cases scale.
- Make model validation, monitoring, and rollback a gate every use case passes through, with an inventory queryable by examiners.
- Consolidate to a small first cohort of sequenced use cases, ship them to supervised production, then re-evaluate the shelved pilots against the new foundation.