A national bank saw its model inventory jump from 180 to 470 in two years as teams shipped AI and machine learning models faster than model risk management could validate them. The validation queue reached 14 months and examiners flagged the backlog. Stratenity rebuilt the model risk operating model around a four-tier risk classification, so scarce senior validators focused only on high-materiality models while lower-tier models cleared through standardized checks. The queue fell to under 4 months, no high-tier model shipped unvalidated, and the tiering held up under regulatory examination.
Model supply outran validation capacity
The client was a national bank with a model risk management function built for a slower world. Its validation process, its documentation standards, and its senior reviewer headcount had all been sized when the bank produced a predictable stream of credit, market, and capital models. Over two years the model inventory climbed from roughly 180 to 470, driven almost entirely by AI and machine learning models coming out of fraud, marketing, collections, and pricing teams that had learned to build and ship in weeks. The validation team had grown by three people in the same period, and every one of the new models carried a validation obligation the old process had no capacity to absorb.
The result was a queue. Median time from model submission to validated production status had stretched to 14 months. Business teams routed around the delay by labeling models as tools or analytics to avoid the validation gate, which meant the bank did not even have an accurate count of its own model risk. A regulatory examination flagged both the backlog and the workarounds, and gave the bank a defined window to demonstrate a credible remediation plan. The executive team had circled the problem for several quarters, alternating between demands to hire more validators and demands to slow the model teams down, neither of which was affordable within the examination window. Hiring senior validators takes a year to recruit and season, and slowing the model teams would have surrendered the fraud and pricing gains that justified the AI investment in the first place. Stratenity was scoped to produce the redesigned operating model and the governance the examiners would accept, not a staffing request.
Tier the models, match the effort to the materiality
The first two weeks rebuilt the inventory into a single defensible list, including the models that had been hidden as tools. That surfaced the core insight. The queue was not slow because every model was hard. It was slow because a low-materiality marketing propensity model went through the same heavyweight validation as a capital model, and the senior validators who should have been on the capital model were buried in propensity models. The redesign replaced the one-size-fits-all process with a four-tier classification that matched validation depth to model materiality and use, so the cost of validating a model tracked the consequence of the model being wrong.
| Tier | Example models | Validation depth | Reviewer | Target cycle |
|---|---|---|---|---|
| Tier 1 | Capital, credit loss, stress testing | Full independent validation plus board reporting | Senior validators | 10 weeks |
| Tier 2 | Fraud, AML, underwriting | Independent validation, standardized scope | Validators plus AI specialist | 6 weeks |
| Tier 3 | Pricing, collections, propensity | Standardized checklist plus challenger test | Analysts, senior sign-off | 3 weeks |
| Tier 4 | Internal tools, non-decision analytics | Registration and monitoring only | Owner attestation | 1 week |
Two design choices made the tiering hold. The first was that tier assignment was governed, not self-selected. A small committee set each model's tier against written criteria, which killed the incentive to mislabel a model as a tool. The second was that the AI and machine learning models got their own standardized validation scope covering data drift, explainability, and monitoring triggers, so validators were not inventing the test for each new model and could reuse the same evidence templates across dozens of similar submissions. A single accountable owner, the Chief Risk Officer, held decision rights over the tiering criteria, which meant the classification could not be quietly negotiated model by model. That concentration of authority was resented by business sponsors who wanted their models tiered lower, and it was the reason the framework survived examination. The examiners could see that depth of scrutiny rose with materiality and that the rise was enforced by governance rather than by goodwill, which is precisely what a supervisory review looks for.
What the redesigned operating model delivered
- Median validation cycle fell from 14 months to under 4 months across the full inventory, with Tier 1 capital models still receiving full independent validation.
- No Tier 1 or Tier 2 model reached production unvalidated during the remediation window, the outcome the examiners were watching most closely.
- The hidden inventory was recovered. Roughly 60 models previously labeled as tools were reclassified into governed tiers, giving the bank an accurate model risk count for the first time.
- Senior validator time on low-materiality models dropped by more than half, redeployed onto the Tier 1 and Tier 2 queue that actually carried regulatory and capital exposure.
- The examination closed with the tiering framework accepted as the bank's target-state operating model, converting a finding into an approved remediation.
What the engagement learned
- The backlog was a triage failure, not a capacity failure. Adding validators without tiering would have added cost and left the capital models still stuck behind propensity models.
- Governing the tier assignment was the whole game. The moment teams could not self-select their tier, the incentive to mislabel models as tools disappeared.
- A standardized AI validation scope turned model-by-model improvisation into a repeatable check, which is what let lower tiers clear in weeks rather than months.
- Two Tier 3 checklists were retired within a quarter when they produced no findings across dozens of models. Effort that catches nothing is not governance, it is theater.
- Concentrating tiering authority in one accountable owner was unpopular and non-negotiable, because a framework that can be renegotiated per model is not a framework.
How to run this pattern
- Rebuild a single honest inventory first, including the models hidden as tools. You cannot tier what you cannot see.
- Match validation depth to materiality, so scarce senior reviewers spend their time only where the risk and the regulator actually are.
- Govern the tier assignment centrally against written criteria, so no team can self-select a lighter path.
- Give AI and machine learning models a standardized validation scope covering drift, explainability, and monitoring, rather than improvising per model.
- Retire any control tier that produces no findings over a meaningful sample. Governance that catches nothing spends capacity without buying safety.