Summary

A consumer goods operator was running two rival forecasts, a statistical model that ignored context and planner spreadsheets that ignored history, and neither won consistently. Stratenity built one governed forecast that fed external demand signals into the model and gave planners a structured override with a logged reason. Over two quarters, weighted forecast accuracy rose from 68 to 82 percent, finished-goods inventory fell 11 percent, and case-fill improved to 97.4 percent. The win came from making judgment auditable rather than replacing it, so the planner became an accountable input to a system the client still owns.

Context

Two forecasts, no owner

The operator ran a mid-market portfolio of roughly 640 active SKUs across food and household categories, sold through grocery, club, and a growing direct channel. Planning was split. A statistical engine produced a baseline forecast every Monday from 36 months of shipment history. Then eleven demand planners overrode it in spreadsheets, adjusting for promotions, weather, competitor activity, and whatever a key account had said on a call. The two numbers rarely agreed, and when they diverged nobody could reconstruct why. Weighted forecast accuracy sat at 68 percent at the four-week horizon, which sounds tolerable until you translate it into the two failure modes it produced at once.

The first was overstock. Slow movers accumulated in three regional distribution centers, tying up about 4.2 million dollars in working capital and driving markdown writeoffs on dated inventory. The second was the opposite: hero SKUs stocked out during promotions, so case-fill on the fastest movers dropped to 91 percent in peak weeks and two national accounts issued chargebacks. Finance blamed the planners for gaming numbers to protect service. The planners blamed the model for ignoring everything they knew. Both were partly right, and the argument had run for six quarters without a resolution because there was no single forecast anyone was accountable for.

The approach

One governed forecast with a logged override

Stratenity did not try to build a better model in isolation. The decision the engagement was scoped to make was narrower and more useful: produce one forecast per SKU per week that carried its own reasoning, so that model math and human judgment became inputs to the same governed artifact rather than competing outputs. External signals were pulled in as features the model could weight, and every planner adjustment had to be entered as a typed override with a reason code and a magnitude cap. Nothing shipped to the supply plan without provenance attached.

The build followed a deliberate sequence rather than a big-bang cutover. The first four weeks instrumented the baseline and locked the accuracy measurement so later gains could be defended against the original numbers, not a moving target.

StageWhat was builtGovernance gateMeasured result
Baseline lockFrozen 36-month accuracy measurement at SKU and category level, MAPE and bias split by horizonMethod signed off by Finance and Supply before any changeConfirmed 68 percent weighted accuracy, plus 6 percent upward bias
Signal fusionExternal features added: promo calendar, regional weather, retail POS pull-through, category price indexEach feature backtested and approved before it could move a forecastModel-only accuracy rose to 75 percent, bias cut to 2 percent
Structured overrideTyped planner override with reason code, magnitude cap at 30 percent, and mandatory noteOverrides above cap route to a planning lead for approvalOverride hit rate tracked per planner, low performers coached
Weekly consensusSingle governed forecast with model, override, and final value visible side by sideLocked forecast versioned each Monday, never overwrittenCombined accuracy reached 82 percent at four weeks
Inventory replanSafety stock recalculated from the new bias and error bands per SKU tierReorder points changed only through the governed modelFinished-goods inventory down 11 percent, no service loss
HandoffCadence, override rules, and accuracy dashboard transferred to the client teamConsulting team removed from the weekly cycle at week 20Client ran two full quarters unassisted

The pivotal design choice was the override cap and the reason code. Planners kept their authority, but authority now left a trail. When an override improved accuracy, the reason code told the team which signals the model was still missing, and several of those reasons became new model features in the following cycle. When an override degraded accuracy, the log made that visible too, and the conversation shifted from blame to a coaching question about a specific decision.

Outcomes

What moved in six months

  • Weighted forecast accuracy at the four-week horizon rose from 68 to 82 percent, beating both the pure-model and pure-planner baselines the engagement had frozen at the start.
  • Finished-goods inventory fell 11 percent, releasing roughly 460 thousand dollars of working capital, with slow-mover markdown writeoffs down by about a third.
  • Case-fill on the top 80 SKUs improved to 97.4 percent, and the two national accounts closed their open chargeback cases within one quarter.
  • Systemic upward bias dropped from 6 percent to under 2 percent, so the supply plan stopped quietly padding demand across the whole book.
  • Override quality became measurable per planner, and the three strongest planners' reason codes seeded four new model features that raised the baseline for everyone.
Lessons

What we would tell the next operator

  • Do not choose between model and judgment. The accuracy gain came from fusing them inside one governed forecast, not from picking a winner in a debate that had no winner.
  • Lock the baseline measurement before you change anything. Every later claim was defensible only because the original 68 percent was frozen and signed off by Finance and Supply first.
  • Cap and log the override. An unbounded, unexplained override is indistinguishable from gaming; a capped, reason-coded one is a data source that improves the model over time.
  • Treat degraded overrides as coaching, not policing. The log makes bad adjustments visible, but the value is in the specific conversation it enables, not in the scoreboard.
  • Recalculate safety stock from the new error bands, or the accuracy gain never reaches working capital. Better forecasts only pay off when the inventory policy actually consumes them.
Replication checklist

Before you start

  • Freeze a weighted accuracy measurement, MAPE and bias, split by SKU tier and forecast horizon, and get Finance and Supply to sign the method.
  • Inventory your external signals: promo calendar, POS pull-through, weather, price index, and confirm each can be delivered weekly and backtested.
  • Define the override contract: reason codes, a magnitude cap, an approval route above the cap, and a rule that every forecast version is kept, never overwritten.
  • Instrument override quality per planner from day one so coaching and feature discovery have data instead of anecdotes.
  • Plan the inventory replan and the handoff as part of scope, with a named date to remove the consulting team from the weekly cycle.