A 340-person engineering org spent eight weeks each quarter setting OKRs and still shipped late, because reliability and platform work kept losing to feature outcomes in one ranking. Stratenity split goal-setting into two tracks: outcome OKRs that move a customer metric, and system OKRs that protect the engine that ships them. Each track got its own budget, owner, and review. Quarterly planning fell from eight weeks to eleven days, on-call pages dropped 44 percent, and reliability work finally stopped being the thing that got cut when features ran late. Two categories, separated cleanly at the start of every quarter.
Reliability kept losing to features in the same ranking
The organization was a 340-person engineering group inside a mid-size SaaS company, structured into 28 squads across product, platform, and infrastructure. Every quarter it ran a single unified OKR process: squads drafted objectives, leadership stack-ranked them into one list, and the list became the quarter. The process took eight weeks of the thirteen-week quarter, which meant nearly two of every three months were spent either planning the next quarter or recovering from planning the last one. Worse, the output did not hold. By week six, roughly 40 percent of committed key results had been quietly abandoned as features slipped and teams reprioritized on the fly.
The deeper failure was structural. Feature outcomes and reliability work were ranked against each other in the same list, and features always won, because a feature has a demo and a reliability investment has an absence of incidents. Platform and infrastructure squads watched their commitments get cut every quarter to make room for feature work that ran late. Technical debt compounded, on-call pages climbed, and the same argument, whether to invest in the engine or the output, consumed executive attention every planning cycle. The leadership team had circled this for three quarters. The engagement was scoped to redesign the cadence so it produced decisions instead of a ranked wish list that fell apart by mid-quarter.
There was a cultural cost underneath the operational one. Every quarter that platform commitments were cut, the message to the infrastructure and reliability engineers was that their work was optional, a reserve to be spent when feature teams overran. Attrition in those squads ran higher than anywhere else in the org, and the people who left were precisely the ones who understood the systems well enough to keep them from failing. The engagement was told, quietly, that fixing the cadence was also a retention problem, because the planning process itself was teaching the most valuable engineers that their work did not count.
Split the two categories before the ranking, not after
The central move was to separate goal-setting into two tracks at the very start, before anything was ranked. Outcome OKRs were defined as goals that move a customer or business metric this quarter: activation, conversion, latency a user feels. System OKRs were defined as goals that protect or improve the engine that ships outcomes: reliability, platform capability, developer throughput, debt reduction. Each track got its own protected budget of engineering capacity, its own accountable owner, and its own review. The two never competed in the same list again, which is what had let reliability lose every time.
The capacity split was set deliberately at 70 percent outcome and 30 percent system, and it was a floor, not a suggestion. A squad could not raid the system budget to rescue a slipping feature. The two reviews ran on different rhythms because the two kinds of work move at different speeds.
Making the split real meant defining each track precisely enough that a squad could not smuggle feature work into the system budget or the reverse. The team wrote a two-column definition, reproduced below, that every squad lead used to classify its own goals before the planning meeting, which removed the reclassification arguments that had eaten days in the old process. When a proposed goal did not clearly move a customer metric this quarter, it went to the system track by default, and the burden was on the squad to argue it out rather than in.
| Dimension | Outcome OKRs | System OKRs | Why they differ |
|---|---|---|---|
| What it moves | A customer or business metric this quarter | The engine that ships outcomes | Different time horizons, different proof |
| Capacity budget | 70 percent, protected floor | 30 percent, protected floor | Neither can raid the other |
| Owner | Product line lead | Platform and reliability lead | Single accountable owner per track |
| Review cadence | Biweekly outcome review | Monthly system review | Reliability compounds slower than features |
| Success signal | Metric moved, measured against baseline | Pages down, lead time down, debt retired | Absence of incidents is a real result |
| Failure mode handled | Overcommitting to demos | Getting cut when features slip | Separation removes the competition |
The biweekly outcome review kept feature work honest against real metrics, and the monthly system review finally gave platform work a forum where its progress was visible and defensible on its own terms, rather than as a line item that features could delete.
Faster planning, fewer pages, reliability that survived the quarter
- Quarterly planning time fell from eight weeks to eleven days, because the two tracks planned in parallel and neither had to be reconciled against the other.
- On-call pages dropped 44 percent over two quarters as protected system capacity finally paid down the debt that generated them.
- Mid-quarter abandonment of committed key results fell from roughly 40 percent to 14 percent, because the plan was smaller and better protected.
- Deployment lead time improved 31 percent once developer-throughput system OKRs were funded on their own budget rather than borrowed against.
- The recurring engine-versus-output argument left the planning cycle entirely, freeing leadership attention for actual product bets.
What the org learned it would keep
- Splitting the two categories before ranking, not after, was the whole trick; once they never shared a list, reliability stopped losing by default.
- Protected capacity floors only work if they cannot be raided; the first time a floor bends for a slipping feature, it stops being a floor.
- Different review cadences matter: forcing slow-compounding system work into a biweekly feature review had been hiding its progress.
- A smaller, protected plan held far better than a large ambitious one; the 14 percent abandonment rate came from committing to less, not trying harder.
- Giving system work its own owner and forum made its results legible to leadership, which is what turned it from a cost into a visible investment.
How to run this in another engineering org
- Separate outcome goals from system goals at the goal-setting stage, before anything gets ranked against anything else.
- Give each track a protected capacity floor, a single owner, and its own review; make the floor non-negotiable.
- Set review cadences to match how each kind of work actually moves, not one shared rhythm for both.
- Define a real success signal for system work (pages, lead time, debt retired) so its absence-of-incidents result is measurable.
- Plan the two tracks in parallel and resist the urge to reconcile them into one list; the parallelism is what collapses the timeline.