ERP go-live is not the finish line. It is when the real work starts. Too many organizations fall back on ad-hoc fixes and tacit knowledge held by a handful of exhausted heroes, and the moment those people move on, restore times climb and the promised benefits quietly leak away. Hypercare was meant to end, not become the permanent operating model. The alternative is a productized run-state: a service catalog, service levels, a clear RACI, incident-to-change discipline, and a steady release train. Get it right and you trade firefighting for dependable outcomes and freed capacity, not a thicker binder.
Why go-live is the start, not the finish
After go-live, many organizations fall back on ad-hoc fixes and tacit knowledge held by a handful of people who happened to survive the project. That is fragile. The moment those people take leave or move on, mean time to restore climbs, workarounds harden into shadow processes, and the promised benefits of the new ERP quietly leak away. Hypercare is meant to end, but too often it simply becomes the permanent operating model, staffed by exhausted heroes.
The alternative is a formal run-state: a service catalog, service level agreements, a RACI, and change governance that together stabilize operations and create capacity for improvement. Treating ERP operations as a productized service is not bureaucracy for its own sake. It is what converts a volatile post-go-live period into a predictable one, where a P1 outage has a known response path and a routine config change follows a pre-approved pattern instead of a debate. The goal is dependable outcomes with freed capacity, not a thicker binder.
The stakes are financial as much as operational. A large ERP program routinely costs tens of millions of dollars and years of effort, yet a meaningful share of that value is realized (or lost) in the run-state that follows go-live. An implementation delivered on time can still fail in operation if a single missed batch cut-off delays the financial close, or a broken interface silently drops orders for a week before anyone notices. Productizing operations is how the organization protects the investment it already made.
The run-state playbook
Design ERP operations as a product: define what is offered (services), how it is measured (SLAs and OLAs), who owns what (RACI), and how it evolves (change rhythm), then embed continuous enablement so the system improves as the business does. This is the same discipline that runs any mature IT service, applied to the process backbone of the enterprise rather than to a single application, which is why borrowing the vocabulary of service management (catalog, SLA, RACI, CAB) is a feature, not overhead. The table sets illustrative service levels and the KPI that proves each one is holding.
| Service | SLA target | Owner | KPI watched |
|---|---|---|---|
| Incident handling (P1) | 15-minute response, 4-hour restore | Service Manager | MTTA and MTTR |
| Integration and interface support | 99.5 percent interface uptime | Integration lead | Interface failure rate |
| Batch and job monitoring | Cut-off met; RPO 15 min, RTO 4 hr | Ops backbone | Missed-cut-off count |
| Master data stewardship | Fix turnaround within 2 business days | Data Steward lead | Data quality score |
| Change and release | Change failure rate below 5 percent | Release Manager | Release success rate |
Underneath the catalog sits the operating discipline: an incident-to-problem-to-change flow, and a steady release train of monthly maintenance drops and quarterly value releases, each with a regression pack and a rollback plan. Worked example: an industrial operator emerging from a chaotic go-live published this catalog, seeded a known-error knowledge base from hypercare, and stood up a monthly service review. Within one quarter, P1 recurrence fell by 40 percent and the change failure rate dropped from 14 percent to under 5 percent, because emergency changes gave way to pre-approved standard ones. Just as important, the monthly service review gave leadership a single place to see whether the run-state was holding: MTTA and MTTR trending down, recurrence rate falling, and release success climbing above 95 percent. Once those numbers were visible and owned, the temptation to route every fix through an emergency change, which is how backsliding usually starts, faded, because the dashboard made the cost of indiscipline obvious.
What CIOs and service managers should do first
- Publish a service catalog and SLAs within the first 30 days: response and restore targets by priority, interface uptime, batch cut-offs, and the KPIs (MTTA, MTTR, recurrence rate, change failure rate) that prove they hold.
- Assign accountable product owners per domain (Finance, Procure-to-Pay, Order-to-Cash, Manufacturing) and stand up the ops backbone of Service Manager, Release Manager, Problem Manager, and Data Steward leads with a clear RACI.
- Install the incident-to-problem-to-change discipline: swarm P1s with business comms templates, run root-cause analysis on chronic issues, and split changes into standard, normal, and emergency with CAB criteria and pre-approved patterns.
- Run a release train of monthly maintenance drops and quarterly value releases, each with a regression pack, cutover steps, and a rollback plan, and require a benefit hypothesis with a named owner for every value release, so no enhancement ships without a stated outcome to measure it against.
- Embed evergreen enablement: role-based micro-learning, a maintained runbook and SOP knowledge base with a quarterly review cadence so it never goes stale, and a business champions community that escalates recurring patterns and spreads good practice across sites.
Where a healthy run-state slips back
- Permanent hypercare. Fix: set a dated exit from hypercare into a defined run-state with named owners, so the heroics stop and the system takes over.
- SLAs with no measurement. Fix: publish MTTA, MTTR, recurrence, and change failure rate on a monthly dashboard, because a target no one watches is decoration.
- Every change treated as an emergency. Fix: pre-approve standard change patterns and reserve emergency handling for genuine outages, and drive the change failure rate below 5 percent by widening the library of standard changes over time.
- Master data drift. Fix: give a Data Steward lead a quality dashboard and a two-day fix SLA, so bad masters do not quietly corrupt downstream processes such as pricing, tax, and financial reporting.
- Access controls that decay. Fix: run a quarterly segregation-of-duties review and high-risk privilege attestations with audit trails, so compliance holds between audits.
Day-1 essentials
- Stand up a 24-by-7 P1 on-call roster and war-room protocol for the first four weeks.
- Monitor every interface with alert thresholds and a contact matrix.
- Publish the service catalog and a ticket routing guide before hypercare ends.
- Seed a known-error knowledge base from hypercare learnings and name its owner.
- Schedule the first monthly service review with published KPIs and A3 problem sheets on the top chronic issues.