Customer Case Study · Telecom · North America
The network that heals itself
How a Tier-1 telecom operator in North America moved from reactive fault handling to a four-stage closed loop: problems discovered from symptoms and anomalies, root causes diagnosed through RCA trees built from rules, ML models and knowledge documents, remediation executed under governance, and every loop closed with quantitative feedback, with the full provenance of each automated fix on record.
- 4stage closed loop, end to endConfigured by the operator, not fixed by us
- 7network data domains correlated into one loopPreviously seven separate systems of record
- ~60%of recurring network faults resolved autonomously, end to endFull loop, no human in the path
- ~50%faster root-cause identification through ML- and rule-driven RCA treesAgainst the prior manual correlation process
- ~35%fewer repeat incidents via top-offender and persistent-root-cause analyticsOnly measurable once closure is measured
- 100%of automated fixes carry full provenance: RCA path, actions taken, closure stateEvery fix, not a sample
The Challenges
The operator's network told its story across seven different data domains, faults, performance counters, configuration, inventory and topology, telemetry, syslogs and OSS streams, and no two of them told it the same way. Root-causing meant human correlation across all of them; fixing meant manual action; and the same faults returned because nothing closed the loop.
- Symptoms without discovery
Anomalies surfaced in one domain at a time; recognizing them as one network problem was manual pattern-matching.
- RCA as tribal knowledge
Diagnosis lived in senior engineers' heads and scattered documents, unrepeatable and unavailable at 3 a.m.
- Manual remediation
Every fix meant a human raising tickets, chasing workflows and executing procedures by hand.
- No loop closure
Whether a fix actually worked was rarely measured; issues were closed on hope, and reopened on evidence.
- Repeat offenders untracked
Persistent, frequent root causes kept returning because nothing aggregated them across the network.
- No provenance
When automation did act, nobody could trace which logic fired, what path it took, or why.
The architecture, and where the loop closes
Seven data domains that described the same network in seven different vocabularies now feed one loop with four configurable stages. Three of those stages are what this category calls self-healing. The fourth is the one almost nobody builds, and it is the reason the other three are worth anything: the loop refuses to record an issue as resolved until quantitative metrics say the fix held.
Autonomy with a paper trail: every automated correction records which RCA path fired, what evidence supported it, what actions ran and how the loop closed, so "the system fixed it" is always a statement with provenance attached.
What was switched on, and what was not
The third stage of the loop is the one an operator argues about, and rightly. This deployment resolves that argument per action rather than globally: remediation is configured stage by stage as autonomous or manual, in the operator hands, not ours.
Around 60 percent of recurring faults now run the full loop without a human. The remaining 40 percent is not the easy 40 percent: it is the novel, the ambiguous and the genuinely cross-domain, and those reach an engineer with the full RCA traversal and evidence trail attached.
Runs without a human
- Discovery across all seven data domains, continuously, on symptoms rather than on alarms alone
- Diagnosis through the RCA trees, including the operator own knowledge documents made executable
- Remediation for the actions the operator has configured as autonomous
- Closure evaluation against quantitative metrics, with a terminal state recorded either way
- Top-offender and persistent-root-cause analytics across the estate
Does not run without a human
- Any remediation the operator has configured as manual, which is theirs to decide per action
- Changing what the loop is allowed to do. Configuration is an operator act, not a platform one
- Deciding an issue is closed when the closure metrics say it is not. A failure reopens it
- Anything whose result cannot be asserted from live signal after the fact
What happens now when a fault appears
The sequence is the same whether the first sign of trouble shows up as a counter, a syslog line or a configuration change, and whether anybody is looking.
The Sentinel loop, as it runs here
Every Opstral deployment runs the same four-stage loop: Observe, Investigate, Act, Optimize. It is the methodology rather than a feature list, and the point of setting it out per deployment is that you can see which stages carried the weight in this one and which did not.
- ObserveTake in every signal the estate produces, normalised and correlated as it arrives rather than after somebody goes looking.HereFM, PM, CM, topology, telemetry, syslogs and OSS records feed discovery, which notices a symptom rather than waiting for an alarm somebody already configured.
- InvestigateWork the signal into a probable cause with the evidence attached, before anyone is notified.HereDiagnosis walks an RCA tree the operator wrote, so the reasoning is inspectable and owned by the customer rather than inferred and unexplainable.
- ActRun the approved procedure where the blast radius allows it, or hand a named human the plan, the evidence and the rollback.HereRemediation runs autonomously or waits for a human, exactly as configured per action. Closure is then decided by numbers rather than by the procedure having finished.
- OptimizeFeed the outcome back so the next run of the loop is better informed than the last.HereThe closure verdict feeds back into discovery, so a pattern that did not actually resolve changes what discovery does next time. Everything carries its provenance through all four stages.
- Discovery notices a symptom, not an alarm
ML models and custom rules watch all seven domains for anomalies. The distinction matters: waiting for an alarm means waiting for a threshold somebody set to be crossed, and the faults that recur are frequently the ones that never cross one cleanly.
PlatformObserve - Diagnosis walks an RCA tree the operator wrote
Trees assembled from rules, ML models and the operator own knowledge documents. This is the step that turns diagnosis from tribal knowledge into something repeatable, and available at 03:00 to whoever is actually on call rather than to whoever wrote it.
PlatformInvestigate - Remediation runs, or waits, exactly as configured
Trouble tickets, change requests, workflow orchestration, notifications and custom procedures. Whether a given action executes autonomously or holds for a person is the operator configuration, set per action rather than as a single global switch.
PlatformAct - Closure is decided by numbers, not by completion
Quantitative metrics are evaluated and a terminal state is recorded: closed, on hold, or failed. A procedure that ran successfully and a fault that actually cleared are different events, and only the second one closes an issue here.
PlatformAct - The verdict feeds back into discovery
Failures reopen. Successes and failures both accumulate, which is what makes top-offender and persistent-root-cause analytics possible at estate level. The fault you have fixed four times stops being something somebody half-remembers and becomes a ranked entry.
PlatformOptimize - Everything carries its provenance
Which RCA path fired, what evidence supported it, what actions ran and how the loop closed. That is what turns "the system fixed it" from a claim into a statement with a trail behind it, and it is the answer to the only question an incident review ever actually asks.
PlatformOptimize
The Impact
Recurring faults stopped consuming the team. The loop discovers, diagnoses, fixes and verifies on its own for the well-understood majority of network problems, and hands engineers the evidence trail for everything else.
| Dimension | Before | After · closed-loop |
|---|---|---|
| Detection | Alarm-chasing, domain by domain | Symptom- and anomaly-driven discovery across seven data domains |
| Diagnosis | Tribal knowledge, unrepeatable | RCA trees from rules, ML and knowledge docs, ~50% faster to root cause |
| Remediation | Manual tickets, workflows and procedures | Governed auto-correction, ~60% of recurring faults resolved end to end |
| Verification | Closed on hope, reopened on evidence | Quantitative feedback with recorded terminal states |
| Repeat incidents | Same causes returning, untracked | Top-offender analytics, ~35% fewer repeats |
| Traceability | Automation as a black box | Full provenance of every RCA traversal and action |
Self-healing is not one clever action, it is a loop that never skips a step: discover from symptoms, diagnose with the operator's own knowledge made executable, act under governance, and refuse to close until the numbers say fixed. The provenance is what makes the autonomy trustworthy.
At a glance
- Industry
- Tier-1 Telecom Operator
- Geography
- North America
- Engagement
- Production deployment
- Scope
- Self-Healing & Closed-Loop Network Automation
- Platform components
- Sentinel AI, operator-authored RCA trees, governed remediation, closure verification
- Evidence basis
- Improvement figures are indicative, reflecting target and expected outcomes of the deployment rather than measured production results. Customer identity withheld by request.
Close the loop on your network operations
Symptom-driven discovery, executable RCA, governed remediation and quantitative closure, configured to your network, running inside your environment.
Customer identity withheld by request and referred to throughout as a "Tier-1 telecom operator in North America." Improvement figures are indicative, reflecting target and expected outcomes of the deployment; exact results vary by network scope, data quality and rollout phase.