Customer Case Study · Telecom · India
27,000+ network devices, one governed NOC
How a carrier-scale Tier-1 telecom operator in India brought fault and performance data from more than twenty-seven thousand routers, switches and network devices, arriving over SNMP, Kafka streams and poll-based collection, into one near-real-time intelligence layer where NOC users see insights, ask Sentinel for live values, and act through governed automation.
- 27,000+routers, switches and devices under one pane, previously no central real-time viewCounted at go-live
- 3ingestion paths unified: SNMP traps & metrics, Kafka streams, poll-based collectionPreviously three uncorrelated silos
- Near real timefault and performance processing, from device event to investigated insightProcessing latency, not report cadence
- 43%reduction in MTTR, measured in production within six months of go-liveMeasured six months post go-live vs documented baseline
- 85%of routine manual operations automated under governance, measured in productionMeasured in production
- ~60%of incidents auto-created with probable cause and context attachedShare of raised incidents
The Challenges
The operator's network estate, more than twenty-seven thousand routers, switches and devices, generated a continuous flood of fault and performance data over SNMP, Kafka and polling mechanisms. There was no central solution monitoring these devices in real time: what was happening inside the network was fragmented across collection silos, and the NOC saw it late, partially, or not at all.
- No central real-time view
Twenty-seven thousand devices, no single place to see their fault and performance state as it happened.
- Three ingestion worlds
SNMP traps and metrics, Kafka event streams and poll-based collection each lived in its own pipeline, with no correlation across them.
- Volume beyond human triage
Carrier-scale telemetry arrived faster than any team could read it; meaningful signals drowned in the stream.
- Investigation was manual archaeology
Root-causing a fault meant hopping between element managers and raw counters, device by device.
- Dashboards without answers
Where views existed, they showed symptoms; getting an actual live value still meant logging into the device.
- Network, infra and services apart
Device faults, server health and service state were monitored, where at all, in separate tools with separate truths.
The architecture, and where the decision sits
Three collection mechanisms that had never met now land in one pipeline, normalised and correlated on arrival rather than joined by a person afterwards. What follows the pipeline is the part worth reading closely, because it is where this deployment differs from an observability rollout: nothing acts on its own judgement, and the boundary is drawn at the procedure rather than at the confidence score.
Governed autonomy, deliberately bounded. Sentinel suggests root causes rather than acting on its own judgement. Automation is limited to predefined, approved MOPs. Every incident, ticket and action carries a full audit trail and rollback. Autonomy is earned scope by scope, never assumed.
What was switched on, and what was not
The most common question a carrier NOC asks about an autonomous platform is not what it can do. It is what it is allowed to do, on a network carrying live subscriber traffic, at three in the morning, without asking anybody. This is where the boundary sits in this deployment today. It is narrower than the marketing category implies, and that is the reason the deployment exists.
Read the right-hand column as the product rather than as a limitation. A vendor arriving at a Tier-1 NOC with an agent that decides on its own judgement does not get a second meeting, and should not. The left-hand column is permitted precisely because the right-hand column exists and is enforced rather than promised.
Runs without a human
- Ingestion, normalisation and correlation across all three collection paths, continuously
- Investigation of a fault against topology, incident history and correlated signals
- Raising an incident with the probable cause and its evidence already attached
- Answering a request for a live value or state from the estate, on demand
- Execution of a predefined, approved MOP for a known pattern, as a reversible and audited Action Ticket
Does not run without a human
- Any remediation outside the approved MOP set, however confident the diagnosis
- Deciding that a probable cause is the cause. Sentinel proposes; an engineer concludes
- Any action whose previous state cannot be restored by a defined path
- Anything the platform cannot verify afterwards from live signal
- Widening its own scope. New procedures enter the approved set by a deliberate act, not by performance
What happens now when a fault appears
The same sequence runs whether the signal arrives as an SNMP trap, a Kafka event or a polled counter, and whether it is 14:00 or 03:00. The step that changed the numbers is the third one, and it is the least glamorous.
The Sentinel loop, as it runs here
Every Opstral deployment runs the same four-stage loop: Observe, Investigate, Act, Optimize. It is the methodology rather than a feature list, and the point of setting it out per deployment is that you can see which stages carried the weight in this one and which did not.
- ObserveTake in every signal the estate produces, normalised and correlated as it arrives rather than after somebody goes looking.HereSNMP traps, Kafka streams and polled counters land in one pipeline and are correlated on arrival, across 27,000 devices. An engineer can also ask for a live value from the estate rather than opening a session on the device.
- InvestigateWork the signal into a probable cause with the evidence attached, before anyone is notified.HereThe fault is examined against topology, prior incidents with the same signature and the correlated signals around it. Nothing is paged while this runs. The incident is created carrying probable cause and evidence.
- ActRun the approved procedure where the blast radius allows it, or hand a named human the plan, the evidence and the rollback.HereWhere the fault matches an approved MOP, ProcBot runs it as a reversible Action Ticket and nobody is paged. Everything novel or cross-domain goes to a named NOC engineer, who starts from a case rather than a console.
- OptimizeFeed the outcome back so the next run of the loop is better informed than the last.Not recorded hereThis deployment does not document this stage as a distinct step, so we do not claim it.
The feedback stage is not separately recorded in this deployment. Where we have it documented as a distinct step, it is on the closed-loop network automation study, where the closure verdict feeds back into discovery.
- The signal lands and is correlated on arrival
It is normalised into the single pipeline and related to whatever else arrived in the same window from the other two paths. Previously this join happened in a person head, if it happened at all, and only once somebody noticed there was something to join.
PlatformObserve - Sentinel investigates rather than notifies
The fault is examined against topology, prior incidents with the same signature, and the correlated signals around it. Nothing is paged while this runs, because a notification carrying a symptom is what the previous process already produced in abundance.
PlatformInvestigate - An incident is raised carrying the cause, not the symptom
Probable cause, the evidence supporting it, and the services in the path arrive together as the incident is created. Around 60 percent of incidents are now auto-created this way. This is where the manual archaeology went: the hop between element managers, the search for which pipeline held the relevant signal, the session opened on a device to read one value.
PlatformInvestigate - Known pattern, approved procedure: it executes
Where the fault matches a predefined MOP that the operator has approved, ProcBot runs it as an Action Ticket. The run is reversible, the full record is written as it happens, and nobody is paged. Eighty-five percent of routine manual operations now run this way.
PlatformAct - Everything else: an engineer decides
The probable cause and its evidence are presented, and a NOC engineer concludes and acts. This is the majority path for anything novel or cross-domain, and it is the intended design rather than a phase to be grown out of. What changed is that the engineer starts from a case rather than from a console.
Named engineerAct - A live value is a question, not a login
At any point in the above, an engineer can ask the platform for an actual current value or state and get it from the live estate. Previously this meant opening a session on the device. It sounds minor next to a data pipeline and it is the change engineers mentioned first.
Named engineerObserve
The Impact
The NOC moved from fragmented, delayed device monitoring to one near-real-time picture of twenty-seven thousand devices, with investigation, values and action available in the same place the fault appears.
| Dimension | Before | After · unified & AI-driven |
|---|---|---|
| Device visibility | No central real-time solution across the device estate | 27,000+ devices in one near-real-time fault and performance pane |
| Data ingestion | SNMP, Kafka and polled data in separate, uncorrelated silos | Three paths, one pipeline, normalized and correlated on arrival |
| Investigation | Manual, device-by-device root-causing across consoles | Sentinel investigation with probable cause suggested, evidence attached, engineer in command |
| Access to truth | Log into the device to read an actual value | Ask Sentinel, live values and states answered conversationally |
| Response | Ad-hoc manual procedures, untracked | 85% of routine manual operations automated as governed MOPs; ~60% of incidents auto-created with context; 43% MTTR reduction, measured |
| Coverage | Network, infrastructure and services in separate tools | One governed layer across devices, infrastructure and services |
At twenty-seven thousand devices, the question is never whether data exists, it always does. The question is whether anyone can see it in time, trust what it means, and act on it safely. Centralizing the ingestion was the start; putting Sentinel and governed MOPs on top is what changed how the NOC works.
At a glance
- Industry
- Tier-1 Telecom Operator (Carrier-Scale)
- Geography
- India
- Engagement
- Production deployment
- Scope
- Network Fault & Performance + Infrastructure & Services
- Scale
- 27,000+ routers, switches & devices
- Platform components
- Sentinel AI, ProcBot, Action Tickets, SNMP / Kafka / poll-based ingestion
- Evidence basis
- Measured in production six months after go-live against the operator’s documented baseline. Customer identity withheld by request.
Bring your device estate into one governed NOC
SNMP, Kafka and poll-based ingestion at carrier scale, near-real-time processing, Sentinel investigation, and governed automation, deployed inside your environment.
Customer identity withheld by request and referred to throughout as a "Tier-1 telecom operator in India." The MTTR reduction and routine-operations automation figures are measured production results from this deployment; remaining improvement figures are indicative, and exact results vary by estate scope, telemetry profile and rollout phase.