Two hundred ATM tickets, one upstream cause
The service desk filled with one ticket per machine, every one of them triaged separately, while the actual fault sat one layer up where nothing was alarming.
What actually happens
The cost of this incident is not the outage. It is the two hundred people who each did the right thing with the wrong picture.
ATMs across a region start declining transactions. Each machine does exactly what it is designed to do: it detects that it cannot reach the authorisation host, and it raises a ticket. Two hundred machines, two hundred tickets, arriving over about twelve minutes.
The service desk triages them the only way it can, which is one at a time. Each ticket looks like an isolated machine fault, because from inside the ticket that is exactly what it looks like. Field dispatch starts rolling crews to the first addresses in the queue.
Nobody looks upstream, because upstream is not alarming. The switch path carrying that region is degraded rather than down, and a degraded path does not trip a device-level alarm. The ATMs are the only things reporting, so the ATMs are where everyone looks.
By the time someone plots the failing machines on a map and notices they all sit behind one aggregation point, several crews are already in vans. The repair, once the right layer is identified, affects one path. The cost already incurred is two hundred triage cycles and a set of dispatches that were never going to fix anything.
The customer-facing damage is worse than the ticket count suggests. Every declined transaction at a cash machine is a person standing in front of a screen, and a meaningful share of them will call.
Two hundred tickets is not two hundred problems. It is one problem, reported two hundred times, by the only devices in the chain that were instrumented to complain.
The same morning, two ways
Correlation is not a reporting feature. It decides whether field crews leave the depot.
Illustrative, not measured. The times below model a scenario built from the patterns we see in production estates. They are not timings recorded at a named customer. The point is the shape of the clock, not the totals: check it against your own last ten incidents.
Today, one ticket at a time
- 09:14First ATM declines. Ticket raised. Triaged as an isolated machine fault.
- 09:14 ↓ 09:26WaitingVolume climbs to about 200 tickets. Each is triaged separately. Queue times climb.
- 09:26Field dispatch begins rolling crews to the earliest addresses.
- 09:26 ↓ 10:35WaitingSomeone maps the failing machines by hand and notices a shared aggregation point.
- 10:35Network team engaged. Degraded switch path confirmed.
- 11:05Path restored. Crews recalled. Tickets closed in bulk.
~2 hours · plus dispatches that could not have helped
With Sentinel correlating on arrival
- 09:14First ATM declines. Sentinel opens an investigation rather than a ticket.
- 09:17Incoming device faults matched against topology as they arrive. Shared upstream aggregation point identified.
- 09:19Switch path interface counters and recent change records correlated. Degraded path confirmed as probable cause.
- 09:20One incident raised with the affected-machine list attached. Field dispatch held automatically pending cause confirmation.
- 09:22Network duty manager notified with the proposed path failover MOP and its blast radius.
- 09:38Approved and executed. Sherlock confirms the machines recover. The 200 device tickets close against the parent.
~25 minutes · no crew dispatched to a machine that was working
Look at where the two hours went. The repair was one action on one path. Almost everything before it was the estate reporting the same fact two hundred times into a queue that could only read it one line at a time.
This is what alarm storms cost. Not confusion, which teams handle, but the sequencing: triage capacity is consumed by symptoms, so the search for the cause starts late and starts from the wrong layer.
Sentinel correlates as events arrive rather than after they are queued. Device faults are matched against topology in flight, so two hundred symptoms become one incident with a cause hypothesis and an affected-machine list attached, before dispatch has a chance to act on the symptom.
Holding dispatch is the part worth arguing about internally. Every truck roll to a machine that was never faulty is cost you can measure precisely, and it is usually a larger number than the outage itself.
Why the number is what it is
Look at where the two hours went. The repair was one action on one path. Almost everything before it was the estate reporting the same fact two hundred times into a queue that could only read it one line at a time.
This is what alarm storms cost. Not confusion, which teams handle, but the sequencing: triage capacity is consumed by symptoms, so the search for the cause starts late and starts from the wrong layer.
Sentinel correlates as events arrive rather than after they are queued. Device faults are matched against topology in flight, so two hundred symptoms become one incident with a cause hypothesis and an affected-machine list attached, before dispatch has a chance to act on the symptom.
Holding dispatch is the part worth arguing about internally. Every truck roll to a machine that was never faulty is cost you can measure precisely, and it is usually a larger number than the outage itself.
Sentinel correlates as events arrive rather than after they are queued.
Who decides to press go
Failing over a switch path serving a live ATM region carries service impact, so it is not something Sentinel does on its own authority.
The Action Ticket executes under policy: pre-check, execute, post-check, verify, with automatic rollback if the post-check fails. Nobody is paged.
The incident is raised and held with the proposed failover MOP, the affected-machine list and the rollback path attached. The network duty manager approves before anything runs.
What Sentinel does take on its own authority is the safe half: holding field dispatch until the cause is confirmed. That reverses cleanly and costs nothing if the hypothesis is wrong.
Device faults arriving from ATMs across one region. Each machine reporting an authorisation host unreachable. No alarm on the layer above.
Faults matched against topology in flight. Common upstream aggregation point identified. Interface counters and recent change records on that path correlated.
One parent incident raised with the affected-machine list. Field dispatch held pending confirmation. Path failover MOP staged and held for approval.
Degraded-path detection added on the aggregation layer so the upstream reports before the endpoints do. Correlation rule retained for the region.
This is a platform capability, not a published customer deployment. The mechanism, which is correlated investigation followed by governed MOP execution with pre-check, post-check, rollback and approval gating, is running in production today in our carrier estates; see governed day-2 operations across 2,000+ nodes and closed-loop network automation. The scenario above is that same mechanism applied to a banking context. The timings shown are modelled, not measured at a named bank.
If the action carries no service impact
The Action Ticket executes under policy: pre-check, execute, post-check, verify, with automatic rollback if the post-check fails. Nobody is paged.
If it carries service impact, as it does here
The incident is raised and held with the proposed failover MOP, the affected-machine list and the rollback path attached. The network duty manager approves before anything runs.
What Sentinel did, step by step
- ObserveDevice faults arriving from ATMs across one region. Each machine reporting an authorisation host unreachable. No alarm on the layer above.
- InvestigateFaults matched against topology in flight. Common upstream aggregation point identified. Interface counters and recent change records on that path correlated.
- ActOne parent incident raised with the affected-machine list. Field dispatch held pending confirmation. Path failover MOP staged and held for approval.
- OptimizeDegraded-path detection added on the aggregation layer so the upstream reports before the endpoints do. Correlation rule retained for the region.
Bring us your last alarm storm
We will show you how many of those tickets were the same incident.