The same incident, for the eleventh time this quarter
Each one was closed correctly, inside its commitment, by a competent engineer. Nobody joined them up, so the thing causing all eleven is still there.
What actually happens
Every individual incident in this story was handled well. That is precisely the mechanism by which the underlying problem survives.
An incident occurs, is diagnosed, is resolved inside its commitment and is closed. The engineer who handled it did a good job.
Three weeks later a similar incident occurs. Different engineer, different shift, possibly a slightly different presenting symptom. Also diagnosed, also resolved, also closed correctly.
This repeats. Over a quarter it happens eleven times. Each occurrence is individually small, individually resolved, and individually unremarkable, and the total cost across eleven occurrences is substantially more than the effort it would take to find the cause once.
It does not get aggregated because aggregation is nobody job in the moment. Problem management exists in the process documentation and requires somebody to notice the pattern, and noticing a pattern across eleven tickets spread over three months, handled by seven different people, is not something the queue view supports.
The signal is genuinely there. The tickets share a cause even when they do not share wording, and the linkage is discoverable from resolution notes, affected components and timing.
Eleven incidents, each closed correctly, each inside commitment. Every quality metric on that queue looks healthy and the underlying problem has never been touched.
The same quarter, two ways
This is measured in occurrences rather than minutes, which is the only honest way to measure a recurrence problem.
Illustrative, not measured. The times below model a scenario built from the patterns we see in production estates. They are not timings recorded at a named customer. The point is the shape of the clock, not the totals: check it against your own last ten incidents.
Today, resolved individually
- Week 1Incident occurs. Diagnosed and resolved inside commitment. Closed.
- Week 4Similar incident, different engineer, slightly different symptom. Resolved and closed.
- Weeks 4 ↓ 12WaitingNine further occurrences across seven engineers and three shifts. Each handled correctly.
- Week 12WaitingNobody has aggregated them. No problem record exists.
- Quarter endWaitingQueue metrics look healthy. MTTR is good. Recurrence is not a tracked metric.
- Next quarterIt happens again.
11 occurrences · and the cause is still in place
With incidents clustered continuously
- Week 1First occurrence resolved and closed as normal.
- Week 4Sentinel matches the second occurrence to the first on affected components, resolution actions and timing rather than on ticket wording.
- Week 4Cluster opened. Both tickets linked. Not yet enough evidence to justify a problem record.
- Week 6Third occurrence. Cluster confidence crosses the threshold. Problem record raised automatically with all three linked.
- Week 6Common factor surfaced across the cluster with the supporting evidence from each occurrence attached.
- Week 8Systemic fix approved and applied under governance. Recurrence watch armed for the following quarter.
3 occurrences · then the cause was addressed
The arithmetic is unusually clean here. Eight avoided occurrences times the cost of handling one, against the cost of one systemic fix. In most estates that comparison is not close.
The reason it does not happen today is not that nobody values problem management. It is that recurrence is not visible from any view the service desk actually uses, and detecting it requires comparing tickets on their substance rather than on their text.
The mechanism is clustering on affected components, resolution actions and timing rather than on ticket wording. Two tickets describing the same underlying fault in different words are the normal case, not the exception, and text similarity misses most of them.
The threshold matters. Raising a problem record on the second occurrence produces noise; waiting for eleven produces the current situation. Three occurrences with a strong component and resolution match is where the evidence starts to justify the investigation.
Why the number is what it is
The arithmetic is unusually clean here. Eight avoided occurrences times the cost of handling one, against the cost of one systemic fix. In most estates that comparison is not close.
The reason it does not happen today is not that nobody values problem management. It is that recurrence is not visible from any view the service desk actually uses, and detecting it requires comparing tickets on their substance rather than on their text.
The mechanism is clustering on affected components, resolution actions and timing rather than on ticket wording. Two tickets describing the same underlying fault in different words are the normal case, not the exception, and text similarity misses most of them.
The threshold matters. Raising a problem record on the second occurrence produces noise; waiting for eleven produces the current situation. Three occurrences with a strong component and resolution match is where the evidence starts to justify the investigation.
The mechanism is clustering on affected components, resolution actions and timing rather than on ticket wording.
Who decides to press go
Raising a problem record changes nothing in production, but the systemic fix that follows it does.
Clustering, linking, problem record creation and evidence assembly all run under policy. None of them change a running system.
The fix is raised as an Action Ticket with the cluster evidence, the blast radius and the rollback path attached, and it goes through change approval like any other change.
The value here is almost entirely in the detection. Once a problem record exists with eleven linked occurrences and a common factor identified, getting the fix approved has never been the hard part.
Individual incidents resolved and closed correctly inside commitment across multiple engineers and shifts. No aggregation view available.
Occurrences matched on affected components, resolution actions and timing rather than ticket text. Cluster confidence accumulated across occurrences.
Problem record raised automatically once evidence justifies it, with all occurrences linked and the common factor surfaced. Systemic fix raised through change approval.
Recurrence watch armed after the fix. Cluster patterns retained so a return of the same signature is detected on the first occurrence rather than the third.
This is a platform capability, not a published customer deployment for this exact scenario. The mechanism, which is cross-system correlation followed by governed MOP execution with pre-check, post-check, rollback and approval gating, is running in production today across managed estates; see governed day-2 operations across 2,000+ nodes and closed-loop network automation. The timings shown are modelled, not measured at a named customer.
If the action carries no service impact
Clustering, linking, problem record creation and evidence assembly all run under policy. None of them change a running system.
If it applies a systemic fix
The fix is raised as an Action Ticket with the cluster evidence, the blast radius and the rollback path attached, and it goes through change approval like any other change.
What Sentinel did, step by step
- ObserveIndividual incidents resolved and closed correctly inside commitment across multiple engineers and shifts. No aggregation view available.
- InvestigateOccurrences matched on affected components, resolution actions and timing rather than ticket text. Cluster confidence accumulated across occurrences.
- ActProblem record raised automatically once evidence justifies it, with all occurrences linked and the common factor surfaced. Systemic fix raised through change approval.
- OptimizeRecurrence watch armed after the fix. Cluster patterns retained so a return of the same signature is detected on the first occurrence rather than the third.
Bring us a quarter of closed tickets
We will cluster them and show you how many were the same problem.