Four thousand subscribers dark, four thousand alarms
The NOC saw a wall of individual faults and the contact centre started filling, while the thing that actually broke sat one layer up and never raised an alarm of its own.
What actually happens
An access network under fault does not send you a problem. It sends you every symptom of that problem, all at once, from the only devices that were instrumented to complain.
An uplink serving a set of OLT shelves degrades. Within about three minutes, 4,120 ONTs report loss of signal. Each report is accurate. None of them describes what happened.
The NOC console fills. The alarms are individually meaningless and collectively overwhelming: there is no parent alarm, because the parent did not fail in a way that produces one. It degraded, and degradation on a transport path frequently does not trip a device-level alarm at all.
The contact centre starts taking calls at roughly the same moment. Those calls are handled as individual service faults, because from the agent script that is exactly what they are, so a second queue starts building alongside the first.
Somebody eventually works upward from the ONT list to the OLT to the uplink. That reasoning is correct and it is also slow, because it is being done by a person reading a console under load, and the console is showing them four thousand rows.
The repair, once the right layer is named, is one action on one path. Everything before it was the estate telling the NOC the same fact four thousand times.
Four thousand alarms is not four thousand problems. It is one problem, reported four thousand times, by the only devices in the chain instrumented to notice.
The same evening, two ways
The value of correlation is not a tidier console. It is where the NOC starts looking in the first ten minutes.
Illustrative, not measured. The times below model a scenario built from the patterns we see in production estates. They are not timings recorded at a named customer. The point is the shape of the clock, not the totals: check it against your own last ten incidents.
Today, working upward from the symptoms
- 20:41Uplink degrades. First ONT reports loss of signal.
- 20:41 ↓ 20:44Waiting4,120 ONTs report loss of signal. No parent alarm. Console saturates.
- 20:44 ↓ 21:10WaitingNOC triages individual alarms. Contact centre queue builds in parallel.
- 21:10Engineer maps affected ONTs to OLT shelves and identifies a shared uplink.
- 21:10 ↓ 21:35WaitingTransport team engaged. Degraded uplink confirmed on interface counters.
- 21:52Path failed over. Subscribers recover. Alarms clear in bulk.
~70 minutes · most of it spent reading symptoms
With Sentinel correlating on arrival
- 20:41First ONT reports loss of signal. Sentinel opens an investigation rather than a ticket.
- 20:43Incoming alarms matched against topology as they arrive. Common OLT shelves and shared uplink identified.
- 20:45Uplink interface counters, optical levels and recent change records correlated. Degraded path confirmed as probable cause.
- 20:46One incident raised. Blast radius sized: affected ONT list, subscriber count, business circuits on the same path flagged.
- 20:47Contact centre given the affected-subscriber list so agents stop opening individual faults.
- 21:04Transport duty manager approves the failover MOP. Path switched. Sherlock confirms ONTs recover.
~25 minutes · one incident, not four thousand
Compare where the time went. The failover itself takes minutes in both columns. The difference is that in the first, the NOC spends half an hour establishing something the topology already knew.
The blast radius matters as much as the cause. Knowing that 4,120 residential subscribers and a handful of business circuits sit on the same path changes the priority, the comms and who gets told, and that information is available at the moment of correlation rather than after the repair.
Sentinel matches alarms against topology in flight rather than after they queue. Four thousand symptoms become one incident with a cause hypothesis and an affected-subscriber list attached, before the NOC has read the second row.
The contact centre effect is the one operators consistently underestimate. Handing agents the affected-subscriber list at minute six, rather than at minute seventy, stops a second queue forming behind the first.
Why the number is what it is
Compare where the time went. The failover itself takes minutes in both columns. The difference is that in the first, the NOC spends half an hour establishing something the topology already knew.
The blast radius matters as much as the cause. Knowing that 4,120 residential subscribers and a handful of business circuits sit on the same path changes the priority, the comms and who gets told, and that information is available at the moment of correlation rather than after the repair.
Sentinel matches alarms against topology in flight rather than after they queue. Four thousand symptoms become one incident with a cause hypothesis and an affected-subscriber list attached, before the NOC has read the second row.
The contact centre effect is the one operators consistently underestimate. Handing agents the affected-subscriber list at minute six, rather than at minute seventy, stops a second queue forming behind the first.
Sentinel matches alarms against topology in flight rather than after they queue.
Who decides to press go
Failing over a transport path carrying live subscribers and business circuits carries service impact, so Sentinel proposes it rather than taking it.
The Action Ticket executes under policy: pre-check, execute, post-check, verify, with automatic rollback if the post-check fails. Nobody is paged.
The incident is raised and held with the proposed failover MOP, the affected-subscriber list, the business circuits on the path and the rollback path attached. The transport duty manager approves before anything switches.
What Sentinel does take on its own authority is the reversible half: collapsing the alarms into one incident and giving the contact centre the affected list. Neither of those can make anything worse.
Loss-of-signal reports arriving from thousands of ONTs inside three minutes. No parent alarm on the layer above. Contact centre volume rising.
Alarms matched against topology in flight. Common OLT shelves and shared uplink identified. Interface counters, optical levels and recent change records on that path correlated.
One parent incident raised with blast radius sized. Contact centre given the affected-subscriber list. Failover MOP staged and held for approval.
Degradation detection added on the uplink so transport reports before the access layer does. Business circuits on shared paths tagged for priority in future correlations.
This is a platform capability, not a published customer deployment for a named fibre operator. The mechanism, which is topology-aware correlation followed by governed MOP execution with pre-check, post-check, rollback and approval gating, is running in production today; see carrier-scale observability across 27,000+ devices and closed-loop network automation. The access-network scenario above applies that same mechanism to a PON estate. The timings shown are modelled, not measured at a named operator.
If the action carries no service impact
The Action Ticket executes under policy: pre-check, execute, post-check, verify, with automatic rollback if the post-check fails. Nobody is paged.
If it carries service impact, as it does here
The incident is raised and held with the proposed failover MOP, the affected-subscriber list, the business circuits on the path and the rollback path attached. The transport duty manager approves before anything switches.
What Sentinel did, step by step
- ObserveLoss-of-signal reports arriving from thousands of ONTs inside three minutes. No parent alarm on the layer above. Contact centre volume rising.
- InvestigateAlarms matched against topology in flight. Common OLT shelves and shared uplink identified. Interface counters, optical levels and recent change records on that path correlated.
- ActOne parent incident raised with blast radius sized. Contact centre given the affected-subscriber list. Failover MOP staged and held for approval.
- OptimizeDegradation detection added on the uplink so transport reports before the access layer does. Business circuits on shared paths tagged for priority in future correlations.
Bring us your last alarm flood
We will show you how few incidents those alarms actually were.