The port was dying for nine days before anyone looked
No alarm fired, because nothing had failed yet. The subscribers on that port just had a slightly worse service every day until some of them stopped being subscribers.
What actually happens
The faults that cost you customers are rarely the ones that take the network down. They are the ones that make it slightly worse for long enough that people leave.
Optical receive power on one PON port begins drifting out of tolerance. It is a slow movement, a fraction of a decibel at a time, and at no point does it cross the alarm threshold that would raise it to the NOC.
Error counters climb on a subset of ONTs behind that port. Retransmissions increase. For the subscribers on those ONTs the service is not down, it is intermittent: video calls drop occasionally, streams re-buffer, and the connection recovers before anyone finishes a speed test.
Those subscribers do not usually raise a fault, because the intermittent version is much harder to report than the down version. A share of them raise something worse instead, which is a churn decision that never gets attributed to this port.
On day nine the port fails properly, which produces an alarm, an emergency dispatch and a service-affecting outage during business hours. The repair is a planned-maintenance job that was forced into an emergency window by the fact that nobody saw it coming.
The signal was present the whole time. It was just present as a trend in a metric that nothing was watching as a trend.
Nine days of degradation produced no alarm and one outage. The alarm threshold was set correctly. It was answering the wrong question.
The same nine days, two ways
This one is not measured in minutes. The unit that matters here is days of warning.
Illustrative, not measured. The times below model a scenario built from the patterns we see in production estates. They are not timings recorded at a named customer. The point is the shape of the clock, not the totals: check it against your own last ten incidents.
Today, waiting for the threshold
- Day 1Optical receive power begins drifting. Well inside tolerance. No alarm.
- Day 1 ↓ Day 8WaitingDrift continues. Error counters climb on a subset of ONTs. Intermittent service. No threshold crossed, no alarm, no ticket.
- Day 9Port fails. Alarm raised. Service-affecting outage during business hours.
- Day 9 +40mWaitingEmergency dispatch mobilised. Field crew pulled off planned work.
- Day 9 +3hPort replaced. Service restored.
- Day 9+Churn from the preceding eight days is not attributed to this port and never will be.
9 days of silent degradation · then an emergency
With Sentinel watching the trend
- Day 2Sentinel flags optical receive power moving steadily away from its own baseline on this port. No threshold involved.
- Day 2ONT error counters, retransmission rates, environmental data and the port maintenance history correlated.
- Day 2Probable cause: port-level optical degradation. Comparable signature found on two previously replaced ports.
- Day 2Projection attached: at the current rate, service-affecting failure is likely inside seven to ten days.
- Day 2Action Ticket raised proposing planned replacement in the next maintenance window, with the affected-subscriber list attached.
- Day 4Approved and executed in a planned window. No emergency dispatch, no business-hours outage.
7 days of warning · planned work instead of an emergency
There is no MTTR figure to quote here, because the point of this pattern is that the incident does not happen. The comparison is between a planned maintenance job and an emergency dispatch with a business-hours outage attached.
The cost difference between those two is something your field operations team already knows precisely: crew utilisation, overtime, the planned work that got displaced, and the truck roll that was not scheduled.
The mechanism is baselining each port against its own history rather than against a global threshold. A port that is inside tolerance but moving steadily in the wrong direction is a signal, and the projection is what converts that signal into a maintenance decision someone can approve.
The churn argument is the one worth taking to your commercial team, and it is also the one to be most careful with. We are not going to put a number on it, because attributing churn to a specific port over a specific week requires data we would be guessing at. What we will say is that the intermittent period is invisible in your fault statistics and visible in your customers experience, and those two facts together are the whole problem.
Why the number is what it is
There is no MTTR figure to quote here, because the point of this pattern is that the incident does not happen. The comparison is between a planned maintenance job and an emergency dispatch with a business-hours outage attached.
The cost difference between those two is something your field operations team already knows precisely: crew utilisation, overtime, the planned work that got displaced, and the truck roll that was not scheduled.
The mechanism is baselining each port against its own history rather than against a global threshold. A port that is inside tolerance but moving steadily in the wrong direction is a signal, and the projection is what converts that signal into a maintenance decision someone can approve.
The churn argument is the one worth taking to your commercial team, and it is also the one to be most careful with. We are not going to put a number on it, because attributing churn to a specific port over a specific week requires data we would be guessing at. What we will say is that the intermittent period is invisible in your fault statistics and visible in your customers experience, and those two facts together are the whole problem.
The mechanism is baselining each port against its own history rather than against a global threshold.
Who decides to press go
Replacing a PON port takes the subscribers behind it out of service for the duration, so it is scheduled rather than executed on detection.
Diagnostic steps run under policy: additional optical sampling, error counter collection, comparison against previously replaced ports. All reversible, none service affecting.
The Action Ticket proposes a planned maintenance window with the affected-subscriber list, the projected failure date and the rollback path attached. Field operations approves and schedules it.
The projection is what makes this approvable. The approver is not asked to act on a hunch about a port, they are shown a slope and a date.
Optical receive power on one PON port moving steadily away from its own baseline. Inside tolerance throughout. Error counters rising on a subset of ONTs behind it.
ONT error counters, retransmission rates, environmental data and port maintenance history correlated. Signature matched against two previously replaced ports.
Action Ticket raised proposing planned replacement, with a projected failure window and the affected-subscriber list. Held for field operations to schedule.
Degradation signature added to the port health model. Baselined optical power promoted to a monitored trend across the PON estate rather than a threshold check.
This is a platform capability, not a published customer deployment for a named fibre operator. The mechanism, which is topology-aware correlation followed by governed MOP execution with pre-check, post-check, rollback and approval gating, is running in production today; see carrier-scale observability across 27,000+ devices and closed-loop network automation. The access-network scenario above applies that same mechanism to a PON estate. The timings shown are modelled, not measured at a named operator.
If the action carries no service impact
Diagnostic steps run under policy: additional optical sampling, error counter collection, comparison against previously replaced ports. All reversible, none service affecting.
If it takes subscribers out of service, as a replacement does
The Action Ticket proposes a planned maintenance window with the affected-subscriber list, the projected failure date and the rollback path attached. Field operations approves and schedules it.
What Sentinel did, step by step
- ObserveOptical receive power on one PON port moving steadily away from its own baseline. Inside tolerance throughout. Error counters rising on a subset of ONTs behind it.
- InvestigateONT error counters, retransmission rates, environmental data and port maintenance history correlated. Signature matched against two previously replaced ports.
- ActAction Ticket raised proposing planned replacement, with a projected failure window and the affected-subscriber list. Held for field operations to schedule.
- OptimizeDegradation signature added to the port health model. Baselined optical power promoted to a monitored trend across the PON estate rather than a threshold check.
Bring us a port that failed last quarter
We will look at its optical history and show you how many days of warning were sitting in it.