Success rate slipped four points. Nobody filed a ticket.
Nothing was down, so nothing alerted. Four points of authorisation rate is not an incident on any dashboard, it is just a slightly worse morning that repeats until someone notices.
What actually happens
The outages get post-mortems. The four-point slips get absorbed into the monthly numbers and explained as seasonality.
Payment success rate runs at about 97.2 percent. Over a morning it settles at 93.1 percent. Every component reports healthy. No acquirer is down, no gateway is throwing errors, and the API is returning 200s on the transactions that do complete.
Four points sounds small until you convert it. On meaningful volume it is thousands of customers who tried to pay and could not, most of whom will not try again immediately and some of whom will not try again at all.
The reason nothing alerts is structural. Success rate is a ratio, and the monitoring on each individual component is looking at that component. One acquirer route quietly increasing its soft decline rate does not break anything, it just declines slightly more often, and the aggregate absorbs it until it does not.
When someone does notice, usually from a business dashboard rather than an operational one, the investigation starts from the aggregate and works downward. Which route, which issuer, which card type, which geography, which BIN range. That decomposition is a day of work with a spreadsheet, and it is the same day of work every time.
The most frustrating version of this is when the degraded route recovers on its own before anyone finishes the analysis, leaving a dip in a chart, no cause, and no confidence that it will not happen again next week.
A four-point drop in authorisation rate is not a system failure. It is revenue leaving quietly, through a component that is technically working.
The same morning, two ways
The question is not how fast you can restore a route. It is how long the leak runs before anyone can name it.
Illustrative, not measured. The times below model a scenario built from the patterns we see in production estates. They are not timings recorded at a named customer. The point is the shape of the clock, not the totals: check it against your own last ten incidents.
Today, found in the business numbers
- 08:00One acquirer route begins soft-declining more often. All components report healthy.
- 08:00 ↓ 13:30WaitingAggregate success rate drifts from 97.2 to 93.1 percent. No operational alert. No ticket.
- 13:30Commercial team queries the drop on a business dashboard.
- 13:30 ↓ 17:00WaitingDecomposition by route, issuer, card type and BIN range, done by hand.
- 17:00One acquirer route identified as the source of the excess declines.
- 17:40Traffic reweighted away from that route. Rate recovers.
~9 hours of leakage · found commercially, not operationally
With Sentinel watching the ratio
- 08:26Sentinel flags the composite: success rate falling against its own baseline for this hour and weekday.
- 08:28Rate decomposed automatically by route, issuer, card type, geography and BIN range.
- 08:31One acquirer route carries the entire excess. Soft decline codes on that route cluster on a single reason.
- 08:32Acquirer status feed and recent routing changes correlated. No published incident on the acquirer side.
- 08:33Action Ticket raised with the reweighting MOP, the projected recovery and the rollback path. Payments lead notified.
- 08:47Approved. Traffic reweighted. Sherlock confirms the rate returns to baseline within the hour.
~20 minutes · leakage measured in minutes, not a working day
Every minute in the first column is a real cost that never appears in an incident report, because no incident was ever opened. That is what makes this pattern durable: it does not fail loudly enough to enter the process that would fix it.
The decomposition is the expensive part, not the remediation. Reweighting traffic away from a route is a routine action. Working out which route, from an aggregate that is only four points off, is the day of work.
Sentinel treats the ratio as the monitored object. Success rate is baselined against its own pattern for that hour and weekday, and when it moves the decomposition runs automatically rather than being commissioned as an analysis task.
This is worth testing against your own history. Take the last twelve months of success rate and mark every dip of more than two points. Then check how many have a corresponding incident record. The ones that do not are the ones this addresses.
Why the number is what it is
Every minute in the first column is a real cost that never appears in an incident report, because no incident was ever opened. That is what makes this pattern durable: it does not fail loudly enough to enter the process that would fix it.
The decomposition is the expensive part, not the remediation. Reweighting traffic away from a route is a routine action. Working out which route, from an aggregate that is only four points off, is the day of work.
Sentinel treats the ratio as the monitored object. Success rate is baselined against its own pattern for that hour and weekday, and when it moves the decomposition runs automatically rather than being commissioned as an analysis task.
This is worth testing against your own history. Take the last twelve months of success rate and mark every dip of more than two points. Then check how many have a corresponding incident record. The ones that do not are the ones this addresses.
Sentinel treats the ratio as the monitored object.
Who decides to press go
Reweighting live payment traffic changes customer outcomes immediately, so the routing change is proposed rather than taken.
The Action Ticket executes under policy: pre-check, execute, post-check, verify, with automatic rollback if the success rate does not recover. Nobody is paged.
The ticket is raised and held with the proposed weighting, the projected recovery, the affected volume and the rollback path attached. The payments lead approves before traffic moves.
The fourteen minutes between the ticket and the reweight is a human deciding whether to move customer traffic. That is a decision that should have a name attached to it.
Aggregate success rate falling against its own baseline for this hour and weekday. All individual components reporting healthy. No error rate breach anywhere.
Rate decomposed by route, issuer, card type, geography and BIN range. One acquirer route carries the excess. Soft decline codes cluster on a single reason. Acquirer status feed shows no published incident.
Action Ticket raised with the reweighting MOP, projected recovery and rollback. Payments lead notified. Nothing moves until approved.
Success rate baselined per route rather than only in aggregate. Soft decline reason codes promoted to first-class signals. Acquirer added to the degradation watch list.
This is a platform capability, not a published customer deployment. The mechanism, which is correlated investigation followed by governed MOP execution with pre-check, post-check, rollback and approval gating, is running in production today in our carrier estates; see closed-loop network automation and carrier-scale observability across 27,000+ devices. The payments scenario above is that same mechanism applied to a transaction estate. The timings shown are modelled, not measured at a named payments provider.
If the action carries no service impact
The Action Ticket executes under policy: pre-check, execute, post-check, verify, with automatic rollback if the success rate does not recover. Nobody is paged.
If it carries service impact, as reweighting does here
The ticket is raised and held with the proposed weighting, the projected recovery, the affected volume and the rollback path attached. The payments lead approves before traffic moves.
What Sentinel did, step by step
- ObserveAggregate success rate falling against its own baseline for this hour and weekday. All individual components reporting healthy. No error rate breach anywhere.
- InvestigateRate decomposed by route, issuer, card type, geography and BIN range. One acquirer route carries the excess. Soft decline codes cluster on a single reason. Acquirer status feed shows no published incident.
- ActAction Ticket raised with the reweighting MOP, projected recovery and rollback. Payments lead notified. Nothing moves until approved.
- OptimizeSuccess rate baselined per route rather than only in aggregate. Soft decline reason codes promoted to first-class signals. Acquirer added to the degradation watch list.
Bring us a dip you never explained
We will decompose it against your own data and show where it came from.