Telecom Circuit Uptime: What SLA Metrics to Verify on Every WAN Link
What the availability number on a telecom SLA really commits to: uptime, MTTR, latency, jitter, packet loss, measurement points, and the credit mechanics to verify on every WAN link.

On this page
- What an Availability Percentage Commits To
- MTTR and MTBF: The Repair Clock and the Relapse Rate
- Latency, Jitter, and Packet Loss: The Degradation Trio
- The Bandwidth Promise: CIR and the Silent Brownout
- Where the SLA Is Measured Decides Who Wins
- Service Credits: What a Breach Actually Pays
- Build the Baseline Before You Need It
Every business circuit is sold with the same reassuring number. The quote says 99.99 percent availability, procurement files it away, and nobody reads the contract again until the night the link drops. When that happens, the discovery begins: the availability percentage is not a promise about your experience, it is a carefully defined claim about a specific measurement, taken at a specific point, over a specific period, excluding a specific list of events. The value of an SLA is decided almost entirely in those definitions, and most teams never look at them until they are already arguing with the carrier who holds the outage records.
This guide covers the SLA metrics to verify on every WAN link, whether the service is MPLS, dedicated internet access, metro Ethernet, or dark fiber: what an uptime figure actually commits to, how the repair clock runs, the latency, jitter, and packet loss thresholds that degrade applications long before a circuit counts as down, the bandwidth guarantee that congestion can silently erode, and the measurement boundaries that decide who wins a dispute. By the end you will know which numbers to check with a calculator and which clauses to check with a lawyer.
What an Availability Percentage Commits To
Your circuit quote leads with a percentage, usually 99.9 or 99.99, and that number is where most SLA conversations begin and end. Availability is best understood as a downtime budget disguised as a compliment. On a 30 day month, 99.9 percent leaves about 43 minutes of allowed unavailability and 99.99 percent leaves just over 4 minutes; spread across a year the same two figures allow 8.8 hours and 52.6 minutes respectively. That is a wide enough spread that the first question to ask any carrier is which tier you are actually buying, because the sales deck and the contract do not always agree.
The budget only means something if you agree on what consumes it. Most contracts define unavailability as a complete loss of service at the demarcation point, which sounds precise and is actually narrow. A link that is up but degraded, congested, or dropping a meaningful percentage of packets does not breach the availability commitment in most contracts, because availability in this sense is a connectivity measurement, not a performance measurement. Latency, jitter, and packet loss live in separate clauses with their own thresholds, which is why a circuit can post a perfect 99.99 percent month while the applications running over it feel noticeably worse.
The exclusions list quietly shrinks the number further. Scheduled maintenance windows with proper notice are usually removed from the calculation entirely, as are force majeure events, power failures at your site, and faults in customer owned equipment. That is reasonable in isolation, but the exclusions accumulate: a month with no outages at all still cannot reach four nines if a two hour maintenance window was subtracted from the denominator. Ask the carrier to show you how the accounting works before you treat the headline number as a floor.
MTTR and MTBF: The Repair Clock and the Relapse Rate
Availability tells you how much downtime is allowed; MTTR, mean time to restore, tells you how long a single incident can run. The clause is usually written as a maximum, four hours for a metro fiber circuit being typical and eight hours for less critical links, with response commitments layered on top, such as acknowledging a ticket within 30 minutes and dispatching a technician by a set deadline. The fine print that matters more than the hours is when the clock starts and whether it runs continuously.
Carriers frame that clock differently. The version most favorable to them starts at dispatch, when a technician accepts the case, which ignores everything between your call and that moment. The version most favorable to you starts at detection, in the carrier’s network management system or at your own ticket open time, whichever is earlier, and it is a clause worth negotiating on any circuit that carries revenue. Equally important is the distinction between business hours and 24/7 restoration. A four hour MTTR with the clock running nine to five can mean a circuit stays down overnight and through the next morning before the carrier’s obligation resets, which is why the contract must state the schedule explicitly.
MTBF, mean time between failures, is the quieter of the two metrics and often the more telling. It describes how often outages happen rather than how long they last, and together the two numbers describe the shape of your risk. A circuit with one 43 minute outage has the same monthly availability as a circuit with ten four minute outages, but completely different operational meaning for a site running time sensitive applications. Ask the carrier for historical outage frequency on the specific path, not just the product brochure, because MTBF is where the real world shows up: the difference between a rural last mile that fails quarterly and a well maintained fiber path that has not blinked in years.
Latency, Jitter, and Packet Loss: The Degradation Trio
When a link stops working you notice immediately; when it degrades you notice slowly, through support tickets about slow pages, poor call quality, and unexplained application timeouts. The three numbers that govern that degradation are latency, jitter, and packet loss, and each has a specific contract question attached. For latency, ask how it is measured and between which points. A commitment of 15 milliseconds to the carrier’s nearest point of presence is not the same as a commitment of 15 milliseconds end to end between your sites, and the headline number on the datasheet is almost always the shorter of the two.
Jitter is the variation in latency between consecutive packets, and it is the metric voice and video feel first. A few milliseconds of variation is normal on a healthy carrier network, commitments of 5 milliseconds or less are common on premium products, and once the variation climbs past roughly 10 to 15 milliseconds, voice quality visibly suffers even when average latency looks fine. Packet loss is the third leg, typically committed at 0.1 percent or better on business circuits, and its real world impact depends on how the metric is sampled more than on the value itself, because a monthly average smooths away the short bursts that actually break sessions.
The same logic applies to every averaging scheme in the contract. A committed loss of 0.1 percent can be met with a flawless month punctuated by a lost burst every time a session needs to cross the network, and a latency commitment framed as a monthly mean can be met while individual minutes are far outside budget. Ask whether each metric is evaluated as a monthly average or in shorter intervals, five minute buckets are the practical standard, and whether percentile reporting such as p95 or p99 is available. An SLA you can only read as a monthly average is an SLA you cannot verify weekly, and verifiability is the entire point.
The Bandwidth Promise: CIR and the Silent Brownout
On Ethernet and carrier Ethernet services, the port speed and the committed information rate are different numbers, and the gap between them is where a carrier can meet every availability target while still shortchanging you. A circuit with a gigabit port and a 100 megabit CIR is fully available every day of the month while delivering nothing above the committed rate, and congestion on the carrier’s network silently turns that gap into reality. If the contract contains no throughput commitment at all, a permanently congested handoff can satisfy every other clause while your users feel the squeeze.
Verify the bandwidth commitment at installation and again at every renewal. Service activation tests based on ITU Y.1564 exercise the circuit at its full committed rate and measure throughput, frame loss, and delay in both directions, which is the same methodology the carrier’s own installers use, so asking for the results is a fair request rather than an adversarial one. A circuit that cannot pass its own turn up test at the committed rate is a circuit you should document now, while the goodwill around the sale still exists, because the same clause that let it underperform at install will cover it for the life of the contract.
Ask explicitly whether the SLA governs throughput at all. Many product SLAs cover connectivity only, with no commitment about how much traffic the link can carry, which makes your own utilization records the only real evidence you can lean on from the first billing month. Capture them consistently, because at renewal time a circuit running against its CIR supports a request for a higher committed rate or a discounted price, while a claim made without utilization data reads as anecdote no matter how correct it is.
Where the SLA Is Measured Decides Who Wins
The single most consequential location in an SLA is not a percentage, it is a coordinate. Everything the carrier commits is measured at a specific point on the circuit path, and the choice of that point determines whether a real world outage counts against them or not. The typical path runs from your customer edge router to the demarcation point, through the local loop, into the carrier’s point of presence, and across its core backbone, and the measurement point a carrier prefers is the boundary of its own network, where the telemetry is cheap and the coverage is complete.
If availability is verified at the carrier’s point of presence, the local loop between the demarc and the POP is invisible to the measurement, which matters because the local loop is the segment that fails most often: the fiber that gets dug up, the copper that corrodes, the aerial plant that a truck takes down. A carrier measuring at its own edge can post a clean availability number while your site sits through a multi hour local loop outage, because the event never registers in its measurement. Insist that availability be verified at your customer edge, or at minimum that both sides keep independent records and reconcile them monthly.
The same scrutiny applies to who declares an outage and when the clock starts. Some contracts count only outages the carrier detects or acknowledges through its network management system, which leaves a silent failure, a link that is down on your side but invisible to the carrier’s monitoring, uncounted even though your users experienced it. Ask how the carrier detects failures, whether proving a fault requires a truck roll, and whether raw performance reports are available on request. A carrier that streams raw reports is a carrier whose SLA you can audit; a carrier that offers only monthly summaries is a carrier you should audit harder.
Service Credits: What a Breach Actually Pays
The remedy for a missed SLA is a service credit applied to your bill, and reading the credit mechanics is where most of the surprise lives. The typical formula credits a percentage of the monthly recurring charge per outage hour, with steeper tiers for failures that blow past the MTTR commitment, but three details decide what the clause is worth in practice: the de minimis threshold, commonly excluding incidents shorter than 15 to 30 minutes from any credit, the cap, usually between 30 and 100 percent of one month’s recurring charge, and the claim window, often 30 to 90 days from the event.
The de minimis threshold deserves particular attention because it changes the meaning of availability. A circuit with a monthly budget of four minutes can pass through ten brief two minute flaps, exceed its budget fourfold, and still produce zero credit under a contract that excludes incidents under 15 minutes, because the threshold swallows every individual event. The headline percentage and the credit outcome are only connected when you read the incident boundary, and on circuits running applications sensitive to short outages, a lower threshold or per event credits are worth negotiating before you sign rather than after the first bad month.
Credits are also rarely automatic. Most carriers require a claim with evidence of the outage start and end times, which means your outage log, not the carrier’s, is what stands between you and the money. Industry estimates have long suggested that a majority of eligible SLA credits go unclaimed because teams lack the records or miss the filing window, and the difference between collecting and not collecting is almost always documentation discipline rather than carrier goodwill.
Build the Baseline Before You Need It
None of these clauses can be verified retroactively. The work happens when the circuit is ordered, when the tests are run, and when the commitments are written into the record your team will consult on the day of the incident. At turn up, run the activation test and keep the result. At contract signing, write down what matters: the availability tier, the MTTR schedule, the latency, jitter, and loss targets, the CIR, the measurement point, the threshold, the cap, and the claim window. Across the life of the circuit, maintain a continuous outage log with timestamps, so the next quarterly review starts from your evidence rather than the carrier’s summary.
Review every circuit against the carrier’s reports on a regular cadence, quarterly is a reasonable rhythm, and treat a gap between what you measured and what the carrier reported as the beginning of a claim rather than the end of a discussion. Circuits that meet their metrics are negotiation material too, because a path with years of clean records supports asking for a price reduction at renewal, while a path with undisputed breaches supports asking for credits before renewal.
The discipline is an inventory problem as much as a monitoring problem, because a claim needs the commitment sitting next to the incident, per circuit, on record. On a multi site WAN that means recording SLA commitments, circuit identifiers, and outage history beside each circuit and link in the topology, in the same place the rest of the circuit data lives, rather than in another spreadsheet nobody reconciles. With the commitments stored on the circuit record, a quarterly compliance review becomes a reading of the page instead of a reconstruction of the contract, and the credits you collect are simply the compensation for keeping that record straight.
Frequently Asked Questions
How many minutes of downtime is 99.99 percent uptime?
What is the difference between MTTR and MTBF in a circuit SLA?
What packet loss percentage is acceptable on a WAN link?
Do telecom carriers actually meet their availability SLAs?
How are SLA credits calculated for a circuit outage?
Is 99.99 percent uptime good for a business internet circuit?
Stop reaching for a spreadsheet
Obelinf keeps every subnet, device, circuit, and rack in one live source of truth, with audit logs and a topology view. Free for personal use.
Related Articles

Fixed Wireless and LEO Satellite as ISP Backup: When to Use Them Over Broadband
When a second broadband line is not diverse enough, fixed wireless and LEO satellite give you true path independence. Learn where each fits, what it costs, and how to size and test a wireless backup that actually fails over.
Read more
How to Switch ISPs Without Downtime: IP Renumbering, DNS and Cutover Planning
Switch ISPs without downtime by planning IP renumbering, DNS TTL strategy, and a parallel cutover that keeps your network reachable throughout the migration.
Read more
ISP Diversity Audit: How to Prove Two Providers Don't Share the Same Fiber
Two ISP contracts do not guarantee diverse fiber. Learn how to audit physical paths, local loop ownership, and routing to prove your providers are truly independent.
Read more