QoS Explained: Prioritizing Traffic on Your Network
Quality of Service decides which packets go first when a link is full. Learn how classification, DSCP marking, queuing, shaping, and policing fit together, and how to build a policy that holds up under real congestion.

On this page
At some point on almost every network, demand exceeds capacity. A WAN circuit saturates during the nightly backup, an office uplink fills when a hundred people watch a town hall, a wireless radio carries more phones than it can serve at once, or a VPN tunnel turns a fast branch link into a slow one. When that happens, an interface cannot send every packet the moment it arrives. It has to decide what goes now and what waits, and if it makes that decision by default, the answer is simply whichever packet happened to arrive first.
Quality of Service is the set of tools that turns that decision into a deliberate one. QoS does not add bandwidth, and it cannot make a link faster than the circuit or the radio allows. What it does is choose which traffic is protected when the link is full, which traffic can wait, and which traffic is dropped first. This guide walks through the mechanisms in the order they operate, explains where marking stops mattering and queuing begins, and shows how to assemble those pieces into a policy that behaves the way you intended under real congestion.
What QoS Can and Cannot Do
Start with the limits, because misunderstanding them leads to policies that never fire or that promise more than the hardware can deliver. QoS is a congestion management discipline, which means it only matters when a queue is building. On an idle gigabit link it changes nothing observable, because the packets leave in the same order and at the same speed. If your links are never congested, QoS is overhead. If they are congested, QoS is the difference between a dropped call and a file transfer that simply slows down.
The reason prioritization matters at all is that applications respond to stress differently. A bulk transfer uses TCP, which interprets delay and loss as signals to slow down and backs off gracefully. A voice call or a live video stream does not adapt that way. It produces packets at a steady rate and treats delay, jitter, and loss as damage to the call itself. Buffering is the wrong answer for the second group: filling a deep queue to avoid drops lowers loss while raising latency, the classic bufferbloat trade that makes a conversation feel broken even though no packets were lost. The goal of QoS is to keep sensitive flows in a short, protected queue, let elastic flows share the rest, and avoid so much buffering that interactive traffic suffers delay.
It also helps to know where QoS earns its keep. The links that benefit are the ones that actually run hot: the carrier or ISP edge, aggregation links where many users share one path, VPN and tunnel headends, and wireless, where a single radio serves every client in range. A well provisioned data center core with large amounts of headroom usually gains little beyond a simple default, because the queues rarely build. Deploy QoS at the bottlenecks and skip it where the link is never full.
Four Steps: Classify, Mark, Queue, Schedule
Every QoS implementation, regardless of vendor, is built from the same four steps. Classification identifies which traffic belongs to which class, using access lists, port numbers, application signatures, or an existing marking. Marking writes the result into the packet, either as a DSCP value in the IP header or a CoS value in the Ethernet tag. Queuing places each packet into the queue for its class. Scheduling decides which queue the interface serves next and, under pressure, which packets it drops.
The important design detail is that these steps do not all run in the same place. Classification and marking belong at the edge you control, close to the source, where you can see the traffic before it blends together: an access switch, a VPN concentrator, or a branch router. Queuing and scheduling belong on every interface that can congest, which is often a different device entirely. Marking is only a label. It has no effect until something downstream reads it and acts on it, and a policy that marks traffic perfectly but never queues it by class has done nothing at all. This is the most common structural mistake in QoS: half a design.
The other half of the design is the trust boundary. You mark as close to the source as possible, but you only trust markings from devices you control. An IP phone marks its own voice as EF and the attached computer’s data as best effort, so an access port that serves a phone can trust the phone. A port that serves a regular computer cannot, because the operating system on that computer can mark every packet as EF and steal the priority queue for itself. The safe default is to re-mark everything entering an untrusted port to best effort and to assign priority only after your own classifier has looked at it.
DSCP and CoS: Marking Traffic With Intent
The two markings you will work with most are DSCP and CoS, and they live at different layers. DSCP is a 6 bit field in the IPv4 type of service byte, mirrored in the IPv6 traffic class field, which gives it 64 values and keeps it in the packet as traffic crosses routers. CoS, also called the priority code point, is a 3 bit field in the 802.1Q VLAN tag with 8 values. It travels only across tagged Layer 2 links and is lost the moment a frame is routed, so it carries priority across a trunk but not across the network. The practical rule is to use DSCP for the end to end intent and to derive CoS from DSCP on switch trunks.
DSCP values are organized into groups called per hop behaviors, and a small number of them cover almost every design. Expedited Forwarding, DSCP 46, is the priority class for voice media, and it expects to be queued briefly and served first. Assured Forwarding defines four classes, AF1 through AF4, with three drop precedences each, so the network can shed the least valuable packets within a class before touching the rest. The Class Selector values mirror the older IP precedence bits, and CS6 is reserved for network control traffic such as routing protocols. Best effort is DSCP 0, the default for anything unmarked.
The Assured Forwarding drop precedences are worth understanding because they are what makes graceful degradation possible. AF41, AF42, and AF43 all belong to the same class, but a congestion avoidance mechanism can drop the higher numbered variants first when the class exceeds its share. That lets you mark a video conference as AF41 at the source and downgrade a less important stream to AF43, so under pressure the network keeps the sessions that matter and trims the marginal ones instead of dropping everything evenly.
Two boundaries keep the map honest. Network control at CS6 exists for routing and management traffic and should never carry user data, because congestion in that class can destabilize the control plane. EF is the highest class for actual user traffic, and it needs to stay small. If you mark every application that matters as EF, you have rebuilt the original problem inside a single queue, because the network can no longer tell your voice call from your software update.
Queuing and Scheduling: Who Goes First
Once packets are marked and classified, an interface has to decide what to send next, and the algorithm it uses determines whether your priorities mean anything. The default is first in, first out: a single queue where every packet waits its turn and a full queue drops whatever arrives next. That is fine for a link with headroom and useless for a congested one, because a large file transfer and a voice packet receive identical treatment.
Strict priority queuing fixes that by serving the high priority queue completely before touching any other, which is exactly what interactive traffic needs. It also carries a real risk. If the priority queue has more traffic than the link can carry, everything else waits indefinitely, a condition called starvation. That is why a strict priority class should be small and policed: you give voice the front of the line, but you cap how much of the line it can take. A common rule is to keep the priority class at or below roughly a third of the link and to police it, so that a flood of EF marked traffic cannot consume the rest.
The remaining classes are handled by weighted algorithms that give each queue a share of the interface during congestion. Weighted fair queuing, deficit weighted round robin, and class based weighted fair queuing all express the same idea in different ways: each class receives a guaranteed minimum bandwidth, and when a class is idle its share becomes available to the others. The important distinction is between bandwidth and priority. A bandwidth guarantee is a floor that applies under congestion and can be exceeded when the link is idle, while a priority queue is served first and capped. Voice needs the second behavior. Video and business applications generally do fine with the first.
Two smaller mechanisms complete the picture. Congestion avoidance, usually implemented as random early detection or its weighted form, drops a percentage of packets before a queue is completely full so that many TCP flows slow down together instead of all hitting the wall at once, a failure mode called global synchronization. The weighted form uses the Assured Forwarding drop precedence to decide which packets to shed first. The low latency voice queue skips these mechanisms on purpose, because a dropped voice packet is never retransmitted, so the queue stays short and predictable rather than deep and adaptive.
Shaping vs Policing: Two Ways to Enforce a Rate
Queuing decides the order of packets that are already allowed onto the link. Shaping and policing decide how much traffic is allowed in the first place, and they differ in how they treat the excess. A shaper buffers traffic above the target rate and releases it later, producing a smooth flow at or below the configured rate. A policer measures traffic against the rate and acts immediately on anything over it, either dropping it or re-marking it to a lower class. Shaping trades delay for smoothness. Policing trades smoothness for low delay.
The choice usually follows the direction of the traffic. On outbound traffic toward a link that is slower than the local interface, shaping is the better tool, because it converts a sudden burst into a paced stream that the far end can absorb and keeps TCP from overreacting to drops. Inbound traffic is harder, because by the time it reaches you it has already crossed the bottleneck. You cannot shape what has not arrived, so you police it, re-mark it, or rely on the provider to enforce the contract for you.
This leads to the single most important implementation detail in enterprise QoS, which is to shape before you queue. If a router has a one gigabit interface connected to a hundred megabit circuit, it can transmit faster than the provider will accept, so the local queue never builds and your carefully designed queues never activate. At the same time, the provider’s policer drops the excess, and the packets it drops are the ones that happened to arrive after the limit, which may include your voice. Configuring a parent shaper at the purchased rate and applying the class based queueing policy inside that shaped envelope forces the prioritization to happen on your equipment, where you control it.
Encryption and encapsulation can undo marking without any error message. A tunnel such as IPsec or GRE carries your original packet inside a new outer header, and unless the device is configured to copy the DSCP value into that outer header, the network in between sees only an unmarked tunnel and treats everything as best effort. Check that your VPN headends preserve DSCP, and remember that the public internet generally ignores it. Treatment with real guarantees along the whole path requires a carrier that exposes a service class and publishes how it maps your markings to its own, usually to MPLS EXP bits at the handoff.
Designing a Policy That Survives Production
A good policy is smaller than most people expect. Platforms rarely offer unlimited queues, and many switches expose a fixed set of hardware queues, sometimes as few as four or eight, that your software classes map onto. Trying to run twelve classes on a platform with four queues forces the device to merge them in ways you did not intend. Start by writing down the handful of behaviors you actually need: a priority class for voice, a guaranteed class for real time video, a guaranteed class for business applications, a best effort class, and a scavenger class for backups and bulk transfers that can be sacrificed first.
Then do the bandwidth arithmetic before you commit. Voice is the class most often sized by guesswork. A single G.711 call is about 64 kilobits per second of payload, plus IP, UDP, and RTP headers, plus the Layer 2 overhead on the access and trunk links, which brings a call to roughly 80 to 100 kilobits per second in a typical VLAN environment. Compressed codecs cut that substantially, but the header overhead becomes proportionally larger. Multiply the per call rate by the number of concurrent calls you expect at the busiest hour, add margin, and that is your voice reservation. If the result is more than about a third of the link, you do not have a QoS problem, you have a capacity problem.
Link speed changes the calculus in a way that is easy to forget. Serialization delay is the time it takes to place a single packet on the wire, and it grows as the link gets slower. A 1500 byte packet takes about 12 milliseconds to serialize on a 1 megabit link, about 24 milliseconds on 512 kilobits, and nearly 50 milliseconds on 256 kilobits, all before it has traveled anywhere. Since interactive voice targets a one way budget well under 150 milliseconds, a single large packet can consume a meaningful chunk of it. On links below roughly 768 kilobits, fragmentation and interleaving break large packets into smaller pieces so a voice packet can slip into the stream, and header compression shrinks the per packet overhead further.
Wireless deserves separate treatment because the bottleneck is a shared radio rather than a wire. Wi-Fi uses its own priority scheme, Wi-Fi Multimedia, with four access categories for voice, video, best effort, and background, and the access point maps the DSCP or CoS value of each frame onto one of them. If your wired policy marks voice as EF but the wireless controller is not mapping that to the voice access category, the phone’s traffic competes with every laptop on the same radio. Verify the mapping end to end rather than assuming that because the LAN marks correctly, the wireless side does too.
Finally, design both directions and document where the policy runs. QoS is applied per interface and per direction, so a policy that protects outbound voice does nothing for the inbound stream returning to your users. A class model is only as reliable as the map of where it is configured, which interfaces carry which classes, and which ports are trusted. Recording that alongside the device and its interfaces, the way device inventory keeps interface records next to the hardware they belong to, means the next person to touch a link can see whether a change preserved the design instead of rediscovering it after a call quality complaint.
Verifying QoS Under Load
QoS that is never tested under congestion is a hypothesis, not a guarantee. An idle link will happily report zero drops in every class and tell you nothing, because no queue filled. To validate a policy you have to create the pressure the policy exists to handle: generate enough traffic to saturate the link, then place a real or synthetic voice and video call on top and watch what happens to each class.
Start with the per class counters on the egress interface. You want to see the offered rate, the drop rate, and the queue depth for each class. A voice class should show zero drops and a shallow queue, while sustained queue depth in any class is the signature of a link that is genuinely saturated. If the priority class reports drops, either the policer is set too low or the traffic marked into it is larger than you thought. If best effort drops while the voice class sits empty, the problem is upstream of the interface, and your classification is probably missing the traffic you meant to protect. Then measure the experience itself. Round trip latency, jitter, and loss give you the numbers that map to call quality, and a mean opinion score estimate turns them into something a helpdesk can act on. Voice typically wants one way latency under 150 milliseconds, jitter under about 30 milliseconds, and loss under one percent, and a policy that meets those targets on a healthy link but not under load has not been finished.
It also helps to know what the network thinks each flow is. Passive flow export such as NetFlow or IPFIX can report traffic volumes per DSCP value, which shows whether your marking is landing where you intended, and active probes can test reachability and latency on a schedule. The failures worth watching for are almost always at the seams: a policy that exists on the WAN router but not on the switch uplink in front of it, a carrier that strips or rewrites DSCP so your classifications never reach the far side, a tunnel that hides the markings, and a priority class so large that it starves the guaranteed classes below it.
What to Take Away
QoS is a way to decide who waits when a link is full, and that is the whole of its promise. It cannot create bandwidth, and it does nothing on a link that is not congested, so the first useful step is to find the links that actually hurt: the carrier edge, the oversubscribed uplink, the shared radio, the tunnel. On those links, build a small class model, mark once at a trust boundary you control, queue and schedule at the egress, and shape to the real rate so the queues engage where you can see them.
Keep the protected classes small and policed, give everything else a guaranteed share, and drop the least valuable traffic first. Then prove it under load, because a policy that has never seen congestion has never been tested. A practical starting point is one link, a priority class for voice, a guaranteed class for video and business applications, and a scavenger class for the traffic you would sacrifice, expanded only when you have evidence that a new class is needed.
Frequently Asked Questions
What does QoS actually do?
What is the difference between DSCP and CoS?
Is QoS necessary on a high bandwidth network?
What is the difference between shaping and policing?
Which DSCP value is used for voice?
Stop reaching for a spreadsheet
Obelinf keeps every subnet, device, circuit, and rack in one live source of truth, with audit logs and a topology view. Free for personal use.
Related Articles

How to Manage IP Addresses in a Homelab
A practical tutorial for planning subnets, assigning addresses, and documenting a homelab IP plan in Obelinf.
Read more
Kubernetes CIDR Planning: A Practical Guide
Plan Kubernetes Pod, Service, Node, and Load Balancer networks with practical CIDR sizing, overlap checks, growth assumptions, and a worked example.
Read more
Multi Cloud CIDR Planning: Avoid VPC and VNet Overlaps
Plan private CIDR space across AWS VPCs, Azure VNets, and Google Cloud VPCs without creating overlaps that break peering, VPNs, or future hybrid connectivity.
Read more