Network Performance: Latency, Throughput, Jitter & Loss
Understand bandwidth, throughput, latency, RTT, jitter, and packet loss with practical measurements, tail percentiles, and production diagnostic guidance.
This guide explains how bandwidth, throughput, goodput, latency, RTT, jitter, and packet loss describe different parts of network performance. It shows how to measure real paths with ping, iperf3, and request timings, then interpret tail percentiles, congestion, and application traces without mistaking a network symptom for the root cause. Use the checklists and failure examples to choose safe tests and diagnose slow or unstable connections.
Network Performance: Latency, Throughput, Jitter & Loss
Introduction
Network performance depends on more than link speed. Latency, jitter, packet loss, throughput, and capacity affect how quickly a workload completes, and averages can hide the tail delays users notice. Measurements only make sense when their endpoints and workload are clear.
This guide defines the main performance metrics, shows how to measure them, and explains how they interact under load. It also covers practical interpretation, failure scenarios, and monitoring checks that help separate network symptoms from application response time.
The measurements that describe a network
Bandwidth and throughput
Bandwidth is the maximum rate a link can carry under stated conditions. A home plan advertised at 500 Mbps describes a nominal capacity, not a promise that every application will transfer data at that rate. Shared links, protocol overhead, Wi-Fi interference, remote server limits, and competing traffic all reduce what an individual flow can use.
Throughput is the data rate actually delivered over a measurement interval. It is usually reported in bits per second: Kbps, Mbps, or Gbps. File sizes, by contrast, are often shown in bytes; 8 bits make 1 byte, before accounting for headers and other overhead. A sustained download at 80 MB/s is roughly 640 Mbps of payload rate.
Application goodput is narrower: useful application data delivered per second, excluding retransmissions and protocol headers. A transfer can have high link throughput but lower goodput when packets are lost or the protocol spends capacity on control traffic.
A large bandwidth number does not guarantee a fast small request. A link with plenty of capacity can still have high latency, and a single TCP flow may take time to ramp up to the available rate.
Latency and RTT
Latency is the time for data to travel between two points. One-way latency is difficult to measure accurately unless clocks at both ends are synchronized. Round-trip time (RTT) measures how long a probe or request takes to travel to a destination and for a response to return. It is commonly measured in milliseconds (ms).
RTT includes more than distance. Routing, queueing, packet processing, and the endpoint’s response all contribute. A network probe may get a different priority or route than application traffic, so treat its number as one observation of a path, not a universal property of every request.
Latency affects how quickly an interaction can start and how long dependent request chains take. If a page makes six serial requests and each spends 40 ms waiting on the network, those waits alone can add about 240 ms. Parallel requests overlap some of their wait time; serial dependencies do not.
Jitter
Jitter describes variation in packet delay over time. A voice packet arriving in 25 ms followed by one arriving in 90 ms is harder to play smoothly than packets arriving at a steady 55 ms, even though the average is similar.
People use different formulas for jitter, such as the variation between consecutive packet delays or a percentile spread over a window. Always check how the tool defines it before comparing dashboards. Real-time media often uses a jitter buffer to absorb variation; a larger buffer smooths playback but adds delay.
Packet loss
Packet loss is the share of packets that fail to reach their destination or are discarded along the path. It is reported as a percentage or count over a time window. A ping tool may show 1 lost packet out of 100 probes, or 1% loss, but that sample is too small to characterize a busy production path on its own.
Loss has different effects by protocol. TCP generally retransmits missing data, which preserves the byte stream but adds delay and can reduce throughput. Real-time UDP traffic may not retransmit late packets; a lost voice frame can instead sound like a brief gap. Some loss also happens at the receiver when the application cannot read data quickly enough, so packet counters should be interpreted with host and application metrics.
Why tail percentiles matter
Averages hide the slow end of a distribution. Suppose 990 requests finish in 30 ms and 10 take 900 ms. The average is 38.7 ms, which sounds fine until a user receives one of the slow requests.
The p50 is the median: half of observations are at or below it. p95 is the value at or below which 95% of observations fall. p99 leaves the slowest 1% above it. Percentiles help answer how often users encounter a delay, but they depend on the sample size and time window. A p99 from 20 observations is not a reliable picture of the slowest one percent.
Track percentiles for the same request class and window. Combining a fast health check with a slow report-generation endpoint can make either distribution hard to understand. For critical services, pair latency percentiles with request volume, error rate, and saturation indicators.
Network metrics are not application response time
A network measurement usually describes one segment or a probe exchange. Application response time measures the user’s whole operation: DNS lookup, connection setup, TLS negotiation, network waits, server queues and code, database calls, response transfer, and sometimes browser rendering.
For example, a 25 ms RTT does not mean an API request takes 25 ms. The request may wait 120 ms in a server queue, spend 80 ms on a database query, and transfer a large response over a slow link. Conversely, a high RTT may have little effect on a cached local interaction.
Use distributed traces or client timing to break a request into spans. Compare the network segment with server processing and downstream dependencies before changing network settings. This pairs naturally with the separate guide to timeouts and network failures, which covers deadlines and safe retry behavior.
A practical measurement example
For a first pass, collect measurements from the same host and network path where the problem occurs. Record the time, destination, protocol, test duration, and whether traffic was idle or under load. A short sample can miss intermittent congestion.
ping can help reveal basic reachability and approximate RTT to a host that responds to ICMP. It cannot prove that an HTTPS API is healthy, show available throughput, or represent every route and traffic class. ICMP may be rate-limited or deprioritized. Pair it with an application-level request and a controlled throughput test.
With permission to generate test traffic, iperf3 can measure capacity between two endpoints you control. On the receiving host, run:
iperf3 --server
From the sending host, run a 30-second TCP test and report intervals:
iperf3 --client 203.0.113.20 --time 30 --interval 5
The example address is reserved for documentation; replace it with your own test endpoint. A result such as 82 Mbits/sec is measured throughput for that test flow, not a guarantee for every user or request. Avoid running throughput tests against services you do not own or during a production incident if they could add load.
To estimate request latency at the application boundary, make repeated requests and retain individual timings rather than just a single ping result:
for i in $(seq 1 20); do
curl -sS -o /dev/null -w '%{time_connect} %{time_starttransfer} %{time_total}\n' \
https://api.example.com/health
sleep 1
done
time_connect is connection establishment time, time_starttransfer measures time to first byte from the start of the request, and time_total is the full transfer duration. DNS and TLS timing can be collected with additional curl fields when needed. The health endpoint still may not exercise the same code path as the slow user request.
A simplified request path helps place the measurements:
graph LR
C[Client] --> D[DNS lookup]
D --> T[TCP and TLS setup]
T --> N[Network transit]
N --> Q[Server queue]
Q --> A[Application and dependencies]
A --> R[Response transfer]
How the metrics interact
These measurements often move together, but none can substitute for another.
- High latency, high bandwidth: large transfers may complete quickly once they start, while small serial exchanges still feel slow. Reduce round trips, reuse connections, or move work closer to users.
- Low latency, low throughput: a constrained link, congestion, a slow sender or receiver, or TCP flow limits may cap transfer rate. Measure with a sustained test and inspect endpoint CPU and socket state.
- Loss and throughput: TCP retransmissions consume time and capacity. Even modest loss can sharply reduce throughput on long-distance paths because recovery takes longer.
- Jitter and media quality: a low average delay can coexist with bursts of late packets. Check delay variation and loss over the same interval as user reports.
- Queueing under load: traffic can fill buffers, causing RTT to rise while throughput approaches the link’s capacity. This is often called bufferbloat. Compare idle and loaded latency rather than judging the line by peak Mbps alone.
A useful diagnosis asks which metric changed, for which users and path, and whether it changed before or after the application became slow.
When to Use
Use network performance measurements when requests differ by location, transfer size, protocol, or time of day; when audio/video quality degrades; or when a service shows rising latency without a corresponding increase in server work. A controlled comparison between a nearby and distant endpoint can help distinguish a local access problem from a broader path issue.
When NOT to Use
Do not use a single ping or one speed test as a release gate for an application. Avoid broad throughput tests on shared production links unless their traffic impact is understood. For user-facing latency, prefer real request timings and traces. For capacity planning, combine sustained tests with traffic profiles and endpoint limits.
Trade-Off Table
| Measurement approach | Useful for | Limitations |
|---|---|---|
| ICMP ping | A quick reachability and approximate RTT sample | May be blocked or deprioritized; does not exercise application processing or prove throughput. |
Controlled iperf3 test |
Sustained throughput between authorized endpoints | Generates load; a test flow may not match user traffic, path, or protocol. |
| Synthetic application request | Repeatable DNS, connection, and response timing | Covers only the chosen endpoint and request path; health checks may be cheaper than real work. |
| Real-user timing and tracing | User impact and request phase attribution | Needs sufficient volume, careful privacy controls, and consistent instrumentation. |
Production Failure Scenarios
| Scenario | What operators may see | Useful response |
|---|---|---|
| Wi-Fi interference or weak signal | Retries, variable RTT, low throughput for users on one access network | Compare wired and wireless samples; inspect access point channel use and signal quality; offer a wired path for critical equipment. |
| Congested uplink | Throughput plateaus, RTT rises during upload, interactive calls become choppy | Measure under load; shape or prioritize traffic where appropriate; reduce competing transfers and review access capacity. |
| Long-distance packet loss | TCP throughput falls despite a high-capacity link; p99 transfer time grows | Compare paths and endpoints, inspect retransmissions, and work with the network provider on routing or transport issues. |
| DNS or connection setup delay | High time-to-first-byte before server work begins; repeated connections are slow | Measure DNS, TCP, and TLS phases separately; use connection reuse and verify resolver health. |
| Server saturation mistaken for network trouble | Ping remains stable while API response time and server queue time rise | Trace requests, inspect CPU, worker queues, and dependency spans; scale or fix the saturated component. |
| UDP media bursts arrive late | Call quality drops while average RTT looks acceptable | Track jitter and loss over short windows; tune the jitter buffer to the delay budget and investigate burst congestion. |
If the measurements show the application is waiting on a dependency or a response arrives after the caller’s deadline, set bounded timeouts and handle the failure deliberately. Retries can add traffic during congestion, so they need the safeguards described in the linked failure guide.
Observability Checklist
- Record client location, network type, destination, protocol, and test window alongside measurements.
- Track p50, p95, and p99 latency with sufficient sample volume; keep endpoint classes separate.
- Measure packet loss and jitter over time windows that can expose bursts, not only daily averages.
- Compare idle and loaded RTT to find queueing delay.
- Separate DNS, connection setup, time to first byte, server processing, and full response time.
- Pair throughput and retransmission data with host CPU, memory, socket, and interface counters.
- Use synthetic probes for consistent baselines and real-user monitoring for actual user paths.
- Alert on sustained user impact and error-budget burn, not one noisy probe.
- Keep test traffic bounded and label it so capacity tests do not masquerade as user demand.
Security and Compliance Notes
Only run throughput tests between systems you own or have authorization to test. High-rate tests can consume bandwidth, fill queues, or look like a denial-of-service attempt. Do not expose an iperf3 server broadly; restrict access with firewall rules and shut it down when the test ends. Follow your organization’s data retention and privacy requirements for telemetry, especially when measurements can be associated with a person, device, or location.
Measurement endpoints can reveal network topology, addresses, and usage patterns. Limit access to detailed telemetry, avoid storing unnecessary client identifiers, and redact tokens or query strings from request logs. Use TLS for application probes that carry credentials, and do not put secrets in shell history or shared command output.
Common Pitfalls / Anti-Patterns
- Comparing Mbps with MB/s without converting bits to bytes.
- Treating advertised bandwidth as delivered throughput.
- Calling RTT “latency” without saying whether the measurement is one-way or round-trip.
- Trusting averages while ignoring p95 and p99 behavior.
- Treating a small ping sample as proof that an application path is healthy.
- Testing from a different region, network, or protocol than the affected users.
- Running a speed test that saturates the same link being investigated.
- Assuming every packet loss counter describes loss on the external network; drops can occur at the host or interface.
- Comparing jitter numbers from tools that use different definitions.
- Blaming the network before separating server queues, database work, and response size.
Quick Recap Checklist
- Bandwidth is a link’s capacity; throughput is the rate a test or transfer actually achieves.
- RTT is a round-trip measurement and includes path plus endpoint behavior.
- Jitter is delay variation; packet loss is data that did not arrive successfully.
- TCP loss recovery can lower throughput, while late UDP packets may harm real-time media directly.
- p95 and p99 reveal slow requests that an average can hide.
- Network metrics describe a segment; application response time covers the whole operation.
- Pair probes with realistic application timings and controlled throughput tests.
Interview Questions
Bandwidth is the maximum capacity of a link under stated conditions. Throughput is the rate actually delivered during a measurement. Congestion, protocol overhead, endpoint limits, and loss can keep throughput below bandwidth.
Ping measures an ICMP probe's round trip to one destination. An application request can use a different protocol or route and also spends time in DNS, TLS, server queues, application code, databases, and response transfer. Use request timings or traces to find the slow phase.
TCP retransmits data that does not arrive and adjusts its sending rate when it detects congestion. Retransmissions and recovery add delay and consume capacity. The effect tends to be worse when the round trip is long because feedback takes longer to return.
An average can look healthy while a smaller share of users waits much longer. Percentiles show where the slow tail begins. Interpret them with the sample count, time window, endpoint type, and request volume so a sparse sample does not create a false conclusion.
Expected answer points:
- Throughput measures the data rate delivered during an interval and can include protocol overhead or retransmitted data.
- Goodput counts useful application data delivered per second.
- A flow can have high throughput but lower goodput when loss and protocol overhead consume capacity.
Expected answer points:
- It holds arriving packets briefly so media can be played at a steadier pace despite variable packet delay.
- A larger buffer can absorb more variation but adds playback delay.
- Choose a buffer that fits the application's delay budget and observe late packets and loss alongside jitter.
Expected answer points:
- Queueing may be adding delay as buffers fill; this is commonly called bufferbloat.
- Compare idle and loaded RTT while measuring throughput and queue or interface drops.
- Check whether the rise appears on the access link, a shared uplink, or an endpoint before changing traffic shaping.
Expected answer points:
- Measure DNS, connection, TLS, time to first byte, and total response time separately.
- Use a distributed trace to compare network spans with server queues and downstream dependency work.
- Compare from the affected client path; a stable ping alone does not rule out application-path delay.
Further Reading
- Network Latency, Timeouts, and Failure explains caller deadlines, bounded retries, and failure handling.
- Network Ports and Firewalls covers ports, listening sockets, and network reachability.
- RFC 3393: IP Packet Delay Variation Metric for IP Performance Metrics (IPPM) defines a framework for packet delay variation.
- RFC 2680: A One-Way Packet Loss Metric for IPPM describes a metric for one-way packet loss.
- iperf3 documentation describes the network testing tool used in the examples.
Conclusion
Network performance is not one number. Bandwidth describes capacity, throughput describes delivered rate, RTT describes a round trip, jitter captures delay variation, and loss records packets that did not arrive. Tail percentiles show how often requests fall into the slow end of the distribution.
Measure from the affected path, use realistic traffic, and compare network timings with application traces. A ping is a useful clue; it is not a performance verdict.
Category
Related Posts
CDN Deep Dive: Content Delivery Networks Explained
A comprehensive guide to CDNs — how they work, PoP architecture, anycast routing, cache invalidation strategies, SSL/TLS termination, and real-world performance trade-offs.
Network Observability: Signals for Reliable Services
Track network health across hosts, DNS, paths, proxies, and requests. Learn which signals help diagnose failures without confusing telemetry with service SLOs.
Packet Capture and Network Troubleshooting: Layered Workflow
Use a layered workflow to diagnose DNS, route, TCP, TLS, and HTTP failures with ping, curl, tcpdump, and Wireshark, then collect useful incident evidence.