Packet Capture and Network Troubleshooting: Layered Workflow
Use a layered workflow to diagnose DNS, route, TCP, TLS, and HTTP failures with ping, curl, tcpdump, and Wireshark, then collect useful incident evidence.
Start with a symptom and trace a request through DNS, routing, TCP, TLS, and HTTP before reaching for a packet capture. This guide shows how to choose narrow tcpdump filters, read captures from both client and server viewpoints, and compare packet evidence with host and application logs. It also covers capture privacy and escalation, so you can collect useful evidence without asking a single trace to prove more than it can.
Packet Capture and Network Troubleshooting: Layered Workflow
Introduction
A packet capture records traffic visible at a particular interface and point in time. It can reveal protocol behavior that application logs do not show, but it only describes what that capture point observed; it may not identify where a missing packet was dropped.
This guide walks through a repeatable troubleshooting sequence, from choosing a narrow filter to reading the resulting trace. It also covers safe capture limits and how to correlate packet evidence with host, application, and network telemetry.
When to Use
Use this workflow when a service is unreachable, intermittently slow, or behaves differently from different machines or networks. It is especially useful when application logs show a timeout but do not reveal whether the request reached the host.
It also helps when a deployment changes DNS, routes, firewall rules, a load balancer, or a listener. A short client-side capture and a matching server-side capture can show where a TCP handshake stops.
When NOT to Use
Do not begin with a broad packet capture if a simpler check already identifies the fault. A failed DNS lookup or a process that is not listening may answer the immediate question without collecting traffic.
Avoid capturing traffic on systems or networks unless you are authorized to inspect it. Do not use a packet capture to collect unrelated user activity, credentials, or application payloads. For high-impact incidents, follow the organization’s incident process and approved evidence-retention rules.
A repeatable troubleshooting order
Record the symptom and check name resolution
1. Reproduce and record the symptom
Write down the exact command or request, the source machine, destination name and port, time with timezone, and whether the failure is consistent. Compare one known-good client if available. Keep the test narrow; repeated retries can add load and blur the timeline.
Use the same hostname and protocol as the failing application. A successful request to an IP address does not prove the hostname path works, because it bypasses DNS and may change TLS certificate selection.
2. Check DNS before the socket
Ask the resolver what address it returns:
dig +time=2 +tries=1 api.example.test A
For IPv6, query AAAA as well. Compare the answer with a known-good resolver only when that comparison is permitted and useful. Note the resolver, answer, TTL, and query time. A stale or split-horizon answer can send two clients to different endpoints.
Confirm the network and classify the failure
3. Verify the route and the listener
On Linux, ask the kernel which route it would use for the destination address:
ip route get 203.0.113.10
On the server, check whether a process is listening on the expected port:
sudo ss -lntp
Look for the right local address and port. A service bound only to 127.0.0.1:443, for example, will not accept connections addressed to the host’s external interface. A listener check on the server does not prove a firewall or load balancer allows traffic to reach it.
4. Classify what the client observes
The failure type narrows the next question:
| Observation | What it suggests | What to check next |
|---|---|---|
| Name lookup fails or returns an unexpected address | Resolver, record, or search-path issue | Query type, resolver, TTL, split-horizon DNS |
| Connection is refused quickly | A host or network device rejected the TCP attempt | Listener address and port, firewall reject rule, service health |
| Connection times out | No timely response reached the client | Route, filtering, packet loss, return path, overloaded endpoint |
| TCP connects, then TLS fails | Transport reached a peer, but the TLS exchange or validation failed | Certificate name/chain/expiry, SNI, protocol policy, TLS alert |
| TLS succeeds and an HTTP status returns | The HTTP service responded | Status, headers, application logs, proxy and routing rules |
These are clues, not verdicts. A reset can come from a firewall or proxy, and a timeout can be caused by a server that never accepted the request. Record the exact error and timestamp before changing anything.
A bounded request can help distinguish connection setup from later work:
curl -sv --connect-timeout 3 --max-time 10 \
https://api.example.test/ -o /dev/null
Verbose output can contain sensitive headers, certificate details, and internal hostnames. Review it before sharing. Do not add -k as a routine fix; it disables certificate verification and can hide the problem you are trying to diagnose.
5. Use reachability tools with care
ping can show whether ICMP echo replies return, but many networks filter ICMP while allowing application traffic:
ping -c 4 203.0.113.10
A route trace can show where probe responses stop, but routers often rate-limit or suppress them. For example, on Linux:
traceroute -n -m 12 -w 1 203.0.113.10
Do not interpret an asterisk as proof that a hop is broken. Treat ping and traceroute as supporting evidence, then verify the application port with the actual client request.
Capture and interpret evidence
6. Capture only what answers the next question
First identify the correct interface and confirm capture authorization. On Linux, ip route get can help identify the outgoing interface. Then capture a limited number of packets with a narrow capture filter:
sudo tcpdump -i eth0 -nn -s 0 -c 200 \
-w /tmp/api-check.pcap \
'host 203.0.113.10 and tcp port 443'
The address is from a documentation-only range; replace it with the approved destination. -c 200 stops after 200 matching packets. -nn avoids name and service lookups, and -w writes packet data to a file. Choose an interface and count that fit the incident; eth0 may not exist on your host. The full-packet setting -s 0 can capture application data, so use it only when that detail is necessary and approved.
Capture filters decide what enters the file. In Wireshark, a capture filter such as host 203.0.113.10 and tcp port 443 follows libpcap filter syntax. Display filters run after capture and use a different language. For example, tcp.flags.syn == 1 && tcp.flags.ack == 0 displays initial TCP SYN packets in an opened capture.
To inspect a saved file in the terminal without changing it:
tcpdump -nn -r /tmp/api-check.pcap
Wireshark is useful for following a TCP stream, checking retransmissions, and inspecting TLS handshake metadata. The official Wireshark User’s Guide documents capture and display filters.
7. Interpret the capture from its vantage point
Ask what the capture point could observe. A client capture can show that the client sent SYN packets; it cannot prove those packets reached the server. A server capture that sees the SYN but sends no SYN-ACK points toward the server host, its local firewall, or listener path. A server-side SYN-ACK that never appears in the client capture shifts attention to the return path or an intermediate filter.
Other caveats matter:
- NIC offload can make locally captured packet sizes or checksums look unusual. Check the capture location and host offload settings before calling a checksum warning a network fault.
- TLS encrypts application payloads. A capture can still show addresses, ports, timing, packet sizes, TCP behavior, and parts of the TLS handshake, but usually not HTTP paths or bodies.
- Packet drops can happen in the network, kernel, capture process, or interface queue. Record tcpdump’s final kernel-drop count when available; a capture with drops is incomplete evidence.
- A capture filter can exclude the packet you needed. If DNS is part of the question, a filter limited to TCP port 443 will omit DNS traffic.
Production Failure Scenarios
One client times out while another succeeds
The failing client resolves the service to 203.0.113.10; the working client resolves a different address. curl from the failing client times out before TLS, while ip route get chooses the expected interface. A narrow capture shows repeated SYNs and no SYN-ACK at the client. The next useful check is whether the destination host received those SYNs, then whether the two DNS answers are intentional. The client capture alone cannot name the dropping device.
Service listens only on loopback
A process appears in ss -lntp, but its socket is bound to 127.0.0.1:8080. Local health checks pass while remote clients time out or receive a refusal. The listener exists, but not on the address remote traffic uses. Confirm the intended bind address and deployment configuration before changing firewall rules.
TLS breaks after a certificate rotation
TCP completes and the client receives a TLS alert or reports a certificate-name error. The packet trace shows that the peer was reachable; it does not make the certificate valid. Compare the requested hostname, SNI, certificate chain, and deployment time. Avoid suppressing verification to make the symptom disappear.
Trade-Off Table
| Tool or evidence | Good for | Limitation |
|---|---|---|
dig |
Resolver answers and record details | Does not prove the application endpoint is reachable |
ip route get |
Local route and outgoing interface choice on Linux | Does not show every network hop or policy device |
ss |
Local socket state and listeners | Does not prove an external path is open |
ping / traceroute |
Basic path clues and loss patterns | ICMP can be filtered or rate-limited |
curl |
Reproducing DNS-through-HTTP behavior with deadlines | Verbose output may expose sensitive metadata |
tcpdump / Wireshark |
Packet timing, handshakes, resets, retransmissions | Captures are vantage-point limited and may contain sensitive data |
Observability Checklist
Before escalation, collect:
- The exact hostname, resolved addresses, destination port, source host, and test command.
- A timestamp with timezone and whether the issue reproduces consistently.
- DNS output, the local route result, and listener output when relevant.
- The exact client error: lookup failure, refusal, timeout, TLS validation or alert, or HTTP status.
- A narrow packet capture from an approved point, with the capture interface, filter, duration or packet limit, and reported packet drops.
- Correlated server, proxy, load balancer, and application logs for the same time window.
- What changed recently, such as a deployment, DNS update, certificate rotation, route, or policy change.
For a useful handoff, state the last step that succeeded and the first step that failed. Include sanitized command output and a short explanation of what the capture can and cannot establish.
Security and Compliance Notes
A PCAP may contain credentials, session cookies, personal data, internal addresses, DNS names, and unencrypted application content. Treat it like production data. Get approval, filter by host and port, cap the packet count or duration, store it in an access-controlled location, and apply the approved retention period. Restrict file permissions before capture when practical:
umask 077
Do not attach raw captures to public tickets or send them through ordinary chat. Share a minimized, redacted excerpt when that answers the question. If full payload capture is unnecessary, use a smaller snap length or collect only headers, then document what detail was omitted. TLS reduces payload visibility but does not make a capture anonymous.
Common Pitfalls / Anti-Patterns
- Declaring the network healthy because ping works, or broken because ping fails.
- Treating
connection refusedandconnection timed outas interchangeable. - Capturing on the wrong interface and concluding that no traffic exists.
- Using a broad filter or unlimited capture by default, which increases privacy exposure and makes analysis harder.
- Assuming one endpoint’s trace proves what happened at another endpoint.
- Treating checksum warnings on a host capture as proof of corrupt packets without considering NIC offload.
- Sharing a raw PCAP without checking for credentials, cookies, or payload data.
- Disabling TLS verification to make a test pass.
A related explanation of the transport behavior is in TCP, IP, and UDP. For timeout budgets and the risks of retries, see Network Latency, Timeouts, and Failure.
Quick Recap Checklist
- Reproduce the failure and record the exact time, source, target, and error.
- Check DNS, then the local route and expected listener.
- Separate lookup, TCP, TLS, and HTTP outcomes.
- Use
pingand route tracing as clues, not verdicts. - Capture a bounded, authorized sample with the narrowest useful filter.
- Compare client and server observations; note capture drops and vantage point.
- Sanitize evidence and include the last successful step in the escalation.
Interview Questions
It can show packets observed at the client, such as repeated SYN transmissions with no SYN-ACK arriving there. It cannot prove whether the SYN reached the server or identify which intermediate device dropped it. A second capture at the server or evidence from network devices can narrow the boundary.
A refusal is a prompt rejection, often a TCP reset from the host or a device. A timeout means the client did not receive a usable response before its deadline. Neither message alone identifies the exact cause, so check the listener and compare packet observations at the endpoints.
Ping tests ICMP echo, while HTTPS needs DNS resolution, a route, TCP port 443, TLS negotiation, and an HTTP response. ICMP may use different firewall rules, and a working route does not prove that the HTTPS listener or certificate is correct.
A capture can expose payloads, credentials, cookies, personal data, DNS names, and internal addresses. Get authorization, narrow the filter and capture window, restrict access, redact before sharing, and delete the artifact under the applicable retention policy.
Expected answer points:
- A capture filter limits which packets are recorded, using libpcap syntax.
- A display filter selects packets for viewing after they have been captured and uses Wireshark's filter language.
- A narrow capture filter can permanently omit traffic needed for a later question, so choose it based on the investigation goal.
Expected answer points:
- The client sent SYNs at its capture point, but the server capture did not observe them.
- The loss or filtering boundary lies somewhere between those capture points, but the captures alone do not identify the specific device.
- Compare timestamps, interfaces, routes, and intermediate firewall or flow logs to narrow the path.
Expected answer points:
- It shows the sender retransmitted data according to its TCP recovery behavior.
- It does not by itself prove where or why the earlier segment or acknowledgment was lost; delay, reordering, or capture vantage can affect interpretation.
- Compare both directions, endpoint counters, and timestamps before assigning a cause.
Expected answer points:
- DNS usually uses UDP or TCP port 53, so a capture filter limited to TCP port 443 will exclude those packets.
- DNS may also have been answered from a local or intermediate cache before the captured connection began.
- Include the relevant resolver traffic or capture a separate DNS query when name resolution is part of the question.
Further Reading
- Wireshark User’s Guide
- TCP, IP, and UDP: Understanding Internet Transport Protocols
- Network Latency, Timeouts, and Failure
Conclusion
Troubleshoot in order: reproduce the symptom, check name resolution, inspect the route and listener, then identify whether the failure is at TCP, TLS, or HTTP. Capture only the traffic needed, from a known vantage point, and treat packet traces as sensitive evidence. When escalating, pair the capture with timestamps, the exact test, and observations from both sides of the connection whenever possible.
Category
Related Posts
Network Encapsulation: Follow a Packet Across the Stack
Follow an HTTPS request from a browser through transport, IP, and link layers, and learn how headers, MTU, routers, and packet captures fit together.
Network Observability: Signals for Reliable Services
Track network health across hosts, DNS, paths, proxies, and requests. Learn which signals help diagnose failures without confusing telemetry with service SLOs.
Network Performance: Latency, Throughput, Jitter & Loss
Understand bandwidth, throughput, latency, RTT, jitter, and packet loss with practical measurements, tail percentiles, and production diagnostic guidance.