Day 02 — Linux Networking & Troubleshooting #
Goal #
Build a rigorous, packet-level mental model of the complete network path:
DNS → Routing → ARP → TCP Handshake → TLS → HTTP Request → Socket Lookup → epoll → Application
Learn to debug network bottlenecks, connection drops, and latency anomalies layer-by-layer without guesswork.
1. What Happens When We Run curl? #
Consider running a standard HTTPS request from a Linux terminal:
curl https://example.com
Before a single byte of HTTP response payload returns, the operating system and network stack must execute a strict sequence of lower-level operations:
flowchart TD
Start["curl https://example.com"] --> DNS["1. DNS Name Resolution\n(Hostname → IP Address)"]
DNS --> Routing["2. Kernel Routing Lookup\n(Which interface & next-hop gateway?)"]
Routing --> ARP["3. ARP / Neighbor Discovery\n(Resolve Next-Hop IP → Next-Hop MAC)"]
ARP --> TCP["4. TCP 3-Way Handshake\n(SYN → SYN-ACK → ACK to Port 443)"]
TCP --> TLS["5. TLS 1.3 Cryptographic Handshake\n(Cert Verification & Session Key Exchange)"]
TLS --> HTTP["6. HTTP Request Dispatch\n(GET / HTTP/1.1 or HTTP/2)"]
HTTP --> TTFB["7. Server Processing & TTFB Wait\n(DB, Redis, Locks, Worker Pools)"]
TTFB --> Stream["8. HTTP Payload Streaming & Connection Closure"]
When an error or latency spike occurs, an engineer must isolate which layer in this chain is failing.
2. DNS (Domain Name Resolution) #
The hostname example.com is a human-readable identifier. IP routing works exclusively with binary IP addresses (IPv4 or IPv6).
# Query DNS records with dig:
dig example.com +noall +answer
Example Output:
example.com. 300 IN A 93.184.216.34
The Architectural Scope of DNS #
DNS answers exactly one question:
"What IP address is associated with this hostname?"
DNS does not:
- Guarantee that the target server is reachable.
- Guarantee that the application is running or healthy.
- Select the network route or physical interface to reach the destination.
3. Kernel Routing & Route Lookup #
Once the destination IP (93.184.216.34) is known, the Linux kernel network subsystem checks its Forwarding Information Base (FIB) / Routing Table.
# Query which route Linux will select for a specific destination IP:
ip route get 93.184.216.34
Example Output:
93.184.216.34 via 192.168.1.1 dev eth0 src 192.168.1.10 uid 1000
cache
The output reveals:
- Next-Hop Gateway:
192.168.1.1 - Outgoing Interface:
eth0 - Source IP:
192.168.1.10(the local IP assigned toeth0)
flowchart TD
DestIP["Destination: 93.184.216.34"] --> RouteCheck{"Is destination IP in local subnet (192.168.1.0/24)?"}
RouteCheck -->|"Yes"| Direct["Send directly to destination MAC"]
RouteCheck -->|"No"| Gateway["Select Default Gateway: 192.168.1.1"]
Gateway --> OutDev["Forward out of device: eth0"]
4. Routing vs ARP / Neighbor Resolution #
These two subsystems solve fundamentally different problems at different layers:
| Subsystem | Network Layer | Question Answered | Data Structure | Diagnostic Command |
|---|---|---|---|---|
| Routing | Layer 3 (Network) | "Where should the packet go next?" (Destination IP → Next-Hop IP) | Routing Table (FIB) | ip route showip route get <IP> |
| ARP / Neighbor | Layer 2 (Data Link) | "What physical MAC address owns the next-hop IP?" | ARP Cache / Neighbor Table | ip neigh show |
# Inspect the Linux ARP / Neighbor Table:
ip neigh show
Example Output:
192.168.1.1 dev eth0 lladdr 52:54:00:12:34:56 REACHABLE
5. Intra-Subnet vs Inter-Subnet Communication #
Suppose:
- Client:
192.168.1.10/24 - Internal Server:
192.168.1.20/24 - Default Gateway:
192.168.1.1
Because 192.168.1.20 shares the same subnet mask (/24), the client does not send the packet to the gateway router:
- Client broadcasts an ARP request: "Who has 192.168.1.20? Tell 192.168.1.10".
- The server responds with its MAC address (
aa:bb:cc:dd:ee:20). - Client encapsulates the IP packet directly in an Ethernet frame with destination MAC
aa:bb:cc:dd:ee:20.
6. IP vs MAC Addresses Across Multiple Hops #
A fundamental rule of packet forwarding:
The Destination IP address remains the final target across the entire internet. The Destination MAC address changes at EVERY Layer-2 hop.
sequenceDiagram
participant Client as Client (192.168.1.10)
participant Router as Local Router (192.168.1.1)
participant ISP as ISP Gateway
participant Server as Target Server (93.184.216.34)
Note over Client,Router: Hop 1: Local Ethernet Segment
Client->>Router: Frame: [Src MAC: Client | Dst MAC: Router] [Src IP: 192.168.1.10 | Dst IP: 93.184.216.34]
Note over Router,ISP: Hop 2: WAN / Fiber Link
Router->>ISP: Frame: [Src MAC: Router | Dst MAC: ISP] [Src IP: 192.168.1.10 | Dst IP: 93.184.216.34]
Note over ISP,Server: Hop N: Data Center Segment
ISP->>Server: Frame: [Src MAC: DC Switch | Dst MAC: Server] [Src IP: 192.168.1.10 | Dst IP: 93.184.216.34]
7. HTTPS and Port 443 #
When invoking https://example.com:
- Protocol: HTTPS (HTTP over TLS).
- Target Port:
443(the default IANA assigned port for HTTPS). - Transport: TCP.
Even if you specify a hostname, the transport layer creates a stream connection to a tuple of 4 elements:
(Source IP, Source Port, Destination IP, 443)
8. TCP 3-Way Handshake #
Before any application payload or TLS handshake messages can be transmitted, the client and server must establish reliable synchronization through the TCP 3-Way Handshake:
sequenceDiagram
participant Client as Client (192.168.1.10)
participant Server as Server (93.184.216.34:443)
Note over Client: State: CLOSED
Note over Server: State: LISTEN (in accept queue)
Client->>Server: 1. SYN (seq = 1000)
Note over Client: State: SYN-SENT
Note over Server: State: SYN-RECEIVED (in SYN backlog)
Server->>Client: 2. SYN-ACK (seq = 5000, ack = 1001)
Note over Client: State: ESTABLISHED
Client->>Server: 3. ACK (ack = 5001)
Note over Server: State: ESTABLISHED (moved to accept queue)
The Kernel Socket Queues #
Under high production load, two kernel queues govern handshake behavior:
- SYN Backlog (
tcp_max_syn_backlog): Holds embryonic half-open connections that have receivedSYNbut are waiting for clientACK. - Accept Queue (
somaxconn/ listenbacklog): Holds fully established 3-way handshakes waiting for the userspace application (e.g. NGINX, Node.js) to callaccept().
Production Incident Warning: If your accept queue overflows (inspectable via
ss -lnt), the kernel will either drop incoming connection handshakes or reply withRSTpackets, leading to intermittent connection timeouts even when CPU usage is low.
9. Connection Refused vs Connection Timeout #
Understanding the precise difference between these two failure modes saves hours during production outages:
flowchart TD
Request["Client initiates TCP connection (SYN)"] --> Outcome{"Kernel / Network Response"}
Outcome -->|"TCP RST packet received immediately"| Refused["Connection Refused\n(Host reached, but port has no listening socket or firewall rejected actively)"]
Outcome -->|"No response received within timeout window"| Timeout["Connection Timeout\n(Packets dropped silently by firewall, security group, or routing blackhole)"]
9.1 Connection Refused (ECONNREFUSED) #
- What happens: The client immediately receives a TCP packet with the
RST(Reset) flag set. - Root causes:
- The destination machine was reached, but no process is listening on the target port (e.g. NGINX or PostgreSQL is stopped or crashed).
- A local firewall (iptables / nftables) explicitly rejected the packet with
--reject-with tcp-reset.
9.2 Connection Timeout (ETIMEDOUT) #
- What happens: The client sends
SYNpackets, retransmits with exponential backoff, and eventually gives up after 30 to 120 seconds. - Root causes:
- Cloud Security Group or AWS VPC Network ACL dropped the packet silently (
DROPrather thanREJECT). - Routing loop or missing route table entry in the network.
- Destination host is powered off or disconnected from the network.
- Cloud Security Group or AWS VPC Network ACL dropped the packet silently (
10. The Layered SRE Troubleshooting Checklist #
When investigating an outage or communication failure, always test from lowest layer to highest layer:
1. DNS Layer: Can the hostname be resolved?
Command: dig +trace <domain>
2. IP Layer: Is the remote host reachable via ICMP?
Command: ping -c 3 <IP>
3. Route Layer: Which interface & gateway are selected?
Command: ip route get <IP>
4. Port / TCP Layer: Is the remote port accepting connections?
Command: nc -zv -w 3 <IP> <PORT>
5. TLS Layer: Is the TLS certificate valid and unexpired?
Command: openssl s_client -connect <IP>:443 -servername <domain>
6. HTTP Layer: Does the web server return valid headers?
Command: curl -Iv https://<domain>
7. Application Layer: Does the application complete the database/queue query?
Inspect: logs, thread dumps, wait events
11. Packet Capturing with tcpdump #
tcpdump is the ultimate ground-truth tool for network troubleshooting. It captures packets before firewall inspection and after driver ingress.
# Capture packets on port 443 with numerical IPs and timestamps:
sudo tcpdump -ni any 'tcp port 443' -c 10
Real Output Example:
10:15:02.102341 eth0 In IP 192.168.1.10.54320 > 93.184.216.34.443: Flags [S], seq 12948102, win 64240
10:15:02.124512 eth0 Out IP 93.184.216.34.443 > 192.168.1.10.54320: Flags [S.], seq 8472910, ack 12948103
10:15:02.124599 eth0 In IP 192.168.1.10.54320 > 93.184.216.34.443: Flags [.], ack 8472911
[S]: SYN flag (Client initiating).[S.]: SYN-ACK flags (Server accepting).[.]: Plain ACK flag (Handshake completed).
12. Packet Traversal Through the Linux Kernel #
When a packet arrives from the physical wire into an Ethernet interface:
flowchart TD
Wire["1. Network Cable / Fiber"] --> NIC["2. NIC RX FIFO Ring Buffer"]
NIC --> IRQ["3. Hardware Interrupt (IRQ)"]
IRQ --> NAPI["4. NAPI Polling Routine & SoftIRQ (NET_RX_SOFTIRQ)"]
NAPI --> SKB["5. Allocate struct sk_buff in kernel memory"]
SKB --> Netfilter["6. Netfilter PREROUTING (iptables/nftables raw, mangle, nat)"]
Netfilter --> Routing["7. Routing Decision: Local process or Forwarding?"]
Routing --> LocalIn["8. Netfilter INPUT chain"]
LocalIn --> TCPStack["9. TCP Protocol Stack Processing & Checksum"]
TCPStack --> SockBuf["10. Enqueue into Socket Receive Buffer (sk_receive_queue)"]
SockBuf --> Wakeup["11. Wake up process blocked in read() or epoll_wait()"]
13. Listening Sockets vs Connected Sockets #
A common point of confusion is how a single web server process can handle tens of thousands of concurrent connections.
A listening socket is a special file descriptor configured via listen():
- It does not exchange application data.
- Its sole responsibility is accepting incoming handshake requests.
- When
accept()is called, the kernel creates a brand new file descriptor for that specific client.
Server Process (e.g. NGINX worker)
├── fd 3 → LISTEN 0.0.0.0:443 (Bound listening port)
├── fd 10 → Connected to Client A (192.168.1.50:52134)
├── fd 11 → Connected to Client B (192.168.1.75:61092)
└── fd 12 → Connected to Client C (10.0.4.12:49811)
Key Takeaway: 10,000 active concurrent connections do NOT require 10,000 separate processes. One single worker process can hold 10,000 file descriptors.
14. Connection Acceptance and accept() #
The sequence of creating a connected socket:
// 1. Create a socket
int listen_fd = socket(AF_INET, SOCK_STREAM, 0);
// 2. Bind to port 443
bind(listen_fd, &server_addr, sizeof(server_addr));
// 3. Mark as passive listener with backlog 1024
listen(listen_fd, 1024);
// 4. Accept completed connection from kernel accept queue
int client_fd = accept(listen_fd, &client_addr, &addr_len);
The TCP 3-way handshake is performed autonomously by the Linux kernel before accept() is even called. accept() merely dequeues the established connection descriptor for userspace processing.
15. epoll and Event-Driven High-Concurrency Servers #
How do servers like NGINX, Redis, and Envoy handle 100,000 concurrent sockets without melting CPU cores?
The Old Inefficient Way: `select()` and `poll()` ($O(N)$) #
Historically, servers used select() or poll(), passing an array of 10,000 file descriptors into the kernel on every iteration. The kernel had to scan all 10,000 descriptors one-by-one to find which ones had data ready. This burned immense CPU cycles.
The Modern Linux Way: `epoll` ($O(1)$) #
epoll registers sockets in a Red-Black tree in the kernel. When a packet arrives for any socket:
- The kernel network driver places the ready socket into a ready list (doubly linked list).
epoll_wait()immediately returns only the ready descriptors in $O(1)$ time.
flowchart TD
subgraph Userspace ["NGINX Worker Process"]
Worker["Event Loop"] -->|"epoll_wait()"| KernelWait["Wait for activity"]
KernelWait -->|"Returns only the 3 sockets with data"| ProcessData["Process HTTP requests"]
end
subgraph Kernelspace ["Linux Kernel epoll Subsystem"]
RBTree["Red-Black Tree: 50,000 registered sockets"]
ReadyList["Ready List (Linked List): Socket 42, Socket 89, Socket 104"]
NICEvent["NIC Driver RX Interrupt"] -->|"Pushes ready event"| ReadyList
end
16. Reverse Proxy → Upstream Backend Troubleshooting #
In production, microservices sit behind a reverse proxy (e.g. NGINX, HAProxy, Envoy).
sequenceDiagram
participant Client as Web Browser / Client
participant Proxy as NGINX Reverse Proxy
participant Backend as Go / Python Backend API
participant DB as PostgreSQL Database
Client->>Proxy: GET /api/v1/orders
Note over Proxy: Proxy selects upstream pool
Proxy->>Backend: TCP Connect & Forward HTTP Request
Note over Backend: Processing query & waiting on DB lock
Backend->>DB: SELECT * FROM orders WHERE user_id = ?
Note over DB: Query blocked by long-running transaction!
Note over Proxy: proxy_read_timeout (60s) expires!
Proxy-->>Client: HTTP 504 Gateway Timeout
When diagnosing upstream failures:
- Check NGINX access and error logs: inspect
$upstream_status,$upstream_response_time, and$upstream_connect_time. - Test network connectivity from proxy host to backend:
bash
nc -zv backend-service 8080 - Check backend thread pool and database connection saturation.
17. HTTP 502 Bad Gateway vs HTTP 504 Gateway Timeout #
These two status codes represent completely different failure scenarios:
| Status Code | RFC Name | What It Means | Typical Root Causes |
|---|---|---|---|
| 502 | Bad Gateway | The proxy reached the backend host, but received an invalid or immediate failure response. | 1. Backend crashed / restarted (SIGSEGV, OOM).2. Backend actively refused connection ( TCP RST).3. Backend closed connection abruptly while writing. |
| 504 | Gateway Timeout | The proxy connected successfully, but the backend took longer than the proxy timeout (e.g. proxy_read_timeout 60s) to reply. | 1. Database table/row lock holding queries. 2. Slow unindexed database query. 3. Backend thread pool exhaustion. 4. Downstream third-party API timeout. |
18. Wall-Clock Duration vs CPU Time #
One of the most essential diagnostic distinctions in systems engineering:
Wall-Clock Elapsed Time ≠ CPU Execution Time
Scenario A: High Wall Time, Tiny CPU Time #
Request Wall-Clock Duration: 30,000 ms (30 seconds)
Process CPU Time: 15 ms (0.015 seconds)
- Diagnosis: The process is 99.95% idle. It is not computing anything. It is blocked in sleep waiting on an external dependency (Database lock, remote API, disk I/O, or thread synchronization).
Scenario B: High Wall Time, High CPU Time #
Request Wall-Clock Duration: 30,000 ms (30 seconds)
Process CPU Time: 29,800 ms (29.8 seconds)
- Diagnosis: The process is CPU-bound. Profile with
perf,top -H, or generate a flamegraph. Look for infinite loops, quadratic regex evaluations, or unoptimized serialization.
19. PostgreSQL Wait Events and Diagnostics #
When an application is waiting on PostgreSQL, senior SREs inspect pg_stat_activity:
SELECT
pid,
state,
wait_event_type,
wait_event,
now() - query_start AS query_duration,
now() - xact_start AS xact_duration,
query
FROM pg_stat_activity
WHERE state != 'idle'
ORDER BY query_duration DESC;
Key Wait Event Types #
Lock: The query is blocked waiting for a table or row lock held by another transaction.IO: The database process is blocked waiting for physical disk blocks to be read intoshared_buffers.IPC: Waiting on inter-process communication between PostgreSQL background processes.Client: Waiting for the client application to send the next query over the socket.
20. The "idle in transaction" Hazard #
-- Dangerous state in pg_stat_activity:
state = 'idle in transaction'
What It Means #
The application executed BEGIN; and some SQL queries, but has not yet executed COMMIT or ROLLBACK. The application thread is currently doing something else (or crashed without closing the connection).
Why It Is Catastrophic in Production #
- Holds Locks: Any row or table locks acquired remain locked, queuing up other transactions.
- Prevents VACUUM (Table Bloat): PostgreSQL MVCC cannot vacuum dead tuples created after this transaction started (
xminhorizon is frozen). Table and index sizes explode. - Exhausts Connection Pools: Holding open database connections starves other web request workers.
21. Mitigating Blocked Backends: pg_cancel vs pg_terminate #
-- 1. Polite: Cancel only the currently executing query (connection remains open)
SELECT pg_cancel_backend(4120);
-- 2. Forceful: Terminate the entire connection process (aborts transaction & releases all locks)
SELECT pg_terminate_backend(4120);
During a production incident where a rogue transaction has locked a critical table:
- Capture metadata from
pg_stat_activityandpg_locksfor post-mortem evidence. - Execute
pg_terminate_backend(blocker_pid)to immediately release locks and restore system availability.
22. Database Defensive Timeouts #
Never deploy production applications without explicit timeout safeguards:
# postgresql.conf or session settings:
# Abort any statement that takes longer than 15 seconds:
statement_timeout = 15000
# Abort query if waiting on a lock for more than 5 seconds:
lock_timeout = 5000
# Terminate connections left 'idle in transaction' for more than 30 seconds:
idle_in_transaction_session_timeout = 30000
23. The "Transaction + External API" Anti-Pattern #
One of the most frequent architectural causes of database lock exhaustion:
# ❌ DANGEROUS PRODUCTION ANTI-PATTERN
with db.transaction():
order = create_order()
# Network call to external payment gateway (takes 10-30 seconds during network lag)
payment_response = requests.post("https://payment-api.com/charge", timeout=30)
order.status = "PAID"
order.save()
While waiting 15 seconds for the external HTTP request:
- The database connection is held hostage.
- Any row locks on the
orderstable are held. - If 50 users click checkout simultaneously, 50 database connections freeze, causing pool exhaustion across the entire cluster.
The Correct Architecture #
Keep database transactions minimal, atomic, and strictly local. Use the Outbox Pattern or async worker queues for external API calls:
# ✅ RESILIENT PRODUCTION PATTERN
# Step 1: Record order in fast local transaction
with db.transaction():
order = create_order(status="PENDING")
# Step 2: Make external network call OUTSIDE database transaction
payment_response = requests.post("https://payment-api.com/charge", timeout=10)
# Step 3: Update order in another fast local transaction
with db.transaction():
order.status = "PAID" if payment_response.ok else "FAILED"
order.save()
24. End-to-End Latency Breakdown #
When an API endpoint responds in 15 seconds, decompose the latency waterfall:
gantt
title HTTP Request Latency Waterfall (15,000ms Total)
dateFormat X
axisFormat %s s
section Network
DNS Resolution :0, 10
TCP 3-Way Handshake :10, 30
TLS 1.3 Handshake :30, 60
section Server Wait (TTFB)
Proxy Dispatch :60, 70
App Lock Wait :70, 14800
App DB Query :14800, 14950
section Response
Data Download :14950, 15000
- Network Connection Time: 60 ms.
- TTFB (Time To First Byte): 14,890 ms (99.3% of latency!).
- Data Transfer: 50 ms.
The issue is definitively in the server/application wait state, not client connectivity or DNS.
25. The SRE Incident Triage Mental Model #
When a network or service degradation incident strikes, adhere to this diagnostic flowchart:
Where is the elapsed time being spent?
↓
Is the process actively executing on CPU or waiting in sleep?
↓
What specific resource is it waiting for?
├── Network Packet (Socket read queue)?
├── Disk I/O (Uninterruptible D-state)?
├── Database Mutex / Row Lock?
└── Downstream Microservice API?
↓
Which process or transaction holds that resource?
↓
Why is it holding it (long-running transaction, deadlock, saturated pool)?
↓
Safest immediate mitigation (terminate blocker, restart worker, scale pool)
↓
Permanent root cause fix (add index, remove external call from transaction)
Key Takeaways #
- DNS resolves names; routing decides the next hop. DNS does not route packets.
- The Destination IP stays constant; the MAC address changes at every hop.
- TCP 3-way handshake is handled by the kernel before userspace calls
accept(). Connection refused(RST) means the host was reached but nothing is listening.Timeoutmeans packets are dropped silently by firewalls or routing.- One process can handle 100,000 connections using
epollbecauseepoll_waitscales in $O(1)$ without scanning idle descriptors. 502 Bad Gatewayvs504 Gateway Timeout: 502 means immediate bad response or crash from backend; 504 means the backend took too long to reply.- Wall-clock time is not CPU time. An endpoint taking 60 seconds with 15ms CPU time is blocked on wait events (locks, I/O, or network).
idle in transactionis a major production hazard. It freezes PostgreSQL vacuuming, holds row locks, and bloats tables.- Never execute external HTTP calls inside a database transaction.
- Always ask: "Where is the time being spent, and what is the process waiting for?"