Linux SystemsDay 03advanced

Linux Deep Dive: Memory Internals, cgroups, Namespaces & I/O Troubleshooting

Build an architectural mental model of Linux memory (VSZ, RSS, PSS), Page Cache dynamics, Swap behavior, Pressure Stall Information (PSI), resource isolation with cgroups v2 and namespaces, file descriptor leaks, storage I/O with iostat, and 15 production SRE incident scenarios.

55 min read
•Foundational Engineering Reference

Day 03 — Linux Deep Dive: Memory, cgroups, Namespaces & I/O Troubleshooting #

Goal #

In this deep-dive lesson, we dissect the core subsystems of the Linux kernel: Virtual Memory & Physical RAM management, resource isolation primitives (Namespaces & cgroups v2), File Descriptors, and Storage Block I/O triage.

The primary objective is not simply to memorize commands, but to construct a rigorous mental model of how Linux behaves under severe saturation, lock contention, and resource starvation. You will learn to isolate real bottlenecks across kernel layers using empirical telemetry and never confuse an isolated metric with the true root cause.


1. Dissecting Memory Metrics: VSZ vs. RSS vs. PSS #

A pervasive mistake in infrastructure engineering is looking at VSZ or even RSS to deduce how much physical RAM an application is actively consuming.

Architecture / Lifecycle Diagram
flowchart TB
    subgraph VirtualMemory ["Virtual Address Space (VSZ)"]
        direction TB
        VM1["Executable Code & Shared Dynamic Libraries (.so)"]
        VM2["Memory-Mapped Files (mmap regions)"]
        VM3["Allocated Virtual Heap (not yet committed to RAM)"]
        VM4["Thread Stacks & Anonymous Committed Pages"]
    end

    subgraph PhysicalRAM ["Physical RAM (Resident Set Size - RSS)"]
        PR1["Active Private Anonymous Pages"]
        PR2["Shared Library Pages in RAM (Multi-counted)"]
    end

    subgraph FairShare ["Proportional Set Size (PSS)"]
        PSS1["100% Private Anonymous Pages"]
        PSS2["Shared Pages ÷ Number of Sharing Processes"]
    end

    VM4 -.-> PR1
    PR1 -.-> PSS1
    PR2 -.-> PSS2

1. VSZ (Virtual Memory Size) #

The total virtual address space (up to hundreds of gigabytes in 64-bit systems) that a process has requested from the kernel via mmap() and brk().

  • Includes executable code, shared .so libraries, memory-mapped files, and reserved but uncommitted address ranges.

Key Rule: VSZ does not represent physical RAM consumed. A JVM or Go runtime allocating a 20 GB heap with 500 MB active data has VSZ = 20 GB and RSS = 500 MB. Physical memory usage is only 500 MB.

2. RSS (Resident Set Size) #

The subset of virtual pages that are currently resident in physical RAM.

  • The RSS Trap: For processes that share code, glibc, or memory-mapped regions (MAP_SHARED), a single shared page is counted in the RSS of every single process referencing it. Summing the RSS of 30 Nginx workers will vastly exceed the physical memory of the entire host.

3. PSS (Proportional Set Size) #

The single most accurate metric for understanding a process's true memory footprint. It accounts for shared pages by dividing their size by the number of processes sharing them. If four workers share a 4 KB page, each process receives 1 KB in its PSS.

Production Inspection Commands: #

bash
# High-level memory status of a target process
cat /proc/<pid>/status | grep -iE 'vmsize|vmrss|data|stk'

# Detailed, proportional accounting with PSS
cat /proc/<pid>/smaps_rollup

Sample output:

text
Rss:              452100 kB
Pss:              210400 kB
Pss_Dirty:        180200 kB
Shared_Clean:     241700 kB
Private_Dirty:    180200 kB
Referenced:       410200 kB
Anonymous:        180200 kB

2. Demystifying `free` vs. `available` Memory #

One of the most frequent false-positive alerts stems from engineers observing low free memory and assuming the server is about to crash.

Architecture / Lifecycle Diagram
flowchart LR
    subgraph TotalRAM ["Physical RAM Capacity"]
        direction LR
        Used["Active Application Memory (Used)"]
        Cache["Page Cache & Dentries/Inodes (Reclaimable)"]
        Free["Unallocated Idle Memory (Free)"]
    end

    Cache -.->|"Instant Reclaim"| Available["Available for Applications"]
    Free -.-> Available

The Linux kernel considers completely unallocated RAM wasted capital. Consequently, it dynamically allocates idle memory to the Page Cache and directory buffers to accelerate file reads and writes.

bash
free -h

Sample output:

text
               total        used        free      shared  buff/cache   available
Mem:            31Gi        12Gi       650Mi       1.2Gi        18Gi        18Gi
Swap:           15Gi       2.0Gi        13Gi
  • free (650 MiB): Bytes that currently contain zero data.
  • buff/cache (18 GiB): Disk blocks, metadata, and files held in memory.
  • available (18 GiB): The kernel's calculated estimate of memory that can be reclaimed instantaneously without forcing the host into swap or dropping dirty buffers.

Engineering Reality: Low free memory alongside high available memory indicates an operating system running with optimal caching efficiency.


3. Dissecting Swap Behavior: Is Swap Usage an Emergency? #

Observing Swap Used = 8 GB in Grafana does not inherently signal an active outage.

Under default kernel tuning (vm.swappiness = 60), the kernel gradually swaps out dormant, cold anonymous pages (e.g., initialization routines or idle background daemon memory) to disk to expand the Page Cache for frequently accessed database files.

Detecting Active, Destructive Thrashing #

The critical signal is never the absolute volume of swap space consumed; it is the rate of swap I/O operations per second:

bash
# Monitor swap activity every second
vmstat 1
text
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st
 2  1 8388608 524288  10240 1048576 4200 6800  1200  2400 1200 4500 18 12  0 70  0
  • si (Swap In): Kilobytes read from swap partition into RAM per second.
  • so (Swap Out): Kilobytes written from RAM into swap partition per second.

Continuous, high values in both si and so indicate Swap Thrashing: CPU threads are perpetually waiting on disk I/O to swap pages in and out, degrading service latency exponentially.


4. Modern Memory Pressure Telemetry: PSI (Pressure Stall Information) #

Introduced in Linux 4.20, PSI revolutionized systems observability by quantifying the exact percentage of time threads are stalled waiting for CPU, memory, or I/O resources.

bash
cat /proc/pressure/memory

Sample output:

text
some avg10=42.50 avg60=31.20 avg300=18.00 total=14205100
full avg10=18.40 avg60=11.10 avg300=05.20 total=4810200
  • some: Percentage of wall-clock time in which at least one thread was stalled on memory allocation, direct reclaim, or page faults.
  • full: Percentage of wall-clock time in which all non-idle runnable threads were completely blocked waiting on memory (total system standstill).
Architecture / Lifecycle Diagram
flowchart TD
    LowAvail{"Is Available Memory low?"}
    LowAvail -->|"No"| Healthy["System memory is healthy"]
    LowAvail -->|"Yes"| ActiveSwap{"Are si / so high in vmstat?"}
    ActiveSwap -->|"No"| Dormant["Memory is occupied by Page Cache; no thrashing"]
    ActiveSwap -->|"Yes"| CheckPSI{"Is PSI full > 0%?"}
    CheckPSI -->|"Yes"| Emergency["Definite Memory Pressure Outage; immediate triage required"]

5. The OOM Killer: Global Host OOM vs. cgroup OOM #

When physical RAM and swap are fully exhausted and the kernel can no longer reclaim clean page cache, the Out-Of-Memory (OOM) Killer is invoked to prevent a total kernel panic.

Architecture / Lifecycle Diagram
flowchart TD
    OOMEvent["Memory Exhaustion Event"] --> Scope{"What is the scope of starvation?"}
    
    Scope -->|"Host-wide RAM + Swap depleted"| GlobalOOM["Global Host OOM Killer"]
    GlobalOOM --> RankProc["Evaluate oom_score = (RAM% * 10) + oom_score_adj"]
    RankProc --> TerminateGlobal["Send SIGKILL to process with highest score"]

    Scope -->|"Container cgroup hits memory.max"| CgroupOOM["cgroup Local OOM Killer"]
    CgroupOOM --> TerminateLocal["Send SIGKILL strictly to process inside that cgroup"]
    TerminateLocal --> NodeSurvives["Rest of host and adjacent containers continue normally"]

1. Global Host OOM Triage #

bash
# Review kernel ring buffer for host-wide OOM events
dmesg -T | grep -iE 'oom|out of memory|killed process'

# Journal inspection
journalctl -k -S "2 hours ago" | grep -iE 'oom-killer'

Signature output:

text
[Mon Oct 05 14:22:01 2026] Out of memory: Killed process 14205 (mysqld) total-vm:18524020kB, anon-rss:14205100kB, file-rss:1024kB, shmem-rss:0kB, oom_score_adj:0

2. Container cgroup OOM #

In modern container runtimes, containers have strict limits (resources.limits.memory). If an application exceeds this ceiling, the kernel kills the container process (Exit Code 137) even if the underlying node has 64 GB of free RAM.


6. Architectural Distinction: Namespaces vs. cgroups #

Mastering container internals requires decoupling visibility from resource accounting:

CharacteristicNamespacescgroups (Control Groups)
Core FunctionIsolation & VisibilityResource Metering & Throttling
Guiding Question"What system entities can this process see?""How much hardware capacity can this process consume?"
SubsystemsPID, Mount, Network, IPC, UTS, User, CgroupMemory, CPU, I/O, PIDs, RDMA, Network Priority
Practical ExampleA process inside a container sees itself as PID 1.A process cannot exceed 2 CPU cores or 2 GiB of memory.

7. cgroup v2 Tree Hierarchy & Effective Limits #

In cgroup v2, all controllers are unified into a single hierarchical tree:

Architecture / Lifecycle Diagram
flowchart TD
    Root["Root cgroup (Host)"]
    Root --> Kubepods["kubepods.slice"]
    Kubepods --> Pod["Pod Sandbox (memory.max = 2 GiB)"]
    Pod --> ContainerA["Container A (memory.max = 1 GiB)\nEffective Limit: 1 GiB"]
    Pod --> ContainerB["Container B (memory.max = 4 GiB)\nEffective Limit: 2 GiB (Constrained by Parent Pod!)"]

The Effective Limit Rule: A child cgroup cannot exceed the resource ceiling imposed by any of its ancestors. If the enclosing Pod has memory.max = 2 GiB, Container B will be terminated upon exceeding 2 GiB, despite its local configuration of 4 GiB.


8. Historical Accounting with `memory.events` in cgroup v2 #

A common engineering puzzle is diagnosing a pod that restarted with Exit Code 137, but whose memory consumption right now is only 300 MiB.

bash
cat /sys/fs/cgroup/<cgroup_path>/memory.events

Sample output:

text
low 0
high 0
max 24
oom 4
oom_kill 2

These numbers are monotonic lifetime counters:

  • max: Number of times memory usage attempted to cross memory.max, triggering synchronous reclaim.
  • oom: Number of distinct times memory exhaustion occurred.
  • oom_kill: Exact count of processes within this cgroup terminated by SIGKILL due to OOM.

Even if memory.current is low at inspection time, a non-zero oom_kill counter proves beyond doubt that the container previously died due to exceeding its memory ceiling.


9. Memory Leak Signatures in Production #

Memory allocators (such as glibc ptmalloc, jemalloc, and tcmalloc) as well as garbage-collected runtimes do not immediately return released memory pages to the operating system via munmap or brk to minimize syscall overhead.

Architecture / Lifecycle Diagram
flowchart LR
    subgraph Normal ["Healthy Baseline Reset"]
        W1["Workload 1"] --> Peak1["Peak: 800MB"] --> Base1["Baseline: 500MB"]
    end

    subgraph LeakPattern ["Memory Leak (Stair-step Baseline)"]
        W2["Workload 1"] --> P1["800MB"] --> B1["Stays at 800MB"]
        W3["Workload 2"] --> P2["1.1GB"] --> B2["Stays at 1.1GB"]
        W4["Workload 3"] --> P3["1.4GB"] --> B3["Stays at 1.4GB"]
    end

The telltale signature of an authentic memory leak is a stair-stepping baseline: following the completion of successive workloads and quiet idle periods, the baseline memory never drops back down. True leaks must be isolated via runtime heap profiling tools (pprof, jemalloc prof, async-profiler).


10. File Descriptors & System Limits #

The UNIX tenet "Everything is a file" means a file descriptor (FD) represents far more than physical files on disk:

text
Types of File Descriptors:
1. Regular disk files
2. TCP and UDP network sockets
3. Unix Domain Sockets (.sock)
4. Anonymous Pipes and FIFOs
5. Standard Streams (stdin, stdout, stderr)
6. Asynchronous Event Channels (epoll, eventfd, timerfd, pidfd)

Production Troubleshooting Commands: #

bash
# Total open FDs for a specific process
ls /proc/<pid>/fd | wc -l

# Full inventory of all open descriptors by type
lsof -p <pid>

# Check kernel limits enforced on the process
cat /proc/<pid>/limits | grep "Max open files"

# View current shell ulimit
ulimit -n

If lsof reveals thousands of descriptors of type sock in state CLOSE_WAIT or unestablished, the issue is not a shortage of disk file slots; it is a Socket Leak caused by unclosed connections or missing timeouts.


11. Decoding I/O Wait (`iowait`) #

When top displays high CPU wait:

text
%Cpu(s): 12.0 us,  4.0 sy,  0.0 ni, 44.0 id, 40.0 wa,  0.0 hi,  0.0 si,  0.0 st

The wa column (iowait) indicates that CPU cores were idle while runnable tasks were queued in an uninterruptible sleep state (TASK_UNINTERRUPTIBLE or D state) waiting on outstanding disk or network storage I/O.

Crucial Rule: High iowait does not automatically indicate physical hard disk failure.
The delay could stem from remote NFS latency, cloud block store throttling (AWS EBS burst balance exhaustion), hypervisor lockups, or severe filesystem lock contention.


12. Storage Performance Analysis with `iostat` #

The standard utility for diagnosing block layer bottlenecks:

bash
iostat -xz 1

Sample output:

text
Device    r/s    w/s    rkB/s    wkB/s    await    aqu-sz    %util
sda    1100.0   50.0   4400.0   200.0    38.50      4.2    99.4%

Key Metrics Decoded: #

  • %util: Percentage of elapsed time during which the device was servicing requests. Above 80% indicates saturation.
  • await: Average time (in milliseconds) for I/O requests to be serviced, encompassing both queue wait time and hardware execution. Modern NVMe drives should maintain < 2ms; values > 30ms indicate a bottleneck.
  • r/s & w/s (IOPS): Read and write operations dispatched to the device per second.
  • rkB/s & wkB/s (Throughput): Data transferred in kilobytes per second.
  • aqu-sz: Average queue length of requests awaiting service.

Throughput vs. IOPS Disconnect: #

text
1100 IOPS × 4 KB (Random Access) ≈ 4.4 MB/s ──> Device is 100% utilized (%util = 99.4%)
10 IOPS   × 40 MB (Sequential Access) ≈ 400 MB/s ──> High throughput, but device queue is empty

A database performing thousands of random index seeks can completely saturate a disk at just 4 MB/s of bandwidth.


13. Anomalous Storage Profiles: High Latency with Low IOPS #

Consider observing the following telemetry:

text
Device    r/s    w/s    rkB/s    wkB/s    await     %util
sdb       8.0    4.0     32.0     16.0   250.00ms   98.0%

The system is issuing merely 12 requests per second (IOPS), yet average latency (await) is an astonishing 250 milliseconds!

This profile immediately points to an external or virtualization issue:

  1. Hypervisor Noisy Neighbor: Another virtual machine on the same physical host is hogging the storage bus.
  2. Network Storage Latency: Packet drops or switch congestion along the SAN, Ceph, or NFS fabric.
  3. Storage Controller Degradation: The hardware RAID controller cache battery failed, forcing writes into write-through mode.

14. Virtualized & Cloud Storage Stacks #

When an application runs inside a virtual machine or container, your mental model must span every layer of the virtualization hierarchy:

Architecture / Lifecycle Diagram
flowchart TD
    App["1. Application Code & Syscalls"] --> FS["2. Guest OS Filesystem (ext4 / xfs)"]
    FS --> VDisk["3. Virtual Disk Driver (virtio-blk / virtio-scsi)"]
    VDisk --> Hypervisor["4. Hypervisor Storage Stack (KVM / QEMU / ESXi)"]
    Hypervisor --> StorageNet["5. Host Storage Fabric (10GbE / SAN / NVMe-oF)"]
    StorageNet --> SAN["6. Storage Controller & Physical Media (Ceph / EBS / SAN Array)"]

If the root bottleneck resides in hypervisor queue saturation or cloud EBS credit depletion, tuning filesystem parameters inside the virtual machine will yield zero improvement.


15. The Core SRE Principle: Metric ≠ Root Cause #

In production incidents, jumping from a single metric to a root-cause conclusion is the hallmark of amateur debugging:

text
❌ Trap: "CPU is only 20%, so compute and application code are healthy!"
✔️ Reality: Threads are stalled waiting on database row locks or external API sockets.

❌ Trap: "Swap Used is 10 GB, the server is out of memory!"
✔️ Reality: Dormant pages were swapped out weeks ago; active si/so is currently zero.

❌ Trap: "iowait is 40%, the hard drive is physically corrupted!"
✔️ Reality: A remote NFS client lost connectivity, blocking filesystem flush operations.

❌ Trap: "Container died with Exit 137, the physical node ran out of RAM!"
✔️ Reality: Only the container's isolated cgroup memory.max was exceeded.

16. Fifteen Production SRE Incident Scenarios #

Review each incident scenario below to build empirical diagnostic intuition.


Scenario 1 — PostgreSQL Row Lock Contention #

Situation: #

A database query experiences severe latency:

sql
UPDATE payments SET status = 'completed' WHERE id = 9812;

The query took 14.8 seconds of wall-clock time. However, CPU time registered for this query was merely 5ms. PostgreSQL telemetry reports:

text
wait_event_type = Lock
wait_event      = transactionid

Diagnostic Question: #

Where is the bottleneck, and how do you locate the blocking transaction?

SRE Lesson & Mental Model: #

Low CPU + high wall-clock latency + Lock/transactionid indicates the query spent 99.9% of its lifecycle dormant, waiting for another transaction to release its lock.

sql
-- 1. Identify active sessions and wait events
SELECT pid, usename, state, wait_event_type, wait_event, query
FROM pg_stat_activity WHERE wait_event_type = 'Lock';

-- 2. Identify the blocking process ID directly
SELECT pg_blocking_pids(<pid>);

-- 3. Inspect active locks held by the blocker
SELECT * FROM pg_locks WHERE pid = <blocker_pid>;

Scenario 2 — Accumulation of `idle in transaction` Sessions #

Situation: #

PostgreSQL reports dozens of connections in state idle in transaction. Database CPU is < 10%, but the application connection pool is fully depleted, blocking new user requests.

Diagnostic Question: #

Why is an idle connection dangerous to system health?

SRE Lesson & Mental Model: #

idle in transaction means an application issued BEGIN but never finalized the transaction with COMMIT or ROLLBACK.

  1. It retains exclusive row and table locks indefinitely.
  2. It pins the database transaction horizon, preventing VACUUM from cleaning up dead tuples, leading to catastrophic table bloat.
  3. It exhausts database and proxy connection pool slots.

Mitigation & Safeguards:

sql
-- Enforce automated session termination in postgresql.conf
ALTER SYSTEM SET idle_in_transaction_session_timeout = '60s';
SELECT pg_reload_conf();

-- Emergency termination during an active incident
SELECT pg_terminate_backend(<pid>);

Scenario 3 — Holding Database Transactions Across Slow External APIs #

Situation: #

Backend application code contains the following sequence:

text
BEGIN;
UPDATE accounts SET balance = balance - 100 WHERE id = 42;
-- Synchronous outbound HTTP call to payment gateway (experiencing 15s latency)
HTTP POST https://payment-provider/charge
COMMIT;

Diagnostic Question: #

What failure mode does this introduce to the database architecture?

SRE Lesson & Mental Model: #

The database transaction remains open for the entire 15 seconds. At 10 requests per second, within 2 seconds all connection pool slots are occupied, locking accounts across the user base and crashing the primary API gateway.

Architectural Law: Never keep database transactions open across network calls. Use asynchronous processing, state machines with pending status, or the Transactional Outbox Pattern.


Scenario 4 — Latency Breakdown & Layer Isolation #

Situation: #

Network telemetry for an endpoint reports:

text
DNS Resolution : 10 ms
TCP Handshake  : 20 ms
TLS Handshake  : 30 ms
TTFB           : 14,800 ms (14.8 seconds)
Content Download: 10 ms

Diagnostic Question: #

Which layer owns the performance anomaly?

SRE Lesson & Mental Model: #

Connection establishment (DNS, TCP, TLS) took only 60ms. Almost the entire duration was spent waiting for TTFB (Time To First Byte). The delay is 100% internal to the backend application, database query execution, or cache lock contention—not the network layer.


Scenario 5 — Container OOM on a Node with Plentiful Free Memory #

Situation: #

A host node possesses 32 GB RAM with 10 GB unallocated. A container deployed with memory.max = 2 GB reaches 2 GB consumption and abruptly crashes with Exit Code 137 and reason OOMKilled.

Diagnostic Question: #

Must the physical host run out of memory for a container to suffer an OOM kill?

SRE Lesson & Mental Model: #

No. Control groups (cgroups) strictly isolate resource allowances. When a cgroup crosses memory.max, the kernel executes a localized OOM kill strictly within that cgroup, leaving the host and adjacent containers untouched.


Scenario 6 — OOMKilled Status with Low Current Memory #

Situation: #

A Kubernetes pod restarted due to OOMKilled (Exit Code 137). An engineer inspects the pod and observes:

text
memory.max = 1 GiB
memory.current = 650 MiB

Diagnostic Question: #

Is this telemetry contradictory? How do you prove the past OOM event?

SRE Lesson & Mental Model: #

memory.current is a point-in-time snapshot. The previous process spiked to 1 GiB, was killed by SIGKILL, and the restarted container currently consumes only 650 MiB.

bash
cat /sys/fs/cgroup/<cgroup_path>/memory.events
# Verify non-zero oom_kill counter

Scenario 7 — Distinguishing cgroup OOM from Global Host OOM #

Situation: #

A pod restarts with Exit Code 137. Simultaneously, node available RAM dropped to 200 MiB.

Diagnostic Question: #

How do you confirm whether the pod was terminated due to its own cgroup limit or sacrificed by the host-wide OOM killer?

SRE Lesson & Mental Model: #

  • If host dmesg contains Out of memory: Killed process <pid>, the host ran out of physical memory and invoked the global killer based on badness score.
  • If host logs are clean but the container cgroup memory.events shows an incremented oom_kill, the container exceeded its isolated quota.

Scenario 8 — Distinguishing a Memory Leak from Allocator Caching #

Situation: #

A service processes a large batch job. Memory increases from 500 MB to 800 MB. After the job completes, memory usage remains parked at 800 MB indefinitely.

Diagnostic Question: #

Does this prove a memory leak?

SRE Lesson & Mental Model: #

No. C memory allocators retain free heaps to avoid the CPU cost of recurring syscalls. A genuine leak displays a stair-stepping baseline across consecutive job runs (500 MB → 800 MB → 1.1 GB → 1.4 GB).


Scenario 9 — Elevated Swap Usage Without Performance Degradation #

Situation: #

Monitoring shows 10 GB of swap consumed on a 32 GB RAM host.

Diagnostic Question: #

Is the host actively impaired by swap thrashing?

SRE Lesson & Mental Model: #

No. Swap usage could reflect inactive pages swapped out during a spike days ago. Validate active paging using vmstat 1. If si and so are 0, performance is unimpaired.


Scenario 10 — Confirming Memory Pressure via PSI #

Situation: #

Users experience elevated latency. Metrics report:

text
Available RAM: 400 MiB
vmstat si/so : Continuously elevated
CPU Usage    : 20%

Diagnostic Question: #

How do you conclusively prove memory exhaustion is causing CPU starvation?

SRE Lesson & Mental Model: #

Inspect /proc/pressure/memory. Elevated some and full percentages demonstrate that threads are stalled on memory allocation, preventing CPU execution.


Scenario 11 — High I/O Wait Without Drive Hardware Failure #

Situation: #

iowait reaches 35% with a load average of 12 on a 4-core machine.

Diagnostic Question: #

Does this prove disk hardware failure?

SRE Lesson & Mental Model: #

No. High iowait reflects threads blocked in TASK_UNINTERRUPTIBLE waiting on block requests. This frequently stems from remote NFS mount stalls, cloud volume bandwidth exhaustion, or hypervisor bus contention. Triage with iostat -xz 1.


Scenario 12 — High Disk Utilization at Low Transfer Bandwidth #

Situation: #

iostat shows:

text
Device    r/s    w/s    rkB/s    wkB/s    await    %util
sda     1200.0   40.0   4800.0   160.0    42.0ms   99.5%

Bandwidth is only 5 MB/s, yet the disk is 99.5% utilized.

Diagnostic Question: #

How can a device be 100% saturated at just 5 MB/s?

SRE Lesson & Mental Model: #

Small, random 4 KB reads consume disk seek operations and queue depth without transferring bulk data. 1200 random IOPS will completely saturate mechanical drives and IOPS-throttled cloud volumes.


Scenario 13 — Extreme Latency with Minimal Request Volume #

Situation: #

iostat reports:

text
Device    r/s    w/s    rkB/s    wkB/s    await     %util
sdb       6.0    4.0     24.0     16.0   280.00ms   98.0%

The device processes only 10 IOPS, yet average request latency is 280ms.

Diagnostic Question: #

Where is the fault located?

SRE Lesson & Mental Model: #

The system is not overloading the device with volume. The storage subsystem itself is degraded: network storage packet loss, RAID controller battery failure, or noisy hypervisor neighbors.


Scenario 14 — Virtualized Storage Stack Troubleshooting #

Situation: #

A virtualized cloud instance reports await > 200ms.

Diagnostic Question: #

Can this issue be diagnosed solely inside the guest OS?

SRE Lesson & Mental Model: #

No. Block requests traverse:

text
App → Guest FS → virtio driver → Hypervisor → SAN/EBS network → Storage Array

If the underlying hypervisor or SAN controller is saturated, guest-level optimizations will have zero impact.


Scenario 15 — The Golden Law: Metric ≠ Root Cause #

Situation: #

During an outage, an engineer claims: "CPU is only 18%, so compute is fine; the issue must be the database because some queries are slow!"

Diagnostic Question: #

What is the flaw in this reasoning?

SRE Lesson & Mental Model: #

A single metric is never the root cause. Low CPU often reflects threads blocked on downstream locks or I/O. Follow the full SRE diagnostic chain:

text
Symptom → Evidence → Bottleneck Resource → Layer → Dependencies → Confirm Hypothesis → Mitigate → Root Cause

Always ask:

"What does this metric measure, and what does it NOT measure?"