ROOT // SYSTEMS
/ ROOT // SYSTEMS / WAR STORIES

WAR STORIES

Things that broke at 03:17.

root@systems:~$ ls
./The Server That Wouldn't Die
./When DNS Wasn't DNS
./The Friday Kernel Upgrade
./rm -rf: A Career Development Program
./The RAID Array That Lied
./That One Broken Symlink
./Production Was Full of Logs
./Why We Don't Edit Production Servers
INITIAL INDEX
01 / WAR STORIES

The Server That Wouldn't Die

A server can survive far more abuse than anyone expects. This story explores the machines that refuse to disappear: services that respawn after being killed, processes that remain attached to unexpected supervisors, failed shutdowns, and systems that mysteriously return after administrators believe they have been stopped. The investigation leads through process trees, service managers, watchdogs, cron jobs and automation—revealing that sometimes the problem isn't how to kill a process, but discovering what keeps bringing it back.

02 / WAR STORIES

When DNS Wasn't DNS

Few infrastructure problems create as much confusion as a failure that looks exactly like DNS. An application cannot connect, a hostname resolves differently from another machine, and every test appears to point toward the resolver. The real cause may instead be /etc/hosts, split-horizon DNS, stale caches, search domains, IPv6, resolver configuration or an application using a different resolution mechanism entirely. This is a practical investigation into why "DNS is broken" is often only the beginning of the diagnosis.

03 / WAR STORIES

The Friday Kernel Upgrade

Kernel upgrades are routine—until the machine doesn't come back. A seemingly harmless update can expose problems involving bootloaders, initramfs images, storage drivers, filesystem modules, third-party kernel modules or hardware compatibility. This story follows the kind of Friday-afternoon failure that turns a routine maintenance window into an incident, demonstrating why kernel changes require rollback planning, console access, tested recovery procedures and a clear understanding of the complete boot chain.

04 / WAR STORIES

rm -rf: A Career Development Program

Recursive deletion is one of the most powerful and dangerous capabilities available from a Unix shell. A misplaced variable, an unexpected empty value, a symbolic link assumption or a command executed from the wrong directory can turn a routine cleanup operation into data loss. This story examines how destructive commands happen, why shell expansion and quoting matter, and how permissions, backups, immutable infrastructure and automation can reduce the consequences of human error.

05 / WAR STORIES

The RAID Array That Lied

RAID can protect against certain disk failures, but it cannot guarantee data integrity or eliminate operational risk. An array may report itself as healthy while another component is silently failing, rebuilds can expose previously unreadable sectors, and hardware RAID controllers can obscure the actual condition of individual disks. This story explores degraded arrays, SMART data, rebuilds, filesystem corruption and the uncomfortable distinction between redundancy and backup.

06 / WAR STORIES

That One Broken Symlink

A symbolic link is only a reference to another path, but a single broken link can have consequences far beyond a missing file. Applications may fail to start, libraries may no longer resolve, configuration directories can point at unexpected locations and deployment systems can silently install files somewhere other than intended. This story follows the investigation of a deceptively small filesystem problem and the lessons it teaches about paths, symlinks, mount points, package management and system configuration.

07 / WAR STORIES

Production Was Full of Logs

Disk exhaustion is one of the oldest and most reliable ways to bring a production system to its knees. Applications continue writing logs, databases consume transaction or journal space, container runtimes accumulate unused layers, and suddenly a filesystem reaches 100 percent utilization. The resulting failures can be surprisingly broad: services cannot write temporary files, databases stop accepting transactions and even administrators may struggle to log in. This story explores how to identify what consumed the space and how log rotation, retention policies, monitoring and capacity planning prevent a repeat.

08 / WAR STORIES

Why We Don't Edit Production Servers

Manually fixing a production server can be remarkably effective—and remarkably dangerous. A change made directly to a live system may solve the immediate problem while leaving the configuration undocumented, unreproducible and vulnerable to being overwritten by the next deployment or automation run. This story examines configuration drift, undocumented fixes, emergency changes and the transition from "fix the server" to "fix the system." Infrastructure as code, immutable infrastructure, configuration management and Git-based workflows provide the mechanisms for making operational changes repeatable and auditable.