Manufacturing company · Jul–Oct 2026
Hands-off reliability for a container platform
Turned manual patching, certificate renewals, and disk cleanup across the Linux estate into tested, fail-closed automation — and fixed the monitoring that was supposed to catch problems.
Business impactRoutine maintenance went from staff-hours on a calendar to unattended runs, with less downtime risk and alerts that actually fire.
Role: Owned the infrastructure automation roadmap and its delivery
- of Linux hosts patched unattended in one window
- 100%
- disk reclaimed by safe automated cleanup
- ~242 GB
- repeat CI incidents from stale base images
- 4 → 0
The problem
Servers were patched by hand, disks crept toward full, certificates were renewed from a calendar reminder, and builds kept failing on outdated base images. Worse, monitors meant to catch failures had quietly stopped firing.
How it works
- 01
Fail-closed patching
The RMM tool only schedules the jobs. Each run takes a lock, moves workloads off the node, checks cluster health before and after, and stops on anything unexpected.
- 02
Safe cleanup
Disk cleanup keeps everything still in use and aborts rather than guess.
- 03
Inventory-driven certificates
Renewals run from a certificate inventory, deploy where it's safe, queue a runbook where a person must approve, verify the result, and roll back on mismatch.
- 04
Monitors that can't fool themselves
Heartbeat monitoring rebuilt so a check can never report success on its own behalf.
What I built
- The patching workflow, with a test suite grown from 28 to 64 cases that caught a bug which would have left a server out of service.
- Automated disk cleanup, certificate renewal, and a rewritten certificate checker.
- On-demand base-image refresh in the build pipeline.
Results
- Every Linux host patched to the same kernel in a single maintenance window.
- ~242 GB reclaimed; servers that were ~80% full dropped to under a third.
- Wildcard certificates moved from a paid commercial CA to free, automated Let's Encrypt.
- Ended a class of build failure that had happened four times.
What I learned
- Automation that can fail open is worse than no automation. Every job here stops on the first thing it doesn't expect.
- Audit the alarms, not just the systems. Silent monitors are the most expensive bug you'll never see.
Stack
- Linux
- Docker
- Bash
- Let's Encrypt / ACME
- CI/CD