Manufacturing company · Jul–Oct 2026

Hands-off reliability for a container platform

Turned manual patching, certificate renewals, and disk cleanup across the Linux estate into tested, fail-closed automation — and fixed the monitoring that was supposed to catch problems.

Business impactRoutine maintenance went from staff-hours on a calendar to unattended runs, with less downtime risk and alerts that actually fire.

Role: Owned the infrastructure automation roadmap and its delivery

of Linux hosts patched unattended in one window
100%
disk reclaimed by safe automated cleanup
~242 GB
repeat CI incidents from stale base images
4 → 0

The problem

Servers were patched by hand, disks crept toward full, certificates were renewed from a calendar reminder, and builds kept failing on outdated base images. Worse, monitors meant to catch failures had quietly stopped firing.

How it works

  1. 01

    Fail-closed patching

    The RMM tool only schedules the jobs. Each run takes a lock, moves workloads off the node, checks cluster health before and after, and stops on anything unexpected.

  2. 02

    Safe cleanup

    Disk cleanup keeps everything still in use and aborts rather than guess.

  3. 03

    Inventory-driven certificates

    Renewals run from a certificate inventory, deploy where it's safe, queue a runbook where a person must approve, verify the result, and roll back on mismatch.

  4. 04

    Monitors that can't fool themselves

    Heartbeat monitoring rebuilt so a check can never report success on its own behalf.

What I built

  • The patching workflow, with a test suite grown from 28 to 64 cases that caught a bug which would have left a server out of service.
  • Automated disk cleanup, certificate renewal, and a rewritten certificate checker.
  • On-demand base-image refresh in the build pipeline.

Results

  • Every Linux host patched to the same kernel in a single maintenance window.
  • ~242 GB reclaimed; servers that were ~80% full dropped to under a third.
  • Wildcard certificates moved from a paid commercial CA to free, automated Let's Encrypt.
  • Ended a class of build failure that had happened four times.

What I learned

  • Automation that can fail open is worse than no automation. Every job here stops on the first thing it doesn't expect.
  • Audit the alarms, not just the systems. Silent monitors are the most expensive bug you'll never see.

Stack

  • Linux
  • Docker
  • Bash
  • Let's Encrypt / ACME
  • CI/CD