Skip to content
Jacob NolletteCloud & software engineer
Topic · 8 case studies

Reliability

Monitoring that watches outcomes instead of jobs, backups that have actually been restored, and the incidents that taught me the difference.

Nine Copies of Every Write

I run a five-node Proxmox cluster with hyperconverged Ceph and a three-node Kubernetes control plane on top of it. It has been up for over a year. Nothing crashes, deployments roll, pods schedule. It was just slow, in a way I could never quite point at. A kubectl...