Two Cards, One Model, and the Ethernet Between Them

I run a small fleet of self-hosted AI agents — they live in Slack, one per job, all of them on a homelab Kubernetes cluster spread across a few Proxmox servers. Their local brain is a pair of Tesla T40s: 24 GB datacenter GPUs, one in each of two physical servers,...