Engineering notes for infrastructure teams.
How we isolate tenant networks, recover unfinished operations and account for shared GPU capacity, with the data behind each result.
Engineering notes
- Who acted in a shared AI desktop session?A shared-desktop prototype labels the source of each input and lets the person override the agent. Next, we are linking each action to the network requests it caused.
- Measuring GPU capacity from recorded allocationsThe meter reports reserved GPU capacity from control-plane records, with resize, stop and restart history and no guest agent. Read how it calculates quantities, recovers missed hours and closes reporting periods.
- Building overlapping tenant networks with CiliumWe extended Cilium so tenants can reuse private address ranges across Virtual Machines and Containers. The article covers the design and a three-node test run.
- Keeping failed cloud deletions visible and retryableA failed delete keeps its record, so cleanup can resume. The article covers retained records, ordered cleanup and the checks made before a delete reports complete.
- Sharing spare host capacity on GPU inference serversIn a DGX B200 simulation, neighboring workloads kept 49 host CPU cores busy with no throughput change beyond the noise. Host CPU and memory sharing is in early access.
- Model checking exposed failures in cloud recoveryA June 2026 review with TLA+ models produced 14 correctness findings. The models exposed recovery failures and changed how retries and completion checks work.
