Sharing must preserve response times
GPU inference still needs host CPU and memory. The host prepares requests, keeps work moving to the accelerator and sends responses back to clients.
Our simulation put substantial spare host capacity to work beside inference. That is the capacity we are opening to Virtual Machines and Containers on the same servers. The GPU stays assigned to inference; the capacity being shared is host CPU and memory.
The constraint is response time. Recovered host capacity only counts if inference stays fast and predictable. Host CPU and memory sharing is in early access.
What one simulated server shows
We used a GPU-node simulator to explore a modeled Kimi K3 workload on a DGX B200 with eight GPUs and 112 host CPU cores. The chart shows one 120-second run, including startup, with 48 agent sessions.
In that run, loading the model uses eight host cores for about 37 seconds. During serving, the simulated median is 0.4 busy cores and the peak is 1.38. Most of the host CPU sits unused.
The gap between 112 available cores and a serving demand under two cores is the capacity colocation is designed to use.
Show the numbers
| Measure | Busy host cores |
|---|---|
| Cold start, 37 s | 8 of 112 |
| Median while serving | 0.4 of 112 |
| 99th percentile while serving | 0.85 of 112 |
| Peak while serving | 1.38 of 112 |
Throughput held within noise at 49 tenant cores
The A/B comparison uses separate runs from the chart. We compared serving alone with serving beside other workloads, three runs with different random seeds per configuration, each with a 15-second warm-up and 45 measured seconds. Percentiles pool across runs; the reported noise is the standard error.
In the preemptible mode, tenant work kept 49 cores busy and added 348 W. The throughput change was −0.60%, with ±1.92% estimated noise. The governed mode (a fixed reservation) gave tenants 56 cores and kept 31 busy, adding 261 W; its throughput change was −0.04%, with ±1.25% noise.
The opportunistic mode matched preemptible exactly. It takes cores from serving only when serving needs more than the 23 cores left free, and serving here needs fewer than two.
That is substantial modeled host work with no resolved throughput effect. The p99 time-to-first-token estimates carry ±29.25% and ±30.96% noise, far larger than the effect, so tail latency is the next thing we measure on hardware.
| Mode | Host cores for tenants | Tenant power | Throughput change | p99 time to first token change |
|---|---|---|---|---|
| Preemptible or opportunistic, 80%, pinned | 89 offered, 49 busy, none taken from serving | +348 W | −0.60% ± 1.92% | −6.13% ± 29.25% |
| Governed, 50%, pinned | 56 reserved, 31 busy | +261 W | −0.04% ± 1.25% | −5.69% ± 30.96% |
Separate cores still share a machine
Assigning inference its own CPU cores prevents other processes from running on those cores. It does not separate every resource they use.
Neighboring workloads can still compete for shared cache, memory bandwidth and communication between CPU sockets. Sharing is therefore judged on inference response time.
Research on host-side interference and CPU-related inference slowdowns maps these effects. Our experiments are built to separate spare capacity from capacity that serving needs during a burst.
Sharing is a per-workload choice
The sharing policy reserves a protected CPU and memory allocation for inference, then lends additional capacity to workloads that can tolerate its return. The aim is to use headroom as demand changes while keeping inference inside its response-time budget.
Batch work can accept interruptions that an interactive service cannot. The sharing policy reflects that difference and leaves unclassified workloads protected.
The amount available to share is measured on the serving stack and hardware in use.
Capacity returns to inference on each server
The approach adjusts neighboring workloads rather than modifying the inference stack. Resource decisions run on each server, with the platform setting the sharing policy. That puts the response close to the workload when inference needs capacity back.
On hardware, we measure how it tells interference from an ordinary change in inference demand, and how quickly capacity returns after the sharing policy changes.
Hardware differs in which resource controls it offers. The design accounts for those differences, including what happens when a control is unavailable.
Each GPU serverResource decisions run here
Lent capacityHost CPU and memoryEarly access
Protected for inferenceCPU and memory
What we test next on hardware
We begin with CPU-only experiments on control behavior and stability, then test inference on GPU servers. Larger runs add work spanning CPU sockets and sustained operation.
Each experiment pairs a serving baseline with a defined neighboring workload. It reports recovered capacity alongside response time, including the slowest responses and recovery after demand changes.
The opportunity is useful host work on servers already bought for their GPUs. The simulation shows enough capacity to go after; the hardware runs set how much we recover while holding response times.
References
- Blink
arXiv 2604.07609, 2026.
- Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
arXiv 2603.22774, 2026.
