Skip to content

Sharing spare host capacity on GPU inference servers

Spare host CPU on a DGX B200 absorbed 49 busy cores with no resolved throughput loss (simulated).

Host CPU and memory sharing is in early access. Next, we take it to hardware and measure tail latency alongside recovered capacity.

Sharing must preserve response times

GPU inference still needs host CPU and memory. The host prepares requests, keeps work moving to the accelerator and sends responses back to clients.

Our simulation put substantial spare host capacity to work beside inference. That is the capacity we are opening to Virtual Machines and Containers on the same servers. The GPU stays assigned to inference; the capacity being shared is host CPU and memory.

The constraint is response time. Recovered host capacity only counts if inference stays fast and predictable. Host CPU and memory sharing is in early access.

What one simulated server shows

We used a GPU-node simulator to explore a modeled Kimi K3 workload on a DGX B200 with eight GPUs and 112 host CPU cores. The chart shows one 120-second run, including startup, with 48 agent sessions.

In that run, loading the model uses eight host cores for about 37 seconds. During serving, the simulated median is 0.4 busy cores and the peak is 1.38. Most of the host CPU sits unused.

The gap between 112 available cores and a serving demand under two cores is the capacity colocation is designed to use.

Cold start024680306090120Simulated secondsBusy host cores (of 112)

Simulated: Kimi K3 on one DGX B200, agent workload, continuous batching and CUDA graphs.
Show the numbers
Busy host cores while serving Kimi K3
MeasureBusy host cores
Cold start, 37 s8 of 112
Median while serving0.4 of 112
99th percentile while serving0.85 of 112
Peak while serving1.38 of 112

Throughput held within noise at 49 tenant cores

The A/B comparison uses separate runs from the chart. We compared serving alone with serving beside other workloads, three runs with different random seeds per configuration, each with a 15-second warm-up and 45 measured seconds. Percentiles pool across runs; the reported noise is the standard error.

In the preemptible mode, tenant work kept 49 cores busy and added 348 W. The throughput change was −0.60%, with ±1.92% estimated noise. The governed mode (a fixed reservation) gave tenants 56 cores and kept 31 busy, adding 261 W; its throughput change was −0.04%, with ±1.25% noise.

The opportunistic mode matched preemptible exactly. It takes cores from serving only when serving needs more than the 23 cores left free, and serving here needs fewer than two.

That is substantial modeled host work with no resolved throughput effect. The p99 time-to-first-token estimates carry ±29.25% and ±30.96% noise, far larger than the effect, so tail latency is the next thing we measure on hardware.

Colocation modes against serving alone, simulated A/B
ModeHost cores for tenantsTenant powerThroughput changep99 time to first token change
Preemptible or opportunistic, 80%, pinned89 offered, 49 busy, none taken from serving+348 W−0.60% ± 1.92%−6.13% ± 29.25%
Governed, 50%, pinned56 reserved, 31 busy+261 W−0.04% ± 1.25%−5.69% ± 30.96%

Separate cores still share a machine

Assigning inference its own CPU cores prevents other processes from running on those cores. It does not separate every resource they use.

Neighboring workloads can still compete for shared cache, memory bandwidth and communication between CPU sockets. Sharing is therefore judged on inference response time.

Research on host-side interference and CPU-related inference slowdowns maps these effects. Our experiments are built to separate spare capacity from capacity that serving needs during a burst.

[1][2]

Sharing is a per-workload choice

The sharing policy reserves a protected CPU and memory allocation for inference, then lends additional capacity to workloads that can tolerate its return. The aim is to use headroom as demand changes while keeping inference inside its response-time budget.

Batch work can accept interruptions that an interactive service cannot. The sharing policy reflects that difference and leaves unclassified workloads protected.

The amount available to share is measured on the serving stack and hardware in use.

Capacity returns to inference on each server

The approach adjusts neighboring workloads rather than modifying the inference stack. Resource decisions run on each server, with the platform setting the sharing policy. That puts the response close to the workload when inference needs capacity back.

On hardware, we measure how it tells interference from an ordinary change in inference demand, and how quickly capacity returns after the sharing policy changes.

Hardware differs in which resource controls it offers. The design accounts for those differences, including what happens when a control is unavailable.

PlatformSets the sharing policy
Sharing policy

Each GPU serverResource decisions run here

Lent capacityHost CPU and memoryEarly access

Neighboring workloadsTolerate the capacity’s return
Returns when needed

Protected for inferenceCPU and memory

InferenceThe GPU stays assigned
The platform sets each GPU server’s sharing policy. Inference keeps its GPU and a protected CPU and memory allocation. Spare host capacity is lent to tolerant workloads and returns when inference needs it. Host sharing is in early access.

What we test next on hardware

We begin with CPU-only experiments on control behavior and stability, then test inference on GPU servers. Larger runs add work spanning CPU sockets and sustained operation.

Each experiment pairs a serving baseline with a defined neighboring workload. It reports recovered capacity alongside response time, including the slowest responses and recovery after demand changes.

The opportunity is useful host work on servers already bought for their GPUs. The simulation shows enough capacity to go after; the hardware runs set how much we recover while holding response times.

References

  1. Blink

    arXiv 2604.07609, 2026.

  2. Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference

    arXiv 2603.22774, 2026.

All research