Skip to content

Host CPU and memory sharing beside GPU inference.

Compare simulated host capacity, power and response times with sharing off and on. Use the results to plan what to test on your hardware.

Loading the simulator in your browser.

How the simulator works

  • Follow a workload over time

    The simulator advances through the model’s computation one step at a time. Requests arrive, wait and complete as the run progresses. The chart shows how much host CPU inference and other work use along the way.

  • Compare sharing policies

    The example starts in Governed mode: a fixed core reservation for other work. Run the off/on comparison for a baseline, then try work that yields to inference or uses idle cores.

  • Check power as well as CPU

    Other workloads add power demand even when inference has enough CPU. In the simulation, exceeding power or cooling limits reduces GPU performance.

  • Compare response times

    The A/B test runs the same scenario with sharing off and on. It reports throughput, response time and estimated variation between runs. Changes smaller than that variation are marked inconclusive.

  • Frame your hardware tests

    This version uses fixed assumptions and a simplified memory model. Use it to shape hardware testing, particularly around response times and competition for shared resources.

What it models

The scenario combines hardware specifications with fixed modeling assumptions. Preemptible work assumes immediate CPU handover. Performance, power and financial results are estimates. Each assumption shows its source: specification, research or model estimate.

Model

Kimi K3, 2.8T total and 104B active parameters

Number formats

MXFP4 model weights; MXFP8 cached context and intermediate values

Serving

Model split across eight GPUs, with continuous batching and shared cached context

Platforms

DGX B200, GB200 NVL72, HGX H200, DGX H100 and MI355X

Host

CPU cores, DRAM, memory bandwidth, PCIe and NVMe

Limits

Estimated demand compared with power and cooling capacity

Simulator questions

Is this measured on real hardware?

No. It is a simulator, and the results come from a model running in your browser. Use them to design an experiment. The research article lays out the hardware testing we run next.

Where does the simulation run?

In your browser. The calculation runs in the background so you can continue using the page.

Why Kimi K3?

Its size leaves little spare GPU memory in the modeled B200 setup. That makes memory pressure and the use of host memory easier to explore.

Why is the simulated host CPU use low?

The model assumes batched GPU execution with relatively little host work during steady serving. Loading the model and processing bursts use more CPU.

Why does the node or rack count change?

GPU memory capacity and supported number formats differ across platforms. The simulator accounts for those differences, then adds enough nodes or racks to hold the model weights. That is why changing the platform can also change the node or rack count.

How should I compare the sharing modes?

The A/B test compares sharing off and on using the same workload. Compare response times and power alongside the capacity each mode makes available.

Does core pinning remove interference?

No. Pinning keeps workloads on separate cores, but they still share cache and memory bandwidth. The simulator includes an estimate of that remaining interference, which you can inspect in the assumptions table.

What does the A/B test tell me?

It compares simulated throughput and response times with sharing off and on. The p99 time to first token is the time within which 99% of simulated requests receive their first token. Estimated noise shows which differences are resolved. Financial estimates use fixed example rates.

Plan how to share your GPU servers.

Host CPU and memory sharing is in early access. Tell us about your GPU servers, serving workload and response-time requirements so we can discuss what to test.