Skip to content

Measuring GPU capacity from recorded allocations

Our meter reports reserved capacity from control-plane records without a guest agent. It records resize, stop and restart history, recovers missed hours after an outage and exports allocation records in an air-gapped deployment.

Busy time and reserved time are different quantities

An idle GPU can still be fully reserved. Another tenant cannot use it just because its current owner has stopped sending work.

That distinction matters when reporting consumption. Allocation tells us how much capacity was reserved and for how long. Utilization tells us how actively that capacity was used. Both are useful, but they answer different questions.

There is also a measurement constraint. When a GPU is passed directly into a Virtual Machine, the guest controls it and host monitoring has limited visibility. Other attachment modes expose different measurements, so the meter starts from records that every mode has.

[1][2]

Start with what the platform can account for

Our meter derives GPU-hours from the control plane’s allocation records. It uses the same basis for vCPU and memory quantities: the amount reserved and the part of each hour covered by the recorded allocation.

This lets the meter report reserved capacity without depending on a monitoring agent inside the tenant’s machine. Keeping a deletion record also preserves the end of an allocation after the resource has gone.

Make a partial hour easy to explain

If one GPU is allocated for half an hour, it contributes half a GPU-hour. If four vCPUs are allocated for fifteen minutes, they contribute one vCPU-hour.

GPU profiles remain separate in the reports, so a full card is not combined with a slice. This matters when teams compare allocations or apply their own rates: the quantity needs to retain the kind of capacity it describes.

We calculate partial intervals with exact arithmetic and round when recording the quantity. Repeating the calculation over the same records produces the same result. That gives operators a reproducible quantity to reconcile with an export.

Partial hours, prorated
AllocationTime allocatedQuantity
One GPUHalf an hourHalf a GPU-hour
Four vCPUsFifteen minutesOne vCPU-hour

A retry should not change the total

The meter can repeat an open-hour calculation without adding a second allocation quantity for the same allocation. After an outage, it resumes from the last completed hour. A failed hour stays unfinished until it can be calculated successfully.

Reports retain the resource and tenant context recorded at the time. A later rename does not rewrite earlier exports, and tenant offboarding does not erase the period before departure. Those are necessary properties for explaining a historical total.

These requirements make metering a record-keeping problem as well as an arithmetic problem. A total is only useful if the platform can explain where it came from.

Closed periods need a clear boundary

Open reporting periods can be recalculated as work is recovered. Closing a period freezes its rows and rejects ordinary recomputation. An outage recovery therefore cannot silently rewrite a total an operator has already closed.

Corrections to a closed period will be recorded as explicit adjustments, kept separate from the original rows.

Tenants and operators can retrieve scoped CSV, JSON or NDJSON exports. The export path does not require an external billing service, so an air-gapped deployment can keep its allocation records and reporting inside its own environment.

Open period

Allocation rowsRecalculated as work is recovered
Close the period

Closed period

Frozen rowsOrdinary recomputation rejected
AdjustmentsCorrections kept separate from the original rowsPlanned
Scoped exports
CSV, JSON or NDJSONNo external billing service required
Open periods can be recalculated as work is recovered. Closing a period freezes its rows. Corrections will be recorded as separate adjustments. Tenants and operators export scoped CSV, JSON or NDJSON without an external billing service.

What operators can use today

The allocation history, ledger, outage recovery, period closure and exports are built. Metered resources include GPUs, vCPUs, memory, storage and network resources. The result is a common record of allocated capacity that operators can inspect before applying prices.

The meter records each resize, stop and restart as it happens, so every part of an hour is metered at the shape it ran. Measured utilization and service-reported quantities come next, labeled so readers can tell a reservation from an observation and an allocation total from an invoice.

Keep monitoring useful on its own terms

OpenCost, Ceilometer and CloudKitty address related questions about resource cost, collection and rating. They are useful context for deciding what a meter should report.

Operational monitoring remains valuable alongside allocation records. It can help an operator understand whether reserved capacity is being used well. A report should make the origin of each quantity clear enough to support that comparison.

[3][4][5]

References

  1. NVIDIA Virtual GPU Software User Guide

    Monitoring GPU Performance.

  2. NVIDIA GPU Operator

    GPU Operator with KubeVirt.

  3. OpenCost Specification

    Workload cost as the greater of request and usage.

  4. OpenStack Ceilometer

    Data collection.

  5. OpenStack CloudKitty

    CloudKitty documentation.

All research