Skip to content

Offer GPU services from your own servers.

Use Neverinstall Private Cloud to provision customer workloads, reserve GPU capacity and set account limits. Send usage records to your billing system, which rates usage and invoices customers.

GPU servers are not yet a service.

Customers expect their own accounts, self-service launches, capacity limits and usage records. Building that layer takes engineering time away from the service itself.

Define the service and who operates it.

Customers receive separate accounts, networks, capacity limits and GPU reservations. They provision workloads through the Console, API or Terraform.

You choose service images, sizes and prices and operate the servers and network.

Neverinstall Private CloudOn your GPU servers
Manages

Customer A

Reserved GPUs
Capacity limits

Customer B

Reserved GPUs
Capacity limits
Sends
Usage recordsAllocation history for your billing system
Neverinstall Private Cloud separates customer accounts on your GPU servers. Each account has GPU reservations and capacity limits. Usage records preserve allocation changes and go to your billing system.
Customer services

Virtual Machines, Containers and Kubernetes

Automation

Console, API and Terraform

Billing

Usage records for your billing system

Pricing

Per node, not per core

Customers, capacity and records.

  • Customer organisations

    Each customer is an organisation with its own isolation tier: shared, dedicated, or dedicated with its own identity provider.

  • Reserved GPUs

    Reserve GPU cards, or whole nodes, for one customer. Each card shows which customer holds it.

  • Capacity ceilings

    Set each customer’s ceiling for vCPU, memory, storage and GPUs. The customer’s admins set limits within it for each of their tenants.

  • Usage records

    Hourly GPU, vCPU, memory and storage records for each customer, exported as CSV or JSON for your billing system.

Connect capacity records to your billing system.

Complete allocation history

Usage records preserve allocation changes across resize, stop and restart. Your billing system rates the records and creates invoices.

Service scope

Confirm supported GPU modes and any workload that combines GPUs across servers.

Test one service through its lifecycle.

  • Customer provisioning

    Launch through the Console, API or Terraform. Verify the GPU profile, quotas and network access.

    For billing reconciliation

    The requested and delivered profile, launch time and permission results.

  • Changes and isolation

    Resize, stop, restart and remove workloads. Check tenant separation and released capacity.

    For billing reconciliation

    Workload events, capacity changes and blocked cross-tenant access.

  • Billing records

    Match the allocation history to each workload event and your billing system’s required fields.

    For billing reconciliation

    A sample export with units and timestamps, reconciled to events with any discrepancies recorded.

Related

Questions from cloud providers.

Can we use the GPU servers we already own?

Yes. Neverinstall Private Cloud runs on x86 and arm64 servers with NVIDIA or AMD GPUs, and one cluster can mix nodes. We review accelerators, storage and networking for the service you plan to offer.

Which GPU sharing modes can we offer?

vGPU, MIG, time-slicing and passthrough depend on the card and workload. Confirm the GPU model, driver and service profile before the pilot.

Can customers use GPUs across servers?

Yes, where their workload supports it. The platform can schedule across GPU servers and support one workload that combines GPUs across servers. Review the application, network and GPU configuration before offering the service.

How are customers kept apart?

Each customer has Tenant Networks, quotas, users and GPU reservations. Disks use keys in the tenant’s vault, with support for customer-provided keys. Test cross-tenant access during the pilot.

Can customers run their own model servers?

Yes. Use the built-in serving layer or a customer’s model server on reserved GPUs. Confirm capacity and network access for the service.

What do the billing records cover?

Records preserve the allocation history across resize, stop and restart events. Reconcile a sample export with the workload lifecycle during the pilot.

Is host CPU and memory sharing available?

Host CPU and memory sharing on GPU servers is in early access. Review it separately from the service you plan to launch.

How is the platform priced?

Neverinstall Private Cloud is priced per node, with no per-core or per-Virtual-Machine charge. GPU servers use the same platform license. Resold desktop sessions use Neverinstall Virtual Desktops pricing per user.

Plan your first customer service.

Tell us the service, GPU hardware and billing workflow you plan to use. Our engineers will help scope the pilot.