Complete allocation history
Usage records preserve allocation changes across resize, stop and restart. Your billing system rates the records and creates invoices.
Use Neverinstall Private Cloud to provision customer workloads, reserve GPU capacity and set account limits. Send usage records to your billing system, which rates usage and invoices customers.
Customers expect their own accounts, self-service launches, capacity limits and usage records. Building that layer takes engineering time away from the service itself.
Customers receive separate accounts, networks, capacity limits and GPU reservations. They provision workloads through the Console, API or Terraform.
You choose service images, sizes and prices and operate the servers and network.
Customer A
Customer B
Virtual Machines, Containers and Kubernetes
Console, API and Terraform
Usage records for your billing system
Per node, not per core
Each customer is an organisation with its own isolation tier: shared, dedicated, or dedicated with its own identity provider.
Reserve GPU cards, or whole nodes, for one customer. Each card shows which customer holds it.
Set each customer’s ceiling for vCPU, memory, storage and GPUs. The customer’s admins set limits within it for each of their tenants.
Hourly GPU, vCPU, memory and storage records for each customer, exported as CSV or JSON for your billing system.
Usage records preserve allocation changes across resize, stop and restart. Your billing system rates the records and creates invoices.
Confirm supported GPU modes and any workload that combines GPUs across servers.
Launch through the Console, API or Terraform. Verify the GPU profile, quotas and network access.
The requested and delivered profile, launch time and permission results.
Resize, stop, restart and remove workloads. Check tenant separation and released capacity.
Workload events, capacity changes and blocked cross-tenant access.
Match the allocation history to each workload event and your billing system’s required fields.
A sample export with units and timestamps, reconciled to events with any discrepancies recorded.
Yes. Neverinstall Private Cloud runs on x86 and arm64 servers with NVIDIA or AMD GPUs, and one cluster can mix nodes. We review accelerators, storage and networking for the service you plan to offer.
vGPU, MIG, time-slicing and passthrough depend on the card and workload. Confirm the GPU model, driver and service profile before the pilot.
Yes, where their workload supports it. The platform can schedule across GPU servers and support one workload that combines GPUs across servers. Review the application, network and GPU configuration before offering the service.
Each customer has Tenant Networks, quotas, users and GPU reservations. Disks use keys in the tenant’s vault, with support for customer-provided keys. Test cross-tenant access during the pilot.
Yes. Use the built-in serving layer or a customer’s model server on reserved GPUs. Confirm capacity and network access for the service.
Records preserve the allocation history across resize, stop and restart events. Reconcile a sample export with the workload lifecycle during the pilot.
Host CPU and memory sharing on GPU servers is in early access. Review it separately from the service you plan to launch.
Neverinstall Private Cloud is priced per node, with no per-core or per-Virtual-Machine charge. GPU servers use the same platform license. Resold desktop sessions use Neverinstall Virtual Desktops pricing per user.
Tell us the service, GPU hardware and billing workflow you plan to use. Our engineers will help scope the pilot.