Skip to content

Serve models and share GPUs on your servers.

Allocate GPUs to models and Virtual Machines through Neverinstall Private Cloud.

Add Neverinstall Virtual Desktops for graphics applications. We review your cards, drivers and workloads before deployment.

Models, GPU Virtual Machines and graphics applications.

Model Serving

Deploy an open-weight model with the built-in serving layer, or bring your own model server on reserved GPUs.

GPU Virtual Machines

Assign a whole card or a supported GPU profile to a Virtual Machine. Confirm the mode against your card and workload in the hardware review.

GPU Desktops

Give CAD and rendering teams Windows or Linux desktops with GPU resources, delivered through Neverinstall Virtual Desktops.

Models and GPUs in the Console.

Choose a mode supported by your hardware.

Isolation depends on the mode, card and driver.

Choose a mode supported by your hardware.
HardwareModes to reviewBefore deployment
NVIDIAvGPU, MIG, time-slicing and passthroughConfirm the card, driver and supported profiles for your workload.
AMDPassthrough and SR-IOV partitionsConfirm support on the specific card in the hardware review.
Across serversJob scheduling or a workload spanning GPUsMulti-server execution depends on the workload and interconnect. Review both before deployment.

Assign GPU resources and track allocations.

GPU server

GPUAllocated GPU resources

AI inference
Graphics workloads

CPU and memoryHost sharingEarly access

Virtual Machines
Containers

Illustrative allocation on one GPU server. Host CPU and memory sharing is in early access.

One GPU server. GPUs are allocated to AI inference and graphics workloads. In a separate CPU and memory area, Virtual Machines and Containers use host sharing, which is in early access.

Assign resources

Share cards with a team and assign GPU resources to its model server, Virtual Machines or desktops.

Set access and limits

Apply team roles, quotas for each GPU profile and network rules for model endpoints.

Review allocations

Track reserved GPU resources by team, with history through resize, stop and restart. Allocation records do not measure GPU utilization.

Size the deployment with your workload.

  1. Check the hardware

    Review cards, drivers, GPU memory, storage and networking against the workload.

  2. Deploy and set limits

    Assign GPU resources, deploy the workload and configure team access and quotas.

  3. Measure and size

    Test latency, concurrency and memory use. Use the results to plan capacity and the quote.

GPU servers are quoted per node.

Your quote defines the configuration and workload scope.

Platform

Per node, with Neverinstall Private Cloud

Desktop delivery

Neverinstall Virtual Desktops is licensed separately

Separate costs

Hardware and any required third-party licenses

Where teams start.

Evaluate in depth.

GPU allocation and host CPU and memory sharing.
  • Inference keeps priority

    Host CPU and memory sharing is in early access. Test usable host capacity and the effect on inference response times with your hardware and workloads.

  • Track GPU allocations

    Track reserved GPU resources per team, including resize, stop and restart history.

GPU questions.

Do you host our models or see our prompts?

Models and inference run on your servers and network. Support access is agreed with your team first.

Can GPU workloads run without internet access?

Yes. Neverinstall Private Cloud can run offline. Agree on the model, driver and update requirements for the disconnected environment.

What does GPU metering record?

It records allocated GPU resources over time, including resize, stop and restart. Use workload measurements to assess utilization and inference performance.

Plan your GPU deployment.

Tell us how many GPU servers you have and which AI and desktop workloads should share them.

We will map the allocation with you.