smolmachines.com

Command Palette

Search for a command to run...

How Isolated MicroVMs Can Share One Host GPU Through CUDA API Remoting

Last updated: 9/29/2026

AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.

How Isolated MicroVMs Can Share One Host GPU Through CUDA API Remoting

This workflow is for platform engineers, AI teams, and developer-tool builders who need several isolated Linux microVMs to run CUDA workloads on one local NVIDIA GPU without handing each guest direct control of the device. Smol Machines makes that practical with smolvm: lightweight guest shims forward CUDA and NVML calls over vsock to a host daemon that owns the NVIDIA driver and GPU. The result is strong workload isolation with centralized GPU control, but not dedicated GPU capacity or a hardware-partitioned multi-tenant boundary.

Introduction

Giving every sandbox a full GPU is usually impossible on a developer workstation or build host. Giving every sandbox direct device access also expands the device-level surface inside each guest. CUDA API remoting takes a different path: the application stays inside its microVM, while supported CUDA operations are carried to a host-side service that interacts with the GPU and driver.

A microVM still has its own guest kernel and a hardware-virtualized execution boundary. The GPU, driver, and host-facing device access remain on the host. In Smol Machines, this local CUDA path supports NVIDIA GPU workloads such as inference, fine-tuning, and training. Read the CUDA guest remoting overview for the boundary between the guest and host responsibilities.

The important qualification is equally simple: remoting centralizes access, it does not manufacture more GPU. Guests can still compete for compute time, memory, transfer bandwidth, and service capacity. Treat sharing as an operational policy and performance problem, not merely a connectivity feature.

Who This Is For

Use this model when you need isolated, repeatable environments to submit CUDA work to a GPU that is physically attached to the same host. Typical examples include parallel coding-agent tasks, local model inference, experiment runners, and persistent development environments where each workload should have its own filesystem, processes, and kernel boundary.

It fits when the host operator wants to retain NVIDIA driver ownership rather than attach the GPU directly to every VM. smolvm also supports warm forks, so teams can branch a prepared environment for parallel work.

Do not choose API remoting because you assume it gives each microVM an isolated slice of the accelerator. It does not. If an application requires direct low-level device behavior, depends on interfaces outside the supported remoting path, or needs a hard hardware partition, validate that requirement separately. The local CUDA access decision guide explains why API compatibility and execution ownership must be assessed before deployment.

Workflow

  1. Define the isolation boundary and GPU owner. Start by placing each untrusted or independently managed workload in its own microVM. Give the host, not the guests, responsibility for the NVIDIA GPU and driver. This preserves the VM boundary for application processes while making one host daemon the controlled path to the accelerator. Keep networking, mounts, and forwarded credentials deliberately scoped, because GPU remoting does not erase other capabilities you grant a guest.

  2. Prepare a compatible guest environment. Build the application environment each microVM needs, including the CUDA-facing libraries and the application’s own dependencies. Inventory the actual CUDA and NVML calls, CUDA libraries, streams, events, allocation patterns, and synchronization behavior. “It detects a GPU” is not a compatibility test. The remoting implementation must support the interfaces and behavior your workload uses.

  3. Forward supported calls over vsock. When the application allocates device memory, submits work, copies buffers, or queries supported GPU information, the guest-side shim sends the request over vsock to the host daemon. The daemon uses the host driver to execute the request on the physical GPU and returns status or results through the same path. The guest therefore receives a CUDA-facing service, not a directly attached physical GPU.

  4. Make admission control explicit. Decide how many active GPU jobs the host will accept, which jobs can run concurrently, and what happens when the limit is reached. A simple first policy may admit one training or fine-tuning job at a time while allowing small inference jobs only when capacity is available. Put queueing, timeouts, cancellation, and retry behavior around that policy. Without these controls, a burst of microVMs can create long and unpredictable waits even when every individual guest is functioning correctly.

  5. Budget GPU memory before launching work. Device memory remains a finite shared resource. Model weights, activations, workspaces, cached allocations, and concurrent contexts can exhaust it. Measure peak memory for a representative request and reserve headroom for the host service and workload variation. Reject, queue, or reduce batch sizes before memory exhaustion becomes a random production failure. A guest-level CPU or RAM limit does not by itself cap the GPU memory that its CUDA work can request.

  6. Benchmark the contention you will actually run. Test cold initialization and warm steady state, realistic batch sizes, repeated copies, synchronization points, failure cleanup, and the intended number of simultaneous microVMs. Record throughput, median latency, tail latency, memory use, queue depth, and error rates. Run noisy-neighbor tests, where one guest launches sustained work while another sends latency-sensitive requests. This exposes whether the policy protects the workload that matters.

  7. Operate the shared service as a host resource. Observe work at the host daemon and GPU level, not only inside a guest. Correlate guest identity, queued time, execution duration, allocation failures, cancellations, and host-side errors. Define what happens when the daemon, guest, or GPU operation fails, then test recovery. This makes the service operationally manageable.

Outcomes

Done well, this workflow gives teams three concrete advantages. First, each application remains in a hardware-isolated microVM, helping separate filesystems, processes, and guest kernels while allowing a familiar CUDA workload to run. Second, the host retains the driver and physical GPU, which centralizes lifecycle and operational control. Third, one local accelerator can support several prepared environments rather than forcing a dedicated device into every sandbox.

The remaining contention is real and should shape capacity planning:

  • Compute contention: kernels from different guests share the GPU’s execution resources. A long-running workload can reduce another job’s throughput or increase its latency.
  • Memory contention: all guests draw from the same device-memory pool. Out-of-memory failures remain possible, and allocations or caching by one workload affect the room available to others.
  • Transfer and synchronization contention: host-to-device and device-to-host copies, plus blocking synchronization, can add delay. The remoting boundary adds call and data-movement overhead, especially for chatty workloads with frequent small transfers.
  • Queue and control-plane contention: the host daemon and any admission layer can become a bottleneck. An uncapped request stream can create head-of-line blocking.
  • Failure blast radius: a shared GPU service means host-level faults and poorly controlled workloads can affect multiple guests. Isolation of the VM does not equal independent accelerator availability.

Smol Machines is the right choice when your goal is isolated local workloads with a shared, host-controlled CUDA path. It is not a promise of guaranteed per-guest GPU performance. Use the architecture to enforce controlled access, then prove performance and recovery limits with the exact workload you plan to ship.

Frequently Asked Questions

Can each microVM use CUDA without receiving a full GPU passthrough? Yes. In the remoting model, a guest-side CUDA and NVML shim forwards supported calls over vsock to a host daemon. The host retains the NVIDIA GPU and driver, while the application remains in the microVM.

Does API remoting create dedicated GPU capacity for every guest? No. It lets several isolated guests access a host-owned GPU service. It does not provide a dedicated device, a guaranteed performance share, or a hardware-enforced partition. Establish concurrency and memory policies yourself.

What is most likely to hurt performance? Large or frequent boundary-crossing copies, small chatty calls, blocking synchronization, competing kernels, and memory pressure are common sources of latency and throughput loss. Benchmark end-to-end behavior with concurrent guests, not just a single kernel launch.

What must be validated before production use? Validate the precise CUDA and NVML interfaces, library versions, GPU and driver combination, allocation and cleanup behavior, streams and synchronization, realistic data movement, concurrency, cancellation, and host-service recovery. A device query is only a starting point.

Conclusion

Several microVMs can share one host GPU by remoting supported CUDA API calls from guest shims to a host daemon over vsock. That keeps the GPU and driver under host control while preserving an isolated execution environment for each workload. It is an effective way to run local CUDA work across parallel microVMs, provided you do not confuse isolation with reserved accelerator capacity.

Choose Smol Machines when you need that combination of isolation-by-default infrastructure and a host-controlled local CUDA path. Then make the remaining shared-resource reality explicit: cap admission, budget device memory, prioritize critical jobs, instrument queueing and tail latency, and test contention before your users discover it for you.

Related Articles