Benchmarks

Every number here was produced by running something, on hardware named by the GPU in it. Where a figure is derived rather than observed, it says so.

The short version: above about 2 ms a frame, a guest renders within 2% of the same machine's bare metal, and costs the same CPU. Below that, the cost of waiting for the GPU dominates a frame that barely exists.

What was measured, and how

nesprobe — the calibrated headless load probe from nesbox (tools/nesprobe), used unchanged. No compositor, no swapchain, no encode: it renders to an offscreen attachment and never presents, so what it reports is the stack's cost and not a window system's. --cost dials the per-frame fragment shader load, which is what lets one probe be submission-bound at one end and GPU-bound at the other.

  • 1920×1080, 30 s per run after an 8 s discard, three reps of every point
  • the figure is the median of three, with the rep-to-rep spread beside it; a difference smaller than the spread is reported as noise, because it is
  • the guest is compared only against its own host, on the same machine, minutes apart

The warm-up is not optional. An idle GPU sits in a low clock state and takes seconds to ramp; a p99 taken without discarding that window is a measurement of the clock ramp. This project has made that mistake and written it up.

The frame time

RTX 3060, driver 595.99.02, host Ubuntu 26.04 on a Ryzen 7 9850X3D. Guest: 2 vCPUs, Linux 7.2, module at its defaults.

--costhost p50guest p50Δrep spread h/g
800038.991 ms38.847 ms−0.4%0.0% / 0.0%
20009.8939.821−0.7%0.0% / 0.1%
4001.9782.011+1.7%0.1% / 0.0%
1000.5100.546+7.1%0.0% / 0.4%
00.0490.069+40.8%0.0% / 2.9%

A frame that takes the host 2 ms or more is within 2% in the guest. A frame that takes 0.05 ms is not, and the reason is not forwarding — it is that the guest sleeps waiting for the GPU, and the wake costs about 0.02 ms whatever the frame cost. A 60 Hz frame is 16.7 ms.

The CPU

A shared GPU is only worth sharing if its guests are cheap. Unpaced at ~100 fps, 12 s, one guest:

framesCPU usedCPU per frame
host, bare metal10090.40 s (3.3%)0.40 ms
guest10160.37 s (3.1%)0.36 ms

Nothing is spent forwarding a render loop because nothing is forwarded. Over a full sweep the guest drew 813,691 frames and the backend served 13,792 messages — one crossing per 59 frames, nearly all of it device setup. The count is the backend's own tally, not an estimate from a log.

Two bugs of our own are worth knowing about, because both were invisible in an ordinary benchmark.

The backend was serving the buffers a guest posts on the event queue as though they were requests — answering each with an error, handing it back, and being kicked again: 7.4 million callbacks in five seconds. It surfaced only when two guests ran at once and one of them failed to start its client, so the storm had the machine to itself. It had been running under every measurement taken since the event queue landed, and cost the guest +726% at cost 0 (now +41%) and every one of the 53 late frames above.

Before that, an earlier build had no .poll on its character devices. A file_operations with a NULL .poll is reported ready by the VFS every time it is asked, so the user-mode driver's wait never waited and the guest spun: 12.26 s of CPU for the same 12 s of work, a whole core per guest. The fix is a real .poll and an event queue that carries the host's readiness across. It is worth knowing that this class of bug is invisible in a frame-rate benchmark — the spinning guest was faster than bare metal.

Several guests on one card

The claim the whole design rests on, and the last one to be checked. Each guest gets its own backend and its own socket; they share one read-only root image.

nesprobe --cost 2000 in each, 30 s after an 8 s discard:

guestseachtotalp50 each
1102.9 fps102.99.72 ms
251.41, 50.95102.3619.576, 19.579 ms
425.84, 26.49, 25.57, 25.79103.6939.165, 39.164, 39.168, 39.165 ms

The total does not move as guests are added, and the split is even to four decimal places. Bare metal on the same card and load is 100.9 fps, so four guests together take nothing off the card that one process does not.

Rendering stays correct under contention — the offscreen regression in four guests at once gives red=11858 blue=53678 other=0 in every one, identical to a single guest.

And four of them encoding at the same time, which is what the product actually does: 575, 576, 577 and 573 frames, each paced at 16.67 ms — 60 Hz exactly — with no late frames and no NVENC session limit reached at four.

Four is what was run, not a limit. Each guest has 2 vCPUs on an 8-core host, so four is also where the host's CPUs are fully committed, and vkcube at 720p is a small load: four guests running a real game is a different measurement.

The whole chain

Not a synthetic load: a Wayland client presenting through a compositor in the guest, captured and encoded on the client's own device by Vulkan Video, with a receiver on the far end of the video socket.

receiver: frames=617 bytes=23643487
gap_ms mean 16.63  p50 16.67  p99 26.03  max 26.16   (60 Hz is 16.67)
ffmpeg -f null -: exit 0, no frame errors

Arrival spacing rather than a frame count, because a count cannot tell a slower pipeline from an on-time one that starts a second later — and here it was the second.

Frames arriving more than 25 ms after the one before, out of ~600: zero, with p99 spacing of 18.0 ms. An earlier version of this file reported 53 and called it an open question; that was a bug of ours in the event queue, described below, and not a property of the design.

A real game, on an A2000

Slime Rancher 2 under Proton, played by a person for about 15 minutes and streamed to a client at 1920×1080@60 in H.265 10-bit, 6 Mbps CBR, captured and encoded on the game's own device. RTX A2000 (70 W) at 615.71.09, an i7-8700K host with 12 threads, one guest with 6 vCPUs and 16 GiB. The game's swapchain is FIFO at 1920×1080.

GPU from nvidia-smi dmon every second (medians), CPU from each process's utime + stime every 5 s (averages over the window). Where a cell has two figures, they are the two client connections:

GPU busyNVENCpowerVM processbackendwhole host
streaming to a client, 10 min39–41%9–10%42–65 W2.10–2.33 cores0.18 cores3.65–4.00 of 12
game running, nobody connected, 3 min40%9%58 W2.31 cores0.18 cores3.93 of 12
  • The card is not the limit. A 60 Hz FIFO game holds the GPU at about 40%, with headroom.
  • Forwarding costs 0.18 of a core under a real game, flat across the session. That is the backend alone; the VM process is everything in the guest — game, Proton, compositor, capture and encode.
  • Encode is about 10% of NVENC at 1080p60 10-bit, rising to 20–26% for a few seconds around a reconnect.
  • No throttling. 718 of 720 seconds show no power violation, and none show a thermal one.

What this does not support: an overhead figure. The game was not run on the host directly, so none of these numbers is a cost of virtualisation — they are what one streamed game uses on this card. Frame rate and frame pacing were not recorded.

Re-taking these

The harnesses are in the private engineering notes rather than here, because they drive specific machines. What they do is simple enough to restate:

# bare metal: the probe, three reps of each cost, 30 s after an 8 s discard
for rep in 1 2 3; do
  for cost in 0 100 400 2000 8000; do
    nesprobe --device <n> --cost $cost --seconds 30 --warmup 8
  done
done

# guest: the same, inside, with the module at its defaults

Then compare each guest median against the host median from the same machine. Do not compare a millisecond figure from one box with a millisecond figure from another; compare the ratios.

Check for a stray VM before believing anything. A benchmark taken beside a guest somebody forgot to shut down looks like an ordinary result — slightly slower each run — and that is how a harness bug here cost a full set of numbers.

What these numbers do not support

  • No comparison with another hypervisor. None was run. "Faster than X" is not a claim this data can carry, and neither is "the fastest".
  • Four guests, not "many". Four ran; eight has not been tried, and neither have four guests running a real game.
  • No overhead figure for a real game. nesprobe is synthetic. One game has been streamed, on the A2000, and what it uses is above — but it has no bare-metal run beside it.
  • One card against bare metal. RTX 3060 at 595.99.02, which resolves to the 595.71.05 ABI profile. An RTX A2000 at 615.71.09 has resource numbers from one game and no bare-metal comparison. The shipped profiles are 535.129.03, 580.178.04 and 595.71.05, and a version older than the first is refused — see the README.

On this page