How to Measure Sandbox Time-to-Ready: Boot, Pull, Setup, and Prepared Artifacts
AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.
How to Measure Sandbox Time-to-Ready: Boot, Pull, Setup, and Prepared Artifacts
This workflow is for platform engineers comparing sandbox approaches for agents, CI, browser automation, or untrusted code, where a running VM is not necessarily ready to work.
Introduction
The direct answer is: measure time-to-ready from a request being accepted to a sandbox passing a workload-specific readiness check. Report VM boot, image pull, dependency installation, and prepared-artifact restoration as separate stages, then report the end-to-end distribution as the decision metric.
VM boot alone answers a narrow question: how quickly does the runtime create a running guest? It does not tell you whether the correct image is available, packages are installed, services have started, credentials or fixtures are in place, and the first real command can succeed. A sandbox that boots quickly but spends minutes downloading dependencies is not ready quickly.
Use percentiles, not one average: p50 for typical experience, p95 for planning, and p99 when fan-out makes tail delays expensive. Keep cold, warm, and prepared-artifact paths distinct so cache hits do not hide the scale-out path.
Who this is for
This measurement workflow is for teams that need repeatable isolated environments and want to make a purchase or architecture decision with evidence. It applies when a task requires a specific operating system image, application runtime, project dependencies, tools, fixtures, or local services before an agent or job can start.
For workloads that execute untrusted code, do not trade away boundary and policy verification for a prettier timing number. Smol Machines runs workloads in hardware-virtualized Linux microVMs with a guest kernel, while networking is off by default. A mounted directory, enabled network route, or forwarded credential is still an explicit capability that needs review. Readiness should include the intended access policy, not just a passing process check.
Workflow
1. Define “ready” as a usable, verifiable state
Write one acceptance test that represents the first meaningful action of the workload. For a coding agent, that might be running the project test command and confirming required tools are on the path. For browser automation, it might be opening the target browser and loading a local fixture. For CI, it might be a clean build of a small representative target.
Start the timer when the control plane accepts the request. Stop it only when this test passes. Record a failure separately rather than excluding it from the data. A fast path that frequently fails to become usable is not a fast time-to-ready path.
Version the test inputs: image digest or artifact version, repository revision, dependency lockfile, machine resources, network policy, and cache policy.
2. Instrument the four timing boundaries
Collect monotonic timestamps around these stages:
- Provisioning and VM boot: request accepted to a guest that can execute a minimal command. Split queue time from actual boot if the platform can expose both.
- Image acquisition: start to end of pulling, unpacking, or making the required base image locally available. Report zero explicitly on a verified local cache hit rather than blending it with misses.
- Dependency and setup work: the time for package installation, language dependency resolution, source checkout, migrations, fixture creation, and service initialization that occur after boot.
- Prepared artifact restoration: request accepted to restoration of a prebuilt environment, followed by the same readiness test. Do not count an artifact as ready merely because it was fetched or mounted.
Where possible, separate download, CPU, and waiting time. An image pull may be network-limited, while installation may be dominated by compilation or a package registry.
3. Run comparable cold, warm, and prepared trials
Run each path with the same workload and readiness test. A practical minimum is three cohorts:
- Cold base image: no local image or dependency cache, then boot, pull, install, and test.
- Warm base image: image and allowed caches are present, then boot, install any remaining dependencies, and test.
- Prepared artifact: restore a versioned, validated environment and run the identical test.
Use enough repetitions to characterize tail behavior. Include forced cache misses and concurrent starts that resemble peak demand, and record host capacity, image and artifact sizes, and registry locality.
A prepared environment should include only what the task needs and be rebuilt from declared inputs. Smol Machines supports a checked-in Smolfile that declares the image, resources, network policy, mounts, ports, and setup commands, making the baseline reviewable before repeated use.
4. Validate the prepared artifact before publishing it
Build the artifact, restore it on a clean compatible host, and run the readiness test. Verify the expected revision, tools, services, permissions, and policy. Confirm that no hidden runtime download or host dependency is completing the work.
For smolvm, a stateful prepared VM can be packaged as a self-contained .smolmachine artifact. The product documents sub-200 ms cold starts for pre-baked artifacts on supported hosts, but your benchmark should treat that as a hypothesis to validate with your artifact size, host architecture, and workload. The relevant outcome is the p95 time until your test passes, not a vendor headline.
A concise guide to prebuilt rollout environments explains the operational advantage: build and verify the environment before demand rather than reinstalling it for every worker.
5. Compare the end-to-end cost, not only the request path
Prepared artifacts shift work earlier. Track build and validation duration, storage footprint, publish and distribution time, rebuild frequency, and failure rate. Compare those costs with cumulative setup time avoided across expected launches.
A simple decision calculation is:
avoided request-path setup time × expected launches per artifact version
Compare that value with preparation and maintenance overhead. Preparation may not amortize if the environment changes every run, but often will when hundreds of short tasks share one validated baseline.
6. Publish a scorecard that makes trade-offs clear
For each approach, publish p50, p95, and p99 end-to-end time-to-ready; the median and p95 of each stage; readiness failure rate; cache state; concurrency; artifact or image version; and the exact readiness command. Add security and operational conditions, including network mode, mounts, secret access, cleanup behavior, and isolation model.
This scorecard turns a vague “fast sandbox” claim into a decision record. If boot is small and setup is large, optimize the environment. If artifact restoration has a long tail, investigate transfer, storage, and host compatibility. Fix artifact validity before celebrating lower latency.
Outcomes
A disciplined benchmark produces four useful outcomes:
- A real service-level metric: teams can state how long users wait before work starts, rather than quoting process creation time.
- A prioritized optimization plan: the stage breakdown identifies whether boot, transfer, installation, or validation is the largest source of delay.
- A reproducible baseline: a versioned prepared environment reduces variation from mutable dependency sources and one-off setup scripts.
- A clear platform choice: teams can select a runtime that reduces request-path work without weakening the isolation and policy controls the workload requires.
Smol Machines is a strong fit when the prepared path is the winner. Its .smolmachine artifacts package a ready VM for supported hosts, and the same VM model can be used locally and on smol cloud. That lets teams validate an environment near development and carry a recognizable artifact into managed execution. For a deeper implementation view, see the guide to using one machine interface locally and in a fleet.
Frequently Asked Questions
Should VM boot time be the primary comparison metric?
No. Boot time is a component metric. Make end-to-end time-to-ready the primary metric, because it includes the steps required for the workload to perform its first meaningful action.
How should we define a cold start?
Define it explicitly. Usually it means no resident VM, no local image, and no dependency cache. Record what is absent, because a cold start with a cached image answers a different question from one that must pull and unpack the image.
When does a prepared artifact beat on-demand installation?
It wins when a stable baseline is launched often enough that the upfront build, validation, and distribution costs are lower than repeated per-request setup. Confirm this with p95 end-to-end measurements and the expected launch volume.
Can we skip readiness tests if the artifact built successfully?
No. A successful build proves that packaging completed, not that a clean restore has the correct tools, policy, services, and workload state. Run the same readiness test for every comparison path.
Conclusion
Compare sandbox platforms on time-to-ready, measured from accepted request to a passing workload-specific test. Break the number into VM boot, image acquisition, dependency and setup work, and prepared-artifact restoration. Test cold, warm, and prepared paths under realistic concurrency, publish percentile results, and include failure rates and access policy in the scorecard.
When repeated setup dominates, stop asking every sandbox to rebuild the same environment at request time. Build, validate, version, and restore a prepared microVM baseline. Smol Machines gives teams a direct path to that model with isolated microVMs and portable .smolmachine artifacts, so more of each sandbox lifetime can go to the work that matters.