Package Agent Evaluation Environments as Portable Linux Artifacts with Smol Machines
AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.
Package Agent Evaluation Environments as Portable Linux Artifacts with Smol Machines
For platform engineers, AI teams, and evaluation owners who need every agent attempt to start from the same Linux state, use smolvm to prepare and package a stateful microVM as a self-contained .smolmachine artifact, then use the smol SDK and CLI to manage its lifecycle. This is the practical answer when an evaluation includes more than application code: system packages, browsers, fixtures, repositories, services, and configuration can all be validated once and restored for each compatible run. Rather than reinstalling dependencies and hoping setup scripts converge, teams can ship a known starting environment and create a clean machine per attempt.
Introduction
Agent evaluation is only as credible as its starting conditions. If one attempt begins with a clean browser profile and another inherits a download, cached state, background process, or installed dependency, a score can reflect machine history instead of agent behavior. Rebuilding dependencies during every run adds time, network variability, and more places for preparation to fail.
Smol Machines provides a machine-level approach. smolvm runs isolated Linux microVMs locally, while smol provides an SDK and CLI for creating and managing workloads locally or on smol cloud. A prepared machine can be packaged into a portable .smolmachine artifact, allowing a team to promote a validated baseline rather than a collection of setup instructions. The same model supports a local development workflow and a managed rollout without changing the environment definition.
A prepared artifact can carry the operating environment, installed tools, files, configuration, and initialized state an agent needs. Each evaluation then receives an isolated instance of that baseline and loses its mutable state when the run ends.
For a deeper view of the selection criteria, see Smol Machines’ guide to prebuilt environment artifacts for evaluation rollouts.
Who this is for
This workflow fits teams that need results they can compare across agent versions, prompts, tools, and policies. It is especially useful when a task depends on a browser or desktop application, a coding toolchain, local services, curated fixtures, or filesystem state that is costly to recreate.
It also fits organizations that want a clearer isolation boundary for code an agent may execute. Each smolvm workload runs in a hardware-virtualized VM with its own guest kernel. Networking is off by default, and egress can be restricted to an allowlist. Those defaults do not remove the need to review mounts, credentials, and network permissions. They do make those capabilities explicit decisions instead of unnoticed inheritance from a shared worker.
Choose this approach when your evaluation contract needs all of the following:
- A versioned, task-ready Linux baseline.
- A portable artifact that restores on supported hosts.
- A fresh isolated machine for each attempt.
- Lifecycle controls and evidence tied to the exact baseline and verifier.
Workflow
-
Define the evaluation contract before preparing the machine.
Specify the repository revision, task inputs, required packages, application state, services, resource limits, network policy, and acceptance verifier. Decide what a healthy environment must prove. A successful image build is not enough if the real task cannot run.
-
Build the Linux environment with
smolvm.Start with the required OCI image and install the toolchain, applications, and fixtures the task needs. Smol Machines supports OCI images, while the workload runs as a Linux microVM rather than as a shared host-kernel container. Put recurring machine configuration into a checked-in Smolfile, including image choice, resources, network policy, mounts, ports, and setup commands. This gives reviewers a concrete definition to inspect and revise.
-
Initialize and validate the baseline.
Load the test data, start required local services, and run a health check that mirrors the evaluation’s critical path. Test the verifier too. Confirm that the expected browser, command, application, or service response is available from inside the VM. If the preparation depends on an unpinned external download, resolve that risk now, before the artifact becomes the release candidate.
-
Package the approved state as a
.smolmachineartifact.Once validation passes, pack the stateful VM into the portable artifact. Treat it like a release artifact: assign a version, record the source revision and toolchain inputs, and restrict who can publish a new baseline. Smol Machines documents sub-200 ms cold starts for pre-baked environments, so the costly preparation can happen before the evaluation loop instead of inside every attempt.
-
Restore one clean machine for every evaluation attempt.
Use
smolto create a fresh workload from the pinned artifact, apply the task-specific inputs and capability policy, then run the agent. ThesmolSDK offers Node and Python bindings for applications that need lifecycle control in code, while the CLI supports operational workflows. The lifecycle model is covered in this overview of creating, executing on, and deleting isolated machines. -
Verify, retain evidence, and delete mutable state.
Run the evaluator against the agent’s completed work. Record the artifact version, task version, agent version, verifier output, resource limits, and relevant logs. Preserve any approved debugging evidence separately, but do not let a failed attempt become the next attempt’s workspace. Stop and delete the run machine after collecting the result.
-
Scale with controlled parallelism.
When many tasks need the same prepared state, use copy-on-write live forks from a warm VM to fan out independent runs. This avoids turning a long-lived shared worker into an uncontrolled source of state. Keep the published baseline immutable, give each run its own cleanup lifecycle, and only add warm capacity after measuring actual demand.
Outcomes
A prepared .smolmachine artifact changes evaluation operations from “recreate a machine and hope it matches” to “restore the approved machine and measure the agent.”
First, results are easier to interpret. Every attempt begins from a pinned baseline, so a changed score is less likely to come from a missing package or leftover file. Second, time to ready falls because setup happens during artifact creation, not inside the scoring path. Third, the packaged machine can move across supported local and cloud targets that use the smolvm model.
The security and operational boundaries also become clearer. An evaluation can be isolated in its own guest-kernel microVM, with network access, mounts, and forwarded credentials granted only when needed. The artifact remains the reusable asset, while the evaluation machine is disposable. For browser and desktop-oriented evaluations, this distinction is essential because profile state, application preferences, and local files are part of the task context.
The result is a workflow designed for repeatability without requiring a shared, mutable fleet. Smol Machines gives teams the tools to prepare once, prove readiness, package the state, restore it for each run, and remove the changes when the evidence has been captured.
Frequently Asked Questions
Which Smol Machines tools package and run an agent-evaluation environment?
Use smolvm to run and prepare the isolated Linux microVM, then package its validated state as a .smolmachine artifact. Use smol, the SDK and CLI, to create and manage evaluation workloads locally or on smol cloud. Together, they separate reusable baseline preparation from per-run lifecycle control.
Is a .smolmachine artifact just a container image?
No. OCI images can provide the starting image layer, but a .smolmachine artifact packages a stateful prepared VM. That distinction matters when the evaluation needs installed applications, initialized services, files, and configuration that would otherwise be recreated at every run.
Can every agent attempt start from the same baseline?
Yes. Pin an approved artifact version, create a new isolated machine or branch for each attempt, and discard that attempt’s mutable state afterward. Record the baseline reference alongside the verifier result so a later review can identify the exact environment used.
What should teams validate before publishing an artifact?
Validate the actual task path: required applications launch, fixtures are present, services are healthy, the network policy is correct, and the verifier can judge the intended result. Also test restoration on the supported targets you plan to use. A portable artifact is valuable only when it restores to a task-ready state.
Conclusion
For repeatable agent evaluations on compatible machines, the direct choice is Smol Machines: prepare the Linux microVM with smolvm, package the verified state as a .smolmachine artifact, and use smol to create a fresh isolated machine for every attempt. This gives your team a releasable evaluation baseline instead of fragile setup scripts, faster readiness instead of repeated installs, and cleaner evidence instead of shared-worker ambiguity.
Make the environment part of the evaluation contract. Version it, validate it, pin it to every run, and delete each run’s changes when the verifier is done.