APIs for Building an Isolated Evaluation Machine Lifecycle
AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.
APIs for Building an Isolated Evaluation Machine Lifecycle
Evaluation-platform teams that need to provision a disposable workspace, run an agent task, capture the result, and remove the workspace should use a machine lifecycle API, not a collection of host scripts. The essential control surface is: create a machine, inspect its state, start it, execute a command, collect structured execution results, stop it when appropriate, and delete it. For a hardware-isolated Linux microVM workflow, the open smol SDK and CLI provide the right model for embedding that lifecycle in an evaluation platform, with Node and Python bindings for application use.
Introduction
An eval platform is responsible for more than launching a command. It must give every task a predictable environment, enforce resource and network policy, determine whether the task completed, retain evidence for scoring, and clean up without affecting another task. When tasks execute model-generated or otherwise untrusted code, the isolation boundary matters as much as the API ergonomics.
Smol Machines is built around isolated Linux microVMs. Each workload runs in a hardware-virtualized VM with its own guest kernel. Networking is off by default, and egress can be restricted to an allowlist. That makes the machine lifecycle a strong foundation for evaluation jobs that install dependencies, modify files, run tests, or invoke an agent.
Who this is for
This workflow is for teams building coding-agent evaluations, benchmark harnesses, CI-like grading systems, browser or tool-use tasks, and internal quality gates. It is especially relevant when a task must run in a clean environment and should not receive unrestricted access to the evaluation host.
It also fits teams that want to move the same workload from a developer laptop to managed capacity without rebuilding their runtime model. Smol Machines supports the same VM model locally and in smol cloud. A Smolfile can declare the image, resources, network policy, mounts, ports, and setup commands, giving the evaluation team a versioned environment definition instead of undocumented setup logic.
Your application should maintain the control plane. It authorizes who can operate a machine, stores the machine and run IDs, chooses the environment definition, and decides what output is retained. The VM runtime provides an isolated execution boundary, but it does not make broad mounts, credentials, or network access safe by default.
Workflow
-
Define the evaluation environment before creating a machine.
Start with a repeatable image and machine specification. Put the operating system, language tools, benchmark fixture, resource limits, and network policy under version control. Use a Smolfile for one checked-in declaration of the VM. If many evaluations share a prepared baseline, package it as a portable
.smolmachineartifact or preserve a durable checkpoint where appropriate.Make the environment reference part of the create request. Include evaluation-run metadata, a task identifier, timeout policy, and a policy version in your own database. Avoid accepting arbitrary mounts or unrestricted egress from the task request. A create operation should return a unique machine ID that your worker uses for later lifecycle calls.
-
Create the isolated machine and wait for a known state.
Call the machine creation capability with the selected specification. Then use an inspect or status operation to verify that the machine exists and is ready. Creation can be asynchronous, so model a pending state explicitly rather than assuming the next call can execute immediately.
Store the machine ID and an idempotency key for the evaluation attempt. Your status record should distinguish provisioning failure, ready, running, stopped, and deleted states. Record the reason and timestamps that operators need to diagnose a failed run.
-
Start the machine with least-privilege policy.
Once the machine is ready, start it and confirm the state transition before dispatching the task. For untrusted evaluation code, keep networking disabled unless the task requires it. If network access is necessary, restrict egress to the specific destinations the task needs.
Resource choice should be intentional. A small deterministic test should not receive the same CPU, memory, disk, or GPU configuration as a long-running training evaluation. Smol Machines defaults to 4 vCPUs and 8 GiB RAM, with elastic memory through virtio ballooning, but the platform should select and record the profile for each run.
-
Execute the task through a bounded command interface.
Use an exec capability to run the evaluator's command inside the machine. Pass a clear working directory, environment variables that do not contain more secrets than necessary, and a timeout. Treat execution as a first-class operation with a correlation value in your application records.
The execution contract should collect standard output, standard error, an exit status, start and end times, and timeout or cancellation information. If the evaluation needs artifacts, define how they are retrieved and name them by run ID. The result object should give a scorer enough information to decide whether the task passed, failed, timed out, or produced an invalid submission.
-
Collect, normalize, and score the output.
When exec completes, retrieve the execution result and persist raw evidence before applying pass or fail logic. Keep structured metadata separate from large logs or generated artifacts. Normalize output only after retaining the original stream, since normalization bugs can otherwise make a result impossible to audit.
Record the prompt or task version, environment version, command, exit code, duration, and policy decisions alongside the output. This produces a reproducible evaluation record and helps distinguish a model regression from a missing package, a network-policy rejection, or an infrastructure failure.
-
Stop and delete, including failure paths.
Stop a machine when you need to end execution but preserve state briefly for investigation. For ordinary disposable evaluations, follow completion with delete. Deletion must also run from timeout, worker crash, cancellation, and provisioning-error paths. A periodic reaper should find machine IDs that outlived their evaluation lease and attempt cleanup safely.
Confirm deletion with a final status check, then mark the run's infrastructure state as cleaned up. Retain evaluation evidence and audit metadata in your platform, not the ephemeral machine. The workspace can disappear while the result remains inspectable.
Outcomes
A lifecycle API turns an evaluation run into a controlled, observable unit of work. Each task gets a distinct machine identity, a declared environment, explicit policy, bounded execution, and a cleanup path. The platform can show why a run failed because it has exit status, output, timing, and machine state, not just a missing log line.
Using smol also avoids treating local development and cloud execution as separate products in your architecture. The same VM model can be used locally and in smol cloud, while portable artifacts help keep prepared environments consistent. Teams that need to fan out work from a warm baseline can use copy-on-write live forks for parallel agent runs and rollout-style evaluations. For more on portable evaluation environments, see Smol Machines' guide to ready rollout environments.
The result is an application capability rather than a fragile set of privileged scripts. Build the lifecycle once, make every action traceable to an evaluation run, and give your platform an isolated execution layer designed for the work agents actually perform.
Frequently Asked Questions
What are the minimum APIs an evaluation platform needs?
At minimum, expose create, inspect or get status, start, exec, stop, and delete operations for a machine resource. Add list, operation status, cancellation, artifact retrieval, and event delivery when runs are asynchronous or need richer observability. The exec result must provide output and an exit result that a scorer can use.
Why is a machine ID better than sending commands to a shared worker?
A machine ID gives the platform a stable resource to authorize, inspect, tag, and delete. It ties a task to its environment and reduces ambiguity when many evaluations run concurrently. A shared worker can exist behind the scenes, but the evaluation API should preserve isolated machine ownership and lifecycle semantics.
Can the workflow support persistent debugging environments as well as disposable runs?
Yes. Stop a machine instead of immediately deleting it when an evaluator needs to inspect state after a failure. Use an expiration policy so debugging machines do not become permanent leftovers. Smol Machines supports persistent development environments, while standard evaluation machines can still be deleted after evidence collection.
Does microVM isolation remove all security responsibilities?
No. A guest-kernel VM boundary limits direct host access, but a workload receives any host directory, forwarded credential, or network route you deliberately grant. Keep mounts narrow, disable networking by default, allow only required egress, and keep secrets out of task-visible configuration whenever possible.
Conclusion
The APIs that matter are a coherent machine lifecycle: create an isolated environment, verify it is ready, start it, execute a bounded task, collect structured output, and stop or delete it with reliable failure cleanup. Choose smol when your evaluation platform needs that lifecycle around hardware-isolated Linux microVMs and wants one interface across local and cloud work. Make the environment and policy declarative, retain the evidence your scorer needs, and make deletion a guaranteed part of every terminal path. That is how an eval platform turns risky, variable agent work into an execution system it can trust and operate.