← Engineering
Engineering · Checkpoints

Giving AI an undo button

You can undo a typo, a deleted email or a bad crop. An AI agent working on a real computer can't undo anything. When it drops the wrong table or spends forty steps heading the wrong way, the only fix is to throw the machine away and lose everything it built. We think agents need the same undo button everyone else gets, so smol machines now gives an agent's whole computer one.

Why agents can't undo today

Undo in an editor works because the editor owns all the state. An agent's state is spread across a whole computer: a half-built repo, a running dev server, a database, a logged-in browser, files in /tmp. Each of the usual tools covers only a slice of that.

ToolWhat it can undoWhat it loses
gitTracked files in one repoProcesses, databases, installed packages, everything outside the repo
Restart the containerEverything, by going back to the imageAll the work since the image, including running processes
VM memory snapshotMemory and CPU stateDisks, unless you manage them separately
smol machines checkpointMemory, CPU state, running processes and disks, togetherNothing on the machine

The last row is the point. Memory and disk have to be saved together, because undoing one without the other leaves a process holding a file handle to a file that no longer matches.

Saving and undoing

You start the machine with --branchable and checkpoint it before anything risky. The machine only pauses while its CPU state and disk boundary are captured, then keeps running while its memory is written out. On a MacBook, a small Alpine machine took 2.7 seconds to checkpoint and was paused for 71 milliseconds of that.

save a point to come back to
smolvm machine start --name agent --branchable

# save before anything risky
smolvm machine checkpoint --name agent --store ./history -o ./before-migrate.checkpoint

Undoing means creating a machine from that checkpoint. Its processes pick up exactly where they were, with files that match. The broken machine stays around until you delete it, so you can still look at what went wrong.

undo
# the migration went wrong: undo it
smolvm machine stop   --name agent
smolvm machine create --name agent-undo --from ./before-migrate.checkpoint
smolvm machine start  --name agent-undo

This is fast enough to use without thinking about it. Our rewind demo runs a full 2 GiB Linux desktop, and jumping it back to a save point takes about 4.5 seconds end to end: 1.3 seconds to restore, 1.8 to start and half a second until the desktop is back on screen. On Linux the restore uses the checkpoint's disk as a shared read-only base instead of copying it, which cut that first step from 1.3 seconds to 0.26 on our test host.

In an agent loop, this means saving before each step and undoing any step that fails, rather than starting over:

an agent loop with undo
while read -r step; do
  smolvm machine checkpoint --name agent --store ./history -o "./before-$step.checkpoint"

  if ! smolvm machine exec --name agent -- ./do-step.sh "$step"; then
    # undo: back to the moment before this step, processes and all
    smolvm machine stop   --name agent
    smolvm machine delete --name agent -f
    smolvm machine create --name agent --from "./before-$step.checkpoint"
    smolvm machine start  --name agent --branchable
  fi
done < plan.txt

Going back further

One step of undo is rarely enough, because the mistake is often ten steps back. So each checkpoint also records its parent: the checkpoint the machine was last saved to or restored from. Save a few times and you have a history. Go back to an old point and keep working, and the history branches instead of being overwritten, much like checking out an old git commit and committing on top of it.

Checkpoint lineage · click a checkpoint, then restore it
9e37793c6ef3daa66d
current position of worker
last command
smolvm machine checkpoint --name worker --store ./ckpt -o ./ckpt-3.checkpoint
smolvm machine checkpoint-log ./ckpt-3.checkpoint
~0   daa66d13daa5  worker  (this checkpoint)
~1   3c6ef3623c6e  worker
~2   9e3779b19e37  worker
see the history, go back two steps
smolvm machine checkpoint-log ./ckpt-3.checkpoint
# ~0   c4c23a7c6d3a  2026-09-28T03:01:05Z  worker  401 MiB data  (this checkpoint)
# ~1   6b3e4f1fde48  2026-09-28T03:00:55Z  worker  401 MiB data
# ~2   58a6bb95baca  2026-09-28T03:00:45Z  worker  401 MiB data

# go back two steps (quote ~2 so zsh doesn't expand it)
smolvm machine create --name rewind --from ./ckpt-3.checkpoint --at '~2'

A history is only useful if cleaning up old checkpoints doesn't break it. Most incremental snapshot formats are chains, where each file depends on the one before, so losing one link breaks every later point. A smol machines checkpoint written with --store doesn't depend on older files. It holds a full index, hard links into a shared content-addressed store, and the indexes of its earlier points, so the newest checkpoint can restore its whole history on its own. Unchanged data is shared, which keeps history cheap: in our demo, three points took 98 MB against 52 MB for one.

You can delete files on both sides here and see what still restores:

Three points in time · delete the older files and see what still restores

Diff chain snapshot per point

  • base.mem
    ABCDEFGH
    full snapshot
  • diff-1.mem
    C'G'
    pages dirtied since base
  • diff-2.mem
    E'H'
    pages dirtied since diff-1
  • rootfs.ext4
    disk state: yours to copy and keep consistent with each point
  1. ✓ point 1 base.mem
  2. ✓ point 2 rebase base.mem → diff-1.mem
  3. ✓ point 3 rebase base.mem → diff-1.mem → diff-2.mem

Stored checkpoints RAM + disk, retained history

  • ./ckpt-1.checkpoint/
    ABCDEFGH
    index + links
  • ./ckpt-2.checkpoint/
    ABC'DEFG'H
    index + links · retains #1
  • ./ckpt-3.checkpoint/
    ABC'DE'FG'H'
    index + links · retains #1, #2
objects/
ABCDEFGHC'G'E'H'
12 unique chunks stored once
  1. ✓ point 1 ./ckpt-3.checkpoint --at '~2'
  2. ✓ point 2 ./ckpt-3.checkpoint --at '~1'
  3. ✓ point 3 ./ckpt-3.checkpoint
Try deleting base.mem, then ./ckpt-1.checkpoint/ and ./ckpt-2.checkpoint/.

Trying two paths at once

Sometimes you don't want to undo at all. You want to try two next steps from the same moment and keep whichever works. That's a branch. The engine pauses the source, freezes its memory and disks into a shared copy-on-write base and resumes it, and each child starts from that moment and only stores what it changes. On a MacBook the pause was about 80 milliseconds for a 1 GiB machine.

branch
smolvm machine branch --from agent --name try-a
smolvm machine branch --from agent --name try-b
# run a different next step in each, keep the one that works

Try it: branch the source while it runs, then write from each side.

Live branch · copy-on-write memory
source counter 0

The source machine is running. Branch it while it works.

machines
1
pages held
32
as full copies
32
still shared
0

private page (written after the branch) read through the shared base

Each child is a new machine with its own name, hostname, machine id and ports, so the copies don't collide, and it records which machine it came from.

How this compares with Firecracker

Most microVM platforms build on Firecracker snapshots. Firecracker gives you the primitive, pausing the VM and writing out its memory and device state, and leaves disks, diff chains and what a copy means up to you. In practice every team on Firecracker ends up building its own undo on top.

Firecrackersmol machines
SavesMemory, device stateMemory, device state, disks
Go back to an older pointRebase base + diffs in order--at ~N on any later checkpoint
Delete older filesLater diffs stop restoringNewest checkpoint still restores all
Copy a running VMSnapshot, then restore it N timesmachine branch; source keeps running
Copies get unique identityUp to the guest (VMGenID reseeds RNG)New name, hostname, machine id, ports
HostsLinux (KVM)Linux (KVM), macOS

Firecracker does win on capture cost. Its diff snapshots write only the memory pages that changed, straight from KVM, while our checkpoints still read all of memory and only skip writing the chunks that haven't changed. If you save every few seconds, dirty-page capture is the cheaper approach.

There are a few limits to know about. A checkpoint restores on a host with the same OS and CPU architecture, and neither approach can move a running machine between architectures. History keeps 32 points by default (--history N changes that), and live branch chains go 32 levels deep. Windows hosts don't support checkpoints or branching yet. Everything here works in smolvm 1.19.1 and later.


Get the runtime on GitHub, or read how checkpoint history works.