Back to blog posts

11 min

Sandbox snapshots: hibernate and resume agents fast

Learn how sandbox snapshots capture filesystem, memory, and process state so AI agents resume in milliseconds instead of rebuilding from scratch every session.

Nicolas LecomteNico is a founder of Blaxel, who usually writes about AI, agentics, and the future of AI runtimes.

Your agent finishes a task. It cloned a repository, installed dependencies, ran an analysis, and wrote results to disk. Then the user closes the session.

On many sandbox platforms, everything in that environment gets deleted. The next session starts from a blank base image. The agent clones the repo again and reinstalls the same packages. It repeats minutes of setup it already ran once.

A snapshot's contents determine what returns on resume, while microVM hibernation determines how execution restarts. Captured data volume and storage location drive resume times from milliseconds to seconds. Snapshot expiration policies determine whether an agent can retain its environment across long idle periods.

Hibernation with a snapshot removes that repeated cost. Instead of deleting the environment, the platform captures its state. The sandbox suspends and stops consuming compute.

When the user returns, the platform restores the snapshot. The agent continues in the same working directory. Mid-execution processes resume at the instruction where they stopped.

TL;DR

  • Snapshots preserve session continuity: Agents can resume completed setup and in-progress work instead of rebuilding their environments.
  • Resume restores the saved environment: Prior setup returns with the snapshot.
  • Speed depends on captured data and storage location: Thin local diffs restore faster than large snapshots from object storage.
  • Perpetual standby avoids expiration rebuilds: Fixed retention windows eventually force agents to start again from the base image.
  • Durable data belongs on a volume: Data that must survive sandbox deletion requires persistent storage.

What a sandbox snapshot captures

Snapshot state includes the writable filesystem layer, guest memory, and each process's execution state. Whatever your platform captures comes back on resume. The agent must rebuild anything the snapshot leaves out.

Three components of sandbox state

The FaaSnap microVM snapshot model (EuroSys 2022) uses two files. One stores "the state of the VM like virtual devices and CPU registers." The other is a guest physical memory file mapped to disk.

The writable filesystem layer comes first. It holds every file the agent created or modified. Cloned repos and installed packages live there. Build output does too. In a layered architecture, the writable layer is a tmpfs mount that lives in RAM. It therefore travels inside the memory file. The read-only base image is a separate artifact and stays out of the snapshot.

Guest memory is the second component. It covers heap allocations, loaded libraries, variable values, and the agent's in-memory data structures. The memory file is a byte-for-byte copy of guest RAM at the moment of hibernation.

At the microVM level, process state is captured implicitly. Program counters and register contents for each virtual CPU (vCPU) go into the state file. Open file descriptors and signal handlers are kernel data structures. They live in guest kernel memory, which is part of the memory file.

On resume, no process restarts. Each process continues from the instruction it was executing. Together, the two files let an agent resume as if no time had passed.

What's excluded from the snapshot and why it matters

The read-only base image is excluded because it never changes between sessions. The platform already stores it on the host. The guest kernel maps program blocks from it on demand. Copying unchanged operating system files into every snapshot would multiply snapshot size for no gain.

Network connections don't survive. AWS Lambda SnapStart uses the same snapshot mechanism. Connection state from before the snapshot "isn't guaranteed" after restore. That warning appears in the SnapStart best practices. Write agent code that reconnects in a post-resume hook. Treat every resume as a fresh connection boundary, and never bind outbound connections to fixed source ports. This avoids bugs that only appear after the first hibernation.

Cached universally unique identifiers (UUIDs) and session tokens resume unchanged. Random number generator (RNG) state does too. Create them after initialization in accordance with AWS's uniqueness guidance. Re-mount external filesystems on startup too. Do not assume they persisted.

Check the guest clock after resume. Do this before the agent checks token expiry or writes timestamped logs. Refresh credentials after resume too. A sandbox may wake weeks after it slept.

How to understand the snapshot mechanism

Hibernation and resume reverse each other's sequence. During hibernation, execution freezes while state is serialized; compute is then released. Resume allocates compute before mapping the saved state and unfreezing execution.

Trace the hibernation sequence

An inactivity timeout or explicit API call can trigger hibernation. Closing the user session can also trigger it. The platform then runs this sequence:

  1. The platform pauses the vCPUs. Register contents and emulated device state are written to the state file.
  2. Guest memory, including the tmpfs writable layer, is serialized to the memory file.
  3. Both files go to the snapshot store.
  4. The platform releases the microVM's CPU and memory. The sandbox is no longer running.

Memory size drives the time cost. A sandbox holding gigabytes takes longer to serialize. One with a few hundred megabytes of changes finishes faster.

Dirty-page tracking reduces the work. A diff snapshot writes only pages modified since the previous snapshot. Sabre (OSDI 2024) compressed diff snapshots two and a half times on average. Smaller diffs leave less data for a restore to read. An agent with a small in-memory footprint pays less at every hibernation and resume.

Reverse the sequence to resume

Resume inverts the sequence:

  1. The platform allocates a fresh microVM.
  2. The read-only base image already on the host is attached.
  3. Loading the state file restores vCPU registers and device state.
  4. The memory file is mapped into the guest's address space.
  5. Unpausing the vCPUs lets processes continue from their saved instruction.

Restoring the state file is cheap. Brooker et al. measured microVM state restoration as low as four milliseconds in favorable conditions. Memory costs more because platforms map the memory file lazily instead of copying it. Each guest access to a page outside host memory triggers a page fault. The faulting vCPU stalls until that page arrives from storage.

After snapshot resume, page-fault servicing consumed 95% of function processing time in REAP (ASPLOS 2021). REAP also cut hello-world processing from 182ms with vanilla serial page faults to 15ms. It records the working set once and prefetches it in a single sequential read. The technique eliminates 97% of page faults on average. Working-set prefetch cuts restore latency most for working sets that are small relative to total memory.

Why resume times vary so dramatically across platforms

Snapshot size and storage location explain most differences between millisecond and multi-second resumes. Both are architectural choices made long before your agent runs.

Snapshot size is the primary variable

A snapshot containing only the memory-resident writable diff is smaller than one holding the whole base image in RAM. Fewer pages mean fewer faults to service. They also mean less data to pull from storage. Layered filesystem architecture makes the thin diff possible.

Reads go to an Enhanced Read-Only File System (EROFS) base. Writes go to a tmpfs upper layer through overlayfs. Only that layer grows the snapshotted memory footprint. Blaxel's engineering team moved a Next.js base image from in-memory initramfs to this layout. Memory usage dropped from four GB to one GB, a 75% reduction. A smaller memory footprint leaves less data to fault in on resume.

Without that layering, memory snapshots grow large. Memory files in the FaaSnap paper range from "a few hundred MB to a few GB." A layout that loads the whole base image into RAM carries those gigabytes into every snapshot. That happens even when the agent changed nothing. A diff layout keeps the unchanged base image out of the snapshot. That keeps resume pages close to the agent's real working set.

Snapshot storage location affects latency

Where the snapshot lives sets the cost of every page fault during resume. AWS documents S3 Standard first-byte latency at roughly 100–200ms. An NVMe drive served reads in 75 microseconds in a USENIX FAST 2026 study. Object storage is slower than local flash by orders of magnitude.

Under lazy loading, that gap is paid per fault, not once. AWS Lambda SnapStart shows the effect in production. Lambda divides snapshots into chunks and serves them from a tiered cache. Chunks available on the worker arrive much faster than chunks fetched from S3. Lambda also records which chunks each function touches. It prefetches them before execution starts. This reduces the storage delays that would otherwise extend resume time.

A memory snapshot restored from object storage adds a network round trip on every cache miss. A resume that touches thousands of pages can't hide that. Local flash brings each fault down to microseconds. With local flash, working-set prefetching can produce millisecond resumes.

Platform benchmarks and what they mean

Academic measurements show how wide the range can be. In FaaSnap, hello-world ran at four milliseconds on a memory-resident VM. Lazy snapshot restore took over 200ms. By my arithmetic, the approaches differed by roughly 50× in that measurement. That gap makes production workload benchmarks essential.

Blaxel reports that sandboxes resume from standby in under 25ms. Filesystem and memory state remain intact. Treat that figure, and every vendor figure, as a reference point. Hibernate a sandbox running your real dependencies. Include the long-running process your agent relies on. On resume, time a request to that process. Repeat until you can read a p50 and a p99.

Include resumes after long idle gaps. A host cache may have evicted an unused snapshot. Multi-gigabyte working sets can expose tail latency hidden by hello-world tests. Ask each provider whether it prefetches the working set. Page faults dominate resume without it.

Hibernation policies, snapshot expiration, and the perpetual standby distinction

What happens to a snapshot over time separates providers more than any millisecond figure. A fast platform can still force full rebuilds by deleting snapshots on a fixed schedule.

Why snapshot expiration breaks agent memory

Some sandbox platforms expire hibernated snapshots after a fixed window and then delete them. The next session starts from the base image. It has no record of what the agent did. AWS Lambda applies the same kind of policy to SnapStart. Java snapshots are deleted after 14 days without an invocation.

When the clock runs out, expiration forces agents that carry state across sessions to rebuild it: a coding agent loses its cloned repo and installed toolchain. A pull request (PR) review agent must recreate the warmed test environment built for a customer's monorepo, while the curated dataset that a data analysis agent loaded into memory last week disappears too.

Some teams keep sandboxes active continuously and pay for idle compute. Others avoid that idle expense by accepting the rebuild cost and adding minutes of latency after each expiration, or by building a state-management layer outside the sandbox. The latter stops relying on the snapshot. Every option imposes a cost. That may be financial or show up as latency and unplanned engineering time. Audit your provider's expiration window before a customer's agent loses context in production.

Perpetual standby as an architectural choice

With perpetual standby, snapshots never expire. A hibernated sandbox stays resumable until the team deletes it. It consumes no compute while waiting. By contrast, some providers delete filesystem state when a sandbox pauses. Others expire hibernated sandboxes after a window measured in days.

Blaxel sets no expiration policy by default. An agent can hibernate after a session. It can resume weeks later with its repo, packages, and processes intact.

Indefinite standby provides session continuity; a Volume provides guaranteed long-term retention in one sandbox.

Per the Volumes documentation, "Files written to the mount path persist even if the sandbox is deleted."

In private preview, Agent Drive provides shared storage for context and files across sandboxes and sessions. Volumes instead provide block storage for raw I/O performance. They also support long-term persistence in a single sandbox. Blaxel's storage guide covers the storage options.

Design agent workloads around perpetual hibernation

Snapshots let environment setup run once. A coding or PR review agent can later resume with its repo cloned and toolchain warm.

Teams that pick a short expiration window often discover the constraint late. A customer's agent loses its context after the window closes. Alternatively, a nightly rebuild adds minutes to every morning session. Either way, the team pays again for setup the agent already finished. Start building at app.blaxel.ai or talk to the team at blaxel.ai/contact.

Blaxel, the perpetual sandbox platform, keeps hibernated Blaxel Sandboxes resumable indefinitely. After approximately 15 seconds of network inactivity, a sandbox moves to standby. It then stops consuming compute. Teams therefore pay compute costs only while the sandbox is active. It resumes with its prior state intact. Volumes provide the durable layer for data that must survive sandbox deletion.

FAQ

What is a sandbox snapshot?

Use a snapshot when the same agent will return to an environment and benefit from preserving installed packages, repository changes, memory, or running processes. It is best suited to session continuity, not guaranteed long-term retention. Keep durable data on a Volume, and prepare post-resume hooks to reconnect networks, re-mount external filesystems, refresh credentials, and check the guest clock.

Why does resume time vary between providers?

Compare providers with the workload users will actually resume, not a hello-world result. Measure the first useful request to the agent's long-running process, then review p50 and p99 results after both short and long idle periods. This test reveals the practical effects of snapshot size, storage distance, host-cache eviction, lazy page faults, and working-set prefetching.

What's the difference between hibernation and deletion?

Hibernation is appropriate when the environment itself remains valuable and should continue later. Its processes resume with the filesystem and memory intact. Deletion is cleanup: the sandbox and snapshot are permanently removed. Before deleting, place required single-sandbox data on a Volume or shared cross-session data in Agent Drive. Recreating the sandbox afterward means restarting processes rather than resuming their saved execution state.

How long does it take for a sandbox to enter standby?

Separate the inactivity trigger from snapshot serialization when evaluating standby timing. A platform may detect inactivity quickly, but it still must pause execution, write CPU and device state, serialize guest memory, and store the snapshot before standby is complete. Larger memory footprints generally take longer. If timing matters operationally, test the transition with the agent's real memory usage rather than an empty sandbox.

Related articles