Back to blog posts

13 min

Persistent storage for AI agents: choosing between in-memory, volumes, and shared filesystems

Compare AI agent storage options: in-memory tmpfs, persistent volumes, and shared filesystems. Use a three-question test to match each data type to the right one.

Nicolas LecomteNico is a founder of Blaxel, who usually writes about AI, agentics, and the future of AI runtimes.

During execution, the agent writes intermediate results and reads them back repeatedly. At creation, every new sandbox needs the same model weights or reference dataset already in place. It should not download them again. When several agents work the same job, they need one work-in-progress dataset without copying it between environments.

The sandbox's tmpfs writable layer keeps working data in RAM for one session. Volumes keep single-sandbox data past deletion. They can also hold data pre-loaded before the agent starts. Depending on the implementation, a shared filesystem may let several sandboxes access the same data, but concurrent read-write access is not guaranteed; some platforms allow it while others limit write access or provide read-only snapshots.

The choices have different performance and cost profiles, and data does not persist equally across them. The wrong choice adds latency and can cause data-loss or synchronization problems.

TL;DR

  • In-memory storage is fastest but ephemeral: tmpfs is a RAM‑backed filesystem that operates at roughly RAM speed, but its contents are volatile and are lost on unmount, reboot, or when the tmpfs instance is removed.
  • Volumes persist single-sandbox data: Block storage attaches at creation and survives deletion. This suits pre-loaded weights and lasting outputs.
  • Shared filesystems support multi-agent work: Several sandboxes mount the same data without copying it.
  • Most pipelines combine two or three: Compute in tmpfs and load dependencies from a volume. Hand results through a shared filesystem.

What storage requirements do agent workloads have?

Web-service storage guidance doesn't fit agents. Agents read, write, and share data differently.

Compare agent access patterns across three storage dimensions

A web service reads and writes small, structured records. Examples include database rows and JSON payloads. Each operation has a tight latency target. An agent reads and writes large, unstructured artifacts. These include cloned repositories, generated files, datasets, model outputs, and tool call results. The latency profile differs too.

During execution, the agent should never stall on file input/output (I/O). Each stall lengthens a response someone is waiting on. Between sessions, durability requirements vary by agent. A pull request (PR) review agent can rebuild its state from the repository. Losing its scratch files therefore costs little.

A report generation agent has failed if its output disappears with the sandbox. Concurrency also depends on architecture. Independent agents never touch each other's data. Planner and executor agents in one pipeline touch the same data constantly.

Storage decisions depend first on access speed, which describes how quickly the agent reads and writes during execution. Persistence concerns how long the data survives; shareability matters when multiple sandboxes need simultaneous access. Persistent data may survive hibernation or sandbox deletion. It may also remain indefinitely.

No single primitive maximizes all three dimensions. In-memory storage wins on speed and loses on persistence. Block storage attached to one sandbox wins on persistence. It generally suits large sequential reads and writes. Shared network filesystems win on shareability but add overhead to every operation. The decision framework matches each workload to its dominant requirement.

In-memory storage: tmpfs and the sandbox writable layer

The fastest storage option sits inside the sandbox itself, trading durability for speed.

Understand how in-memory storage works in a sandbox

Pick in-memory storage when data lives and dies with one session. A sandbox filesystem stacks a writable RAM layer on a read-only base. The base image sits on host storage in the Extendable Read-Only File System (EROFS) format. A tmpfs writable layer on top lives entirely in the sandbox's RAM. OverlayFS directs reads to the EROFS base and writes to tmpfs. This layout keeps active files close to the process while preserving the read-only base image beneath them. It is designed for session-local work rather than guaranteed long-term retention.

  • Consequently, writes go to RAM rather than host disk. The engineering post on sandbox memory usage covers this layout in detail.
  • Hibernation preserves this data. When a sandbox loses its active connection, the platform snapshots its entire state.
  • That state includes the in-memory filesystem and running processes. The sandbox then moves to standby. On resume, the snapshot restores.
  • The files return where the agent left them. Deleting the sandbox erases the tmpfs contents immediately.
  • Anything not already written to a volume or shared filesystem is gone.

Choose when to use in-memory storage

Use in-memory storage for data the agent touches repeatedly within one session and can afford to lose. The practical rule applies across coding, PR review, and data analysis agents. Keep the working set in tmpfs and copy only durable outputs elsewhere. Start by listing the files your agent writes during a task. Then mark which files anyone needs afterward.

  • Active working data: This category includes intermediate results, partial outputs, in-progress computations, and tool call outputs. A coding agent's cloned repository, node_modules directory, and test output all belong here.
  • Data that's cheap to recreate: Losing data costs little if the agent can regenerate it in seconds.
  • Frequently accessed data: This category covers files the agent reads or writes repeatedly.
  • Data that exceeds RAM: The dataset may exceed the sandbox's RAM budget. The 16.07 GB Llama-3.1-8B-Instruct checkpoint would consume 16 GB of tmpfs RAM. That usage comes on top of what processes need. Oversized tmpfs can also deadlock the machine. The out-of-memory (OOM) handler cannot free that memory.
  • Data that must survive deletion: Some data must survive sandbox deletion. Blaxel's best-practices docs rate sandbox filesystem durability low. Deletion or a crash can erase it.
  • Data shared across sandboxes: More than one sandbox may need the data simultaneously. A tmpfs layer lives inside one sandbox. No second sandbox can see it without a copy.

Volumes: persistent block storage for single-sandbox workloads

When data needs to outlive the sandbox but stays tied to a single agent, volumes provide durable block storage.

Understand volumes, attachments, and templates

Volumes answer one question: does single-sandbox data need to outlive its sandbox? A volume is a persistent block storage device attached to a sandbox at creation time. You pass the volume name and a mountPath in the sandbox's volumes property. Attaching or detaching after creation isn't supported, per the volume attachment rules. The volume must also live in the sandbox's region.

  • Unlike tmpfs, a volume survives sandbox deletion. It can reattach to a new sandbox with its files intact.
  • Each volume attaches to one sandbox at a time. Concurrent multi-sandbox access is the shared filesystem's job.
  • Volumes use block storage. For data kept weeks or longer, a volume costs less than a standby snapshot.
  • The volumes docs describe this cost distinction.
  • Volume templates are blueprints for creating pre-populated volumes. You build one from a directory of content and deploy it.
  • Every sandbox created with that volume starts with the same content. You can preload weights and reference data. The same applies to dependency caches.
  • Templates remove per-session download cost and latency.
  • Suppose the Llama-3.1-8B-Instruct checkpoint is served from AWS. Egress past the free tier runs $0.09 per GB. That egress charge recurs on every sandbox creation. A template pays it once.
  • Provision 20–30% extra space beyond the template's content. Creation fails if the volume is smaller.
  • Deploy the template first. Then reference it in your sandbox setup's volume creation call.

Choose when to use volumes

Pick a volume when data must outlast the sandbox and only one sandbox needs it at a time. Data analysis agents often follow a two-step pattern. They write to tmpfs during the task. They then copy final artifacts to the mounted volume path as the last step. The working set retains RAM-speed access. The output is also guaranteed to survive.

  • Pre-loaded model weights or large datasets: The agent needs the same files on every sandbox creation. Store them once, and every new sandbox starts with the data in place.
  • Outputs that must survive the sandbox: An agent may generate a report, migration script, or cleaned dataset. It should write that output to a volume. If the sandbox is deleted afterward, the output stays.
  • Long-term retention: This category covers data kept for weeks, months, or indefinitely. Volume data stays until you delete the volume.

Route only durable files to the mounted path. Temporary computation can then continue using faster in-memory storage.

Shared filesystems: concurrent multi-agent data access

When multiple sandboxes need to read and write the same data at once, a shared filesystem removes the copy step from every handoff.

Understand shared filesystems and Agent Drive

Only a shared filesystem lets several sandboxes mount the same data at once. In-memory storage and volumes are both single-sandbox primitives. One sandbox writes, and no other sandbox reads the same bytes without a copy operation. Suppose a planning agent produces artifacts that execution agents consume. Every handoff becomes an upload, download, or transfer through an intermediate bucket.

  • With a shared filesystem, the planning agent writes a task definition to a path. The execution agent reads it from the same path immediately. No transfer step is required.
  • Agent Drive supports concurrent reads and writes. It reaches the sandbox through a Filesystem in Userspace (FUSE) client.
  • A FUSE passthrough filesystem over Ext4 had no throughput loss on sequential solid-state drive (SSD) reads. This result came from a 2017 USENIX File and Storage Technologies (FAST) study. However, file creation lost up to 82.5% of throughput. Plan read-heavy sharing of checkpoints and datasets on the drive. Keep small-file churn in local tmpfs.
  • Blaxel's Agent Drive is a distributed shared filesystem. It mounts to multiple sandboxes simultaneously and uses a POSIX-compliant FUSE client.
  • Drives support concurrent read-write access (RWX) and built-in replication for durability.
  • A drive can attach to an already-running sandbox at any mount path. Volumes cannot do that.
  • Capacity scales automatically with no fixed limits.
  • Agent Drive is available only in us-was-1 during private preview. Both the drive and sandbox must be created there.
  • Access requests and mount syntax are in the private preview docs.

Identify use cases that need a shared filesystem

Multi-agent architectures are the main reason to reach for a shared filesystem. Start with one shared mount for handoff artifacts. Keep each agent's scratch work in its own tmpfs. For collaborative editing, assign each agent its own directory or another application-level boundary. Blaxel's storage guide covers the available storage patterns.

  • Multi-agent pipelines: Planning agents write task definitions, and execution agents read them. Completion agents write results, and an aggregation agent reads everything afterward.
  • Shared dependency caches: A package cache mounted across all agent sandboxes means installs happen once.
  • Shared context histories: Conversation logs, tool call outputs, and prior session results stay in one place. Every agent can read them without replication.
  • Collaborative code editing: Two agents working on the same repository read and write one tree. They do not need to copy it between environments.

Favor read-heavy shared assets and move small-file churn to each sandbox's local tmpfs.

How to choose: a decision framework

Every file your agents produce belongs on one primitive. The following questions identify which primitive fits.

Run the three-question decision test

Answer these three questions in order for each data type your agent produces:

  1. Does the data need to survive sandbox deletion? Yes means a volume or a shared filesystem. No means in-memory storage can be enough.
  2. If persistence is required, does more than one sandbox need the data simultaneously? Yes means a shared filesystem. No means a volume.
  3. Is access speed during execution the primary concern? Yes means in-memory storage, provided the data fits in RAM. Otherwise, use the answer from the first two questions.

Run the test per data type, not per agent. Record each answer beside the file so the routing decision remains explicit. A PR review agent's cloned repository answers no, no, and yes. It therefore stays in tmpfs. Its review report answers yes and no, so it lands on a volume. A planner's task queue answers yes and yes, so it belongs on a shared filesystem. This file-level test prevents one storage choice from governing the entire agent. It also exposes files with conflicting speed, durability, or sharing requirements. Revisit the answers when a file's consumers or retention needs change.

Combine primitives by data type

Production agent architectures often combine primitives. A single agent usually produces several data types. Latency-sensitive computation stays in RAM, while durable or shared files go to storage designed for those requirements. This also avoids forcing every file into one primitive's cost, performance, and concurrency profile.

  • In-memory plus volume: Active working data lives in tmpfs. Long-term outputs go to the volume before sandbox deletion. This pattern fits single-agent coding and data analysis workloads.
  • Volume plus shared filesystem: Pre-loaded model weights sit on a single-sandbox volume created from a template. The shared work queue sits on Agent Drive for multi-agent access.
  • In-memory plus shared filesystem: Fast computation happens in tmpfs. Results then move to the shared filesystem for downstream agents.

For example, a coding agent can keep its cloned repository, dependencies, and test output in tmpfs while writing the completed report to a volume. In a multi-agent pipeline, it can instead place only handoff artifacts on the shared filesystem for downstream agents.

Design the storage layer around each data type's access pattern. Do not apply one model to the whole pipeline. Start with a table. Give each file its own row and each decision question its own column.

Comparison: the three primitives at a glance

The table below summarizes how tmpfs, volumes, and Agent Drive differ across speed, persistence, and sharing.

In-memory (tmpfs)VolumesShared filesystem (Agent Drive)
SpeedRAM-speedBlock storage (slower than RAM)Network filesystem
Persists through hibernationYes (snapshot)YesYes
Persists through sandbox deletionNoYesYes
Multi-sandbox accessNoNoYes
Best forActive working dataPre-loaded data, long-term outputsMulti-agent collaboration

Match storage primitives to access patterns

Matching each data type to its primitive avoids latency and data loss caused by a single storage model. It also prevents synchronization failures. Teams that keep must-survive data in tmpfs lose it on the first deletion. Teams that misuse volumes for shared workloads must copy data during every handoff.

Blaxel is the perpetual sandbox platform for agents that execute code in production, providing the storage primitives covered here. Its sandboxes hold tmpfs state in standby and resume in under 25 milliseconds. Volume templates support persistent single-sandbox storage, with Agent Drive available for shared multi-agent access. Coding agents, PR review agents, and data analysis agents fit this combination most directly.

Start building at app.blaxel.ai or talk to the team at blaxel.ai/contact.

FAQ

Q: What's the difference between sandbox hibernation and persistent storage?

Use hibernation when a session may resume and you need its filesystem, memory, and running processes restored together. Treat it as resumable execution rather than a durability guarantee because deletion or a crash can still remove sandbox state. Put must-survive files on a volume for one-sandbox use or a shared filesystem when several sandboxes need them.

Q: Can I use multiple storage types in one agent pipeline?

Yes. Assign each storage type a distinct mount path and route files according to their lifecycle. Configure a volume when creating its sandbox, because it cannot be attached later. Agent Drive can mount to a sandbox that is already running. Keep temporary computation in tmpfs, durable single-sandbox outputs on the volume, and shared handoff artifacts on the drive.

Q: What happens to in-memory data if the sandbox crashes?

Files that exist only in tmpfs can be lost before the agent finishes. Set a checkpoint policy based on how much work the pipeline can afford to repeat. Copy important intermediate state and final outputs to a volume or shared filesystem at task boundaries, after expensive computations, or before deleting the sandbox. Leave easily regenerated scratch files in memory.

Q: How large can a volume be?

Volume size is set in megabytes at creation. You can increase it afterward with VolumeInstance.update in the SDK, but you cannot decrease it. Size the initial volume for current data plus expected template overhead, then expand it as retention needs grow. Check limits with the team before designing for unusually large capacity or demanding input/output operations per second (IOPS).

Related articles