15 min
How to share files between AI agents working on the same task
Learn how to share files between AI agents using shared filesystems, naming conventions, and concurrency controls for reliable multi-agent pipelines.

You're building a multi-agent pipeline. A planning agent splits a large job into subtasks. It writes each subtask to a file. Five execution agents work on those subtasks in parallel. Each agent needs the planner's output. A review agent then reads all five outputs and checks them for consistency.
With isolated sandboxes and no shared storage, none of that coordination happens automatically. Each agent reads from its own filesystem. The writer must upload a file, and the reader must download it. File-sharing mechanisms differ in how they move artifacts between agents. A shared filesystem lets agents use ordinary file operations, while naming and status conventions prevent write collisions.
That round trip is tolerable for a handful of handoffs. As agent and file counts grow, repeated copy cycles slow the pipeline.
The review agent waits on five downloads before it can start. Every exchange needs a message-passing layer or intermediate API. It also needs a collision-free naming scheme.
Examples use Blaxel's Agent Drive, in private preview. Blaxel Batch Jobs can provide compute for parallel fan-out workloads, while Agent Drive provides shared storage.
TL;DR
- Isolated sandboxes don't share filesystems by default: Each agent's sandbox has its own writable layer. Files remain invisible to other sandboxes until you add a sharing mechanism.
- A shared filesystem is the right primitive for agent collaboration: Several sandboxes mount the same directory tree. One agent's completed writes become available to the others.
- Object storage fits occasional transfers: Tight collaboration loops incur transfer requests and object-key coordination on every handoff.
- Concurrent writes need coordination: Partitioned namespaces prevent writers from targeting the same path. Atomic rename can publish a complete file, but it does not preserve every competing writer's result.
- The shared filesystem is the inter-agent communication layer: Define the directory structure and its naming and status-signaling rules while designing the pipeline. Don't wait for the first collision.
Why isolated sandboxes need an explicit file-sharing mechanism
The isolation that protects agents also keeps their files apart. Four mechanisms bridge the gap.
Sandbox isolation is intentional and creates a coordination gap
Each Blaxel sandbox is isolated from its neighbors and has its own filesystem. The boundary is a security property. If generated code turns hostile, the agent still can't read a neighboring sandbox's files. That isolation protects unrelated workloads and limits the reach of unsafe code. It also means filesystem paths are local unless you explicitly mount shared storage.
The same boundary blocks legitimate sharing. A review agent needing five outputs has no path to them by default. An executor writing result.json to local disk has produced nothing the reviewer can see. The reviewer's run then stalls or reports missing inputs. Parallel execution alone doesn't solve that coordination problem. Blaxel Batch Jobs can run fan-out work, but outputs still need an exchange mechanism. Agent Drive supplies shared storage when those parallel tasks need common files.
The application layer must bridge that gap deliberately. The right mechanism depends on exchange frequency and file size. Pick the mechanism before writing the directory layout and agent code.
Also decide who creates files, who consumes them, and how completion becomes visible. Those choices form the pipeline's file protocol. Making them early prevents local-path assumptions from spreading through every agent.
The options: object storage, databases, message queues, and shared filesystems
Four mechanisms cover nearly every inter-agent file exchange:
- Object storage: Examples include Amazon Simple Storage Service (S3) and Google Cloud Storage (GCS). Agents upload and download files over HTTP. Amazon S3 offers strong read-after-write consistency. Simultaneous writes to one key overwrite earlier writes. Each transfer is a separate request with its own latency. That suits occasional handoffs, not a tight collaboration loop.
- Databases: Agents read and write structured records to a shared database. That works for JSON outputs, metadata, and status rows. It's the wrong home for large binary artifacts, model outputs, or unstructured documents.
- Message queues: Agents pass task definitions and result pointers through a queue. The message carries "here's the path to the file." The file itself still needs somewhere to live.
- Shared filesystem: Multiple sandboxes mount the same directory tree and use plain file I/O. No upload or download step exists. A written file occupies one location for every mount.
For file-centric collaboration, the shared filesystem removes the copy step required by the other three mechanisms. The filesystem works alongside databases and queues. Structured state can remain in a database, and queues can still schedule work. The shared mount handles artifacts that agents must read or modify through ordinary file operations. This separation keeps task delivery distinct from artifact storage.
How a shared filesystem works for multi-agent sandboxes
Mount semantics determine write visibility and how agents can access one file concurrently.
Mount semantics: one filesystem, many sandboxes
A shared filesystem presents one directory tree to many sandboxes. They can mount it simultaneously. A write from one sandbox lands in the storage layer. Every other mount can then read it without a copy. Agents use ordinary paths rather than object keys or transfer APIs. The mount point should remain consistent across the pipeline. Consistent paths prevent each agent from needing a different configuration or translation layer.
Visibility timing depends on the cache model. Linux Network File System (NFS) clients cache directory attributes for up to 60 seconds by default. A new file can remain invisible to other clients during that interval. With attribute caching, Amazon Elastic File System (EFS) can take up to three seconds to show another instance's write. The exact delay depends on NFS client settings. Status conventions must tolerate those visibility gaps.
In the introduction's pipeline, the planner writes subtasks to /tasks. The executors and reviewer read the same tree. Agents should not assume, though, that every filesystem has identical cache behavior. Completion markers, retries, and bounded polling make the protocol resilient to delayed visibility.
Concurrent read-write access and consistency semantics
Concurrent reads do not conflict. Any number of sandboxes can open the same file simultaneously. POSIX (Portable Operating System Interface) does not specify how concurrent writes to a regular file behave. It only promises that each write is atomic. Applications must provide concurrency control. Page-rounded buffered writes can also disappear across non-overlapping regions on NFS.
Most multi-agent pipelines can avoid same-file writes by design. Give each agent its own subdirectory. Coordination then happens by reading other agents' directories. Large distributed filesystems chose the same rule. Their single-writer semantics avoid serializing writes from multiple writers. Partitioning removes the need for a distributed lock manager. It also removes lock leases and related cache races.
Shared read-write access therefore describes mount capability. Safe application behavior requires ownership rules for files and directories. One writer per output is the simplest rule. Other agents can read that output after its completion signal appears. If multiple contributors must create an aggregate, they should produce separate fragments. A coordinator can then combine those fragments in a controlled step.
Agent Drive: Blaxel's shared filesystem primitive (private preview)
Blaxel's Agent Drive is a distributed filesystem in private preview. It is available in us-was-1. Multiple sandboxes can mount one drive simultaneously. It supports concurrent read-write access and includes storage-layer replication for durability. Both the drive and each mounting sandbox must exist in us-was-1.
A POSIX-compliant FUSE (Filesystem in Userspace) client is added directly to each sandbox. Agent code therefore sees a normal directory. Drives can attach to running sandboxes at any mount path. Recreating those sandboxes is unnecessary. A sandbox can mount the whole drive or a selected subdirectory. It can also mount the drive in read-only mode.
Written files are immediately readable. However, no formal consistency or locking model is published yet. Namespace partitioning therefore remains appropriate. Rename behavior also needs testing before a workflow depends on it.
Access control is currently workspace-level. Fine-grained path and label controls are not yet available for drive permissions. Any workload with workspace access may reach the drive unless its mount is read-only or limited to a selected subdirectory. Directory conventions and subdirectory mounts do not create path-level authorization during the preview. Use separate workspaces when a security boundary requires stronger isolation.
How to implement shared file access in a multi-agent pipeline
A working shared filesystem requires provisioning, consistent mounts, naming conventions, and concurrency controls. The directory and naming decisions apply to any shared filesystem.
1. Provision the shared filesystem and define the directory structure
In Blaxel's Agent Drive preview, bl drive create --name my-drive --region us-was-1 creates the drive. Design the directory tree before any agent runs. Drive size isn't configurable during the preview. Drives scale automatically and have no fixed capacity limit. Pick one mount path for every sandbox, such as /mnt/data.
Split the tree by concern:
/tasks: The planning agent writes task definitions here./work/{agent-id}: Each execution agent keeps its in-progress files here./outputs: Agents write completed results here./status: Agents write status signals here./context: The coordinator writes shared inputs here for every agent to read.
Treat the directory layout as an application protocol rather than an authorization boundary. Give each agent explicit path configuration and one assigned namespace. The coordinator should be the only application component that writes /context.
Whole-drive read-only mounts can protect a sandbox that never needs to write. A reviewer, for example, can mount the drive with readOnly when it only consumes completed outputs. An executor needing /outputs can mount that selected subdirectory read-write, while /context can be provided through a separate read-only mount. Because access control remains workspace-level, use a separate workspace when that separation must be a security boundary. Otherwise, validate paths inside the application and limit workspace membership.
2. Mount the shared filesystem in each sandbox
Every sandbox mounts the same filesystem at the same path. The planner and all downstream agents follow this rule. The orchestrator then calls sandbox.drives.mount(). In TypeScript:
await sandbox.drives.mount({
driveName: "my-drive",
mountPath: "/mnt/data",
drivePath: "/",
});
driveName and mountPath are required. drivePath mounts a subdirectory. readOnly blocks writes from that sandbox. Mount the drive root once per sandbox. This layout keeps a claim rename within one filesystem. On Linux, a rename across separate mounts fails. Expect separate drivePath mounts to break the claim step. Test that behavior before deployment.
Your orchestrator makes the SDK calls from outside the sandbox. Agent code inside doesn't need the Blaxel SDK. Batch Jobs can execute parallel fan-out tasks, and Blaxel also provides shared Agent Drive mounts for artifacts. The compute scheduler and storage protocol remain separate concerns.
Use identical configuration keys across every participant. A planner and reviewer should not hard-code different mount paths. Consistent configuration also makes local testing easier. A temporary local directory can be the mount during development. Production configuration can then replace only the root path.
3. Define naming conventions for task and result files
Naming conventions are the inter-agent protocol. Define them before the first agent writes a file. Name tasks by universally unique identifier (UUID) in the format task-{uuid}.json. Signal completion with task-{uuid}.done. Namespace outputs by agent with the format output-{agent-id}-{task-uuid}.json.
Use UUIDv7 as the uniqueness mechanism. Agents writing within the clock's resolution can produce an identical wall-clock timestamp and collide if it is used alone. UUIDv7 places a 48-bit millisecond timestamp in the leading bits. It fills the remaining 74 bits with random data. The result provides sortable filenames without relying on timestamps alone.
Store the convention in pipeline configuration. It is the contract every new or updated agent must honor. Include file extensions and completion-marker rules in that contract. Define whether task identifiers remain stable across retries. Stable identifiers simplify deduplication, while unique attempt identifiers preserve failed outputs for diagnosis.
Names should carry enough information for routing without duplicating full metadata. Put larger task details inside the JSON file. Keep identifiers safe for filesystem paths and consistent across languages. Agents should never derive ownership from an untrusted free-form filename. Validate each identifier before joining it to the mount path.
4. Handle concurrent reads and writes
Apply one-writer ownership: each agent writes only to its own namespace and reads completed outputs from the others. This model prevents most collisions without locks and makes partial failures easier to inspect. Each output has one responsible writer and one predictable location.
Sometimes a shared target file is unavoidable. An aggregate receiving contributions from several agents is one example. In that case, write to a temporary file. Then rename it into place. Linux guarantees that rename atomically replaces an existing target. Readers therefore never see a half-written file.
That guarantee covers atomic publication. If several writers rename files to the same target, one complete file can replace another. Rename must also remain within one filesystem. Object-storage-backed FUSE mounts can't rename directories atomically. Agent Drive's rename atomicity is not formally specified during the private preview. Test the pattern before depending on it. The test should race multiple writers and verify that readers observe only complete files.
Batching helps too. Have each agent collect its contributions before writing once. Don't append many small updates to a shared file. For a true multi-writer aggregate, use separate contribution files. Let one coordinator create the final result after all completion markers appear.
Multi-agent file-sharing patterns
Common pipeline patterns separate work ownership from shared visibility.
The producer-consumer pattern
One producer writes task files into /tasks. Many consumers scan that directory for work. The claim step renames a task into the consumer's /work/{agent-id} directory. If two consumers race to rename the same source path to distinct targets, at most one rename should succeed; self-renames or renames where source and target already resolve to the same file are no-ops that may also return success. The other fails with ENOENT because the source entry is gone.
Test that race using the rename validation described in step four. Claiming by rename needs no application lock when the filesystem provides the expected behavior. Each task also needs its own uniquely named file. Consumers should classify a failed rename as a normal lost race. They can return to scanning for another available task.
When a consumer finishes, it writes its result to /outputs. It then writes a .done marker beside the result. The producer or downstream aggregator lists /outputs for markers. Those markers reveal which tasks are complete. Write the result before the marker. A marker's presence should always mean the data is complete.
Retries need the same ownership rules. A retry can reuse the task identifier while adding an attempt identifier to its work directory. The coordinator should accept only one successful completion for each task. Failed attempt files can remain available for debugging without becoming valid outputs.
The shared context pattern
Agents working on different parts of one task often need the same inputs. Examples include the original brief and prior conversation history. Retrieved documents or a shared knowledge base can also provide context. Put that context at a well-known path such as /context. Have every agent read it at startup.
When the coordinator retrieves a document, it writes the file once. The same process applies when it revises the brief. Every agent sees the update on its next read. No message fan-out is required. Version context files when agents must use one consistent snapshot. An immutable version directory prevents a mid-run update from changing inputs unexpectedly.
Anthropic's research subagents use a related approach. They store work in external systems and pass back lightweight references. This reduces token overhead because agents exchange paths instead of repeatedly embedding complete artifacts.
In an ensemble, each agent reads /context and produces its result. It writes that result to its own namespace under /outputs. A coordinator reads all outputs and synthesizes them. If storage must enforce a read-only /context boundary, place it on a separate read-only drive in a workspace whose membership reflects that boundary.
The pipeline handoff pattern
In a sequential pipeline, Stage One writes its output to /outputs/stage-1. It then drops a status file in /status. Stage Two starts and reads that output. It then writes its own output. Each stage reads the shared filesystem instead of receiving a payload from the previous agent.
Stages are decoupled in time. Stage Two doesn't need to run while Stage One works. During the private preview, Agent Drive contents survive sandbox deletion. A failed stage can also be retried alone because its inputs remain on disk. Batch Jobs can schedule parallel or asynchronous stages when the workflow requires fan-out.
Suppose Stage One dies while writing /outputs/stage-1. No status file exists yet. Stage Two therefore never reads partial data. That protection depends on ordering. Write each status file only after its output is complete. A temporary file and rename can further prevent readers from observing partial output.
On restart, a stage without a status file reruns from its inputs. The crash therefore costs one stage's work. Idempotent stages make that retry safer. A stage should either replace its incomplete attempt or write a new attempt directory. The coordinator can then select the completed attempt indicated by the status file.
Build inter-agent coordination on a shared filesystem
For coding and code-generation agents first, as well as other multi-agent workloads, a shared filesystem fits when agents need low-latency access to each other's outputs. The directory tree becomes the coordination protocol: mounts provide visibility, while ownership conventions and whole-mount controls define safe application behavior.
Object storage remains useful for occasional or cross-region handoffs. Batch Jobs can provide parallel compute for large fan-out workloads. Agent Drive remains the shared-storage mechanism for shared context and tasks. Resulting artifacts stay there as well.
Blaxel, the perpetual sandbox platform, pairs Sandboxes with Agent Drive for this collaboration pattern. It gives sandboxes one replicated filesystem with concurrent read-write access. For pre-loaded datasets on one sandbox, volume templates remain the better fit. Request preview access through blaxel.ai/contact, or explore the platform at app.blaxel.ai.
FAQ
Can two AI agents write to the same file simultaneously?
Yes, but overlapping concurrent writes have no defined final result. Use one writer per path. Atomic rename can publish a complete file without exposing partial data, but competing renames can still replace one another.
What's the difference between Agent Drive (private preview) and a volume for file sharing?
A volume is block storage attached to one sandbox at creation, with access limited to that sandbox. Agent Drive is a distributed filesystem that multiple sandboxes can mount simultaneously, including on running sandboxes, with read-only and subdirectory-mount options. Use volumes for private, high-I/O data owned by one sandbox; use Agent Drive for shared artifacts across agents or sessions. Agent Drive also includes storage-layer replication.
Do agents need special code to use a shared filesystem?
Agent code uses ordinary file operations on paths under a configured mount such as /mnt/data, so existing path-based libraries usually work unchanged and no special file-transfer code is required. The orchestrator uses the Blaxel SDK or API to attach the drive; sandbox code does not need that SDK for file access. A readOnly mount rejects writes, while writable mounts still require application-level directory ownership conventions.
How does file sharing work for agents in different regions?
During the private preview, Agent Drive and its mounting sandboxes must be in us-was-1. Cross-region handoffs use object storage for the artifact and a queue carrying a stable key, task identifier, and completion state. The consumer then downloads it. Keep tightly coupled collaborators in one supported region when possible. For multi-region architectures, contact the Blaxel team at blaxel.ai/contact.
Related articles

Guides
How to schedule recurring AI agent tasks with cron expressions
Learn how a cron job as a service runs recurring agent tasks without a scheduler to manage. Configure Blaxel Batch Jobs with cron, alerts, and retries.
10 min read

Guides
Persistent storage for AI agents: choosing between in-memory, volumes, and shared filesystems
Compare AI agent storage options: in-memory tmpfs, persistent volumes, and shared filesystems. Use a three-question test to match each data type to the right one.
13 min read

Guides
Sandbox snapshots: hibernate and resume agents fast
Learn how sandbox snapshots capture filesystem, memory, and process state so AI agents resume in milliseconds instead of rebuilding from scratch every session.
11 min read