Back to blog posts

12 min

How to run batch jobs across many isolated AI agent sandboxes

Learn how to schedule AI agent batch jobs with task-level microVM isolation. Step-by-step guide covering configuration, monitoring, failure handling, and output collection.

Nicolas LecomteNico is a founder of Blaxel, who usually writes about AI, agentics, and the future of AI runtimes.

Your agent pipeline has a large set of documents to process before morning. Each one needs its own run: the agent parses the file, extracts fields, validates them, and writes a result. One sandbox working through the list in sequence finishes in days. A pool of shared workers finishes faster, until one task fills a worker's disk. Every task queued behind it fails too.

Agent batches carry a risk that a nightly ETL run doesn't. Each task executes code a model wrote, or processes a file nobody reviewed. A malicious PDF in one task can't be allowed to touch the next. A hallucinated shell command or a loop that never exits has to die alone.

Task-level microVM isolation contains those failures. Configuration bounds each run, and output planning keeps its results durable. Scheduling controls recurring runs, while concurrency tuning and partial-failure handling keep them manageable.

So the isolation boundary has to sit at the task. One job, many tasks, each in its own microVM.

Examples use Blaxel Batch Jobs, which implement this model directly. The platform provisions each isolated task and cleans it up after execution, up to the configured concurrency. Scheduling governs recurring runs; concurrency tuning and partial-failure handling control their operation.

Best-fit workloads start with coding and code-generation agents that fan out asynchronous work, followed by large dataset processing, test execution across multiple configurations, ETL, data validation, and scheduled analysis.

TL;DR

  • Each task is its own isolation boundary: One job fans out into many tasks. Each runs on isolated microVM-based compute. Provisioning one sandbox per task and tearing it down afterward is the recommended pattern.
  • Failures are task-scoped: A malicious or runaway task can be logged, retried, or reviewed independently while the remaining tasks continue.
  • The platform provisions on demand: Submit the job with its task inputs and a concurrency limit. The infrastructure spins up sandboxes to that limit, drains the queue, and returns to zero. There is no pool to pre-size.
  • Batch jobs are built for async work: Tasks run from minutes to hours. Blaxel documents a boot time of about 30 seconds per task. That startup profile is designed for throughput-oriented work lasting minutes to hours.
  • Outputs need a plan before the first task starts: Sandbox storage disappears when the task ends. Route results to storage that outlives the sandbox.

Why AI agent batch processing needs task-level isolation

Per-task isolation prioritizes containment for asynchronous work. Worker reuse is the mechanism behind the risk. A single process handles many tasks, and filesystem state, memory, and open handles carry forward between them. With a microVM per task, a failure ends at that task's kernel boundary and the queue keeps draining.

Traditional batch workers share failure domains

A worker pool pulls tasks from a queue, and each worker handles many in sequence. When a worker crashes, every task in flight on it fails. A task that leaks memory or leaves files behind damages the next task on that worker. One leaking process degrades other workloads on the same machine. On shared workers, resource leaks can cause excessive paging, kill innocent processes, and force node reboots, as documented in a USENIX OSDI 2022 paper (RESIN).

For trusted, deterministic code, this model holds up. The code is predictable, crashes are rare, and shared state stays clean. Containers remain the standard for running your own software that way.

Agent batch tasks break that assumption. Each task runs code a model generated or handles content nobody reviewed. Meta's CyberSecEval benchmark rated about 30% of code suggestions across seven models as insecure. On a shared worker, a vulnerability that executes reaches whatever that worker can reach. The blast radius grows with every task you add to the batch.

Isolated execution means each task is its own sandbox

The alternative gives every task its own microVM, provisioned when the task starts and destroyed when it finishes. Firecracker, the open-source microVM technology, runs one process per microVM. Each guest gets its own operating system kernel and virtual hardware with separate page tables. AWS Lambda runs on the same technology and limits each environment to one concurrent invocation. One unit of work per microVM is production practice.

If one task crashes, the next keeps running. The two share no filesystem and no memory. A runaway loop stays inside its own kernel and filesystem. Its process table is isolated too.

Batch Jobs are designed for asynchronous tasks that run from minutes to hours, where per-task provisioning fits the workload. Interactive workloads need a different execution model. For those, use Blaxel Sandboxes, which resume from standby in under 25 milliseconds with state intact. That resume speed keeps interactive agent sessions responsive while Batch Jobs handle asynchronous work.

Architecture of isolated batch execution

Three entities structure every batch run: job, execution, and task. A task's lifecycle maps onto a sandbox's lifecycle, which is what makes per-task isolation possible.

Job, tasks, and sandbox lifecycle

A job is the code definition. It specifies the batch processing task and accepts optional input parameters. An execution is one run of that job at a given timestamp, made up of many tasks running in parallel. A task is a single instance of the job definition inside an execution. Processing a large document set means one job, one execution, and one task per document.

In this model, each task runs as an isolated microVM workload and receives its input object. It runs the code, writes its output somewhere durable, and terminates. Storage attached to the job is ephemeral by design. Job volumes are temporary disk-backed storage, created when the job starts and destroyed when it completes. That constraint shapes step four below.

The lifecycle runs in order: execution submitted, tasks distributed, isolated workloads provisioned, code executed, output written, workloads terminated. The status enum includes queued, pending, running, succeeded, failed, cancelled, cancelling, and timeout. In Blaxel Batch Jobs, you define the job code and provide the task inputs. The platform handles distribution, provisioning, and cleanup up to the configured concurrency.

Scaling model with elastic provisioning and no pre-allocation

Sandboxes are provisioned as tasks dispatch, so there's no worker pool to size. Submit an execution with a large task set and a concurrency limit. The platform ramps to the configured number of concurrent sandboxes and works through the queue. It returns to zero when the last task exits. Dispatching a single job across many sandboxes is the fan-out pattern our companion guide walks through.

Blaxel meters job compute in GB-seconds, charged against task runtime. The concurrency limit mainly changes how quickly the work finishes. You are billed for job compute runtime, so idle periods between batches carry no compute charge.

Flexera's State of the Cloud 2026 puts wasted cloud spend at 29%. That's the first rise in five years, and Flexera attributes it to AI workload cost complexity. An always-on worker pool sized for a nightly peak is exactly that kind of waste.

The only sizing decision left is the concurrency limit.

How to configure and run a batch job

Four steps take a job from a function to collected results. The examples use the Blaxel CLI and blaxel.toml; other batch systems expose the same knobs under different names.

Step one: Define the job code and input schema

Write a function that handles one unit of work and nothing more. It receives one input object and produces one result. The function might parse one document or analyze one dataset. Code generation should likewise be scoped to one task. Blaxel supports Python and TypeScript for job code, with a generated entrypoint at src/main.py or src/index.ts.

Define the input schema next. Every task receives a JSON object. The full set of tasks is passed as an array under the tasks key. For example, a document job might carry a document URL per object. A code analysis job might carry a repository reference and a file path. Keep each task's input self-contained so the isolated workload does not depend on state left by another task. Test locally before deploying:

bl run job <<JOB-NAME>> --local --data '{"tasks": [{"name": "John"}]}'

Make the function idempotent. Every message may be retried, so a retry that reprocesses a document should overwrite the same output key, never append a second copy.

Step two: Submit the job with task inputs and concurrency settings

Concurrency is set at deployment time in blaxel.toml under [runtime] as maxConcurrentTasks:

[runtime]
memory = 1024
maxConcurrentTasks = 10
maxRetries = 0

Raise the value to finish sooner at higher peak compute. Lower it to spread cost over time. Then submit an execution through the CLI:

bl run job <<JOB-NAME>> --data '{"tasks": [{"name": "John"}, {"name": "Jane"}]}'

Executions can also be submitted through the API or console.

Set the execution id for idempotency. A client that retries a timed-out submission without one can launch the entire batch twice. The same call over REST is a POST to /jobs/{your-job-name}/executions with the tasks array in the body. The CreateJobExecutionRequest also accepts memory overrides, env overrides, and an id field. Once accepted, the platform distributes tasks across sandboxes up to the concurrency limit.

Step three: Monitor job progress and handle task failures

Track progress through GET /jobs/{jobId}/executions/{executionId}. It returns a JobExecutionStats object with total, success, failure, running, cancelled, and retried counts. Use those counts to distinguish normal queue progress from a growing failure pattern that needs intervention. The console dashboard shows active jobs, failures, and request counts across deployments.

Failed tasks stay isolated. Each one is logged with its error output while the rest continue. Pull a single task's logs with the bl logs command. Add --follow to stream in real time:

bl logs job my-job exec-abc123 task-456

Decide the failure strategy before you submit. Configure automatic retries with maxRetries in the runtime config, or collect failed inputs for review. Set a threshold if the job should stop after too many failures. If you retry, always use randomized exponential backoff. Synchronized retries against a shared API turn one failure into a spike.

Step four: Collect outputs from completed tasks

Output has to leave the sandbox before the task ends, because job storage is ephemeral. Each task POSTs its result to an external API or object store. Or it hands results to a sandbox that has shared storage mounted.

For shared storage, Agent Drive is a distributed filesystem with concurrent read-write access from multiple sandboxes and built-in replication. In private preview in the us-was-1 region, a drive attaches to an already-running sandbox at any mount path. Permission rules support mounting at specific paths and read-only mode. Drive access control is currently workspace-level, with fine-grained access control coming soon.

Mount the drive from a sandbox. Then treat it as the collection point a downstream sandbox reads after the execution finishes.

Results that must survive sandbox deletion should be copied onto a Volume. Settle this before the first task runs. A large execution with no output plan loses every result at termination.

Operational patterns for batch processing at large scale

A job that runs once needs the four steps above. A nightly job also needs scheduling and concurrency limits. It needs a reprocessing path for failures as well.

Schedule recurring batch jobs

Nightly document processing and weekly report generation fit a cron trigger inside the job definition. Daily pipeline runs do too. Blaxel accepts cron as a trigger type on Batch Jobs. Define it in blaxel.toml:

[[triggers]]
id = "cron-trigger"
type = "cron"

[triggers.configuration]
schedule = "0 * * * *"
tasks = [ { p_id = "1234", arg2 = "foo" }, { p_id = "5678", arg2 = "bar" } ]

The platform fires the execution on schedule, and Blaxel's pricing page lists cron scheduling as included. Compute consumed by cron-triggered tasks is billed at the normal job rate.

An in-platform trigger leaves no external orchestrator to keep alive and patch. No credentials go to a separate scheduler either. The trigger ships in the same blaxel.toml as the job code, so schedule and code deploy together.

The documented cron example carries a static tasks array. When the input list changes every run, have an upstream agent build the list. A small preceding job can do the same and submit the execution through the API. Cron expression syntax itself is covered in our companion guide on scheduling recurring agent tasks.

Tune concurrency for cost and throughput

Higher concurrency shortens the batch and raises peak cost. Lower concurrency does the reverse. A large document job due by morning wants the maximum. Weekly analytics with a broad completion window can run at a fraction of that.

Downstream rate limits set a ceiling that compute doesn't. Unsuccessful retries count against OpenAI's per-minute limits. Tasks hammering a throttled endpoint burn quota and produce nothing. Once a rate limit caps throughput, adding tasks only lengthens the queue. Set maxConcurrentTasks at or below the rate limit divided by each task's calls per minute.

For batches that tolerate delay, route model calls through provider batch endpoints. OpenAI's Batch API runs at half of synchronous pricing with a separate rate-limit pool. Anthropic's message batches and Google's Gemini batch API offer comparable endpoints. Set the limit from the deadline you actually have, not the one the batch could theoretically hit.

Handle partial failures in large batches

Some tasks will fail, and most won't fail for the same reason. Unique production conditions accounted for 71% of partial failures in an NSDI 2020 study. Those conditions include bad input, resource contention, a flaky disk, or a faulty remote process. Capture the failing input rather than retrying uniformly.

Log every task's input parameters next to its completion status. After the execution finishes, read the failure count from the stats object. Pull those tasks' inputs from the logs and resubmit them as a smaller execution. The original successes never rerun.

Then set a threshold that stops the job when failures cluster. Step Functions Distributed Map defaults to zero tolerated failures and fails on the first child error. S3 Batch Operations fails a job once failures pass half, after 1,000 tasks. Agent batches belong between those.

A failure spike usually means bad input data or a provider outage. Every further task burns compute on the same fault. Pause, inspect the failed inputs, fix the cause, resubmit.

Run many isolated tasks without building the infrastructure

Batch processing that executes untrusted, arbitrary, or AI-generated code requires isolation at the task level. Teams that run agent batches on pooled workers accept a blast radius that grows with every task added. Containment is decided at provisioning time, when you choose whether each task gets its own kernel. It cannot be patched in later through application error handling.

Blaxel, a perpetual sandbox platform, implements this model in Batch Jobs. You get microVM-level isolation per task and concurrency scaling to the configured limit. Cron triggers support recurring runs. None of it requires custom orchestration.

Start building at app.blaxel.ai or talk to the team at blaxel.ai/contact.

FAQ

When should I use Batch Jobs instead of Blaxel Sandboxes?

Batch Jobs create and tear down task-scoped compute as a queue drains. Blaxel Sandboxes are better when the same environment must remain available between interactions, including its filesystem and memory state.

When should I use Agent Drive, a Volume, or an external store for batch outputs?

Use Agent Drive for shared filesystem access across concurrently running sandboxes. Choose a Volume when one sandbox needs raw I/O performance and retained data, or an external API or object store when systems outside Blaxel need to consume results directly.

Related articles