10 min
How to schedule recurring AI agent tasks with cron expressions
Learn how a cron job as a service runs recurring agent tasks without a scheduler to manage. Configure Blaxel Batch Jobs with cron, alerts, and retries.

Your data analysis agent needs to run every night at two am. It pulls customer data and generates a summary report. It then writes the result to your data warehouse. Kaelio builds natural-language analytics agents.
Right now, someone triggers the job manually every morning. When they forget, the report doesn't exist. When they're on vacation, nobody knows how to run it. Cron syntax configures the recurring Blaxel Batch Job. Monitoring, failure handling, and timezone settings govern its unattended runs.
Recurring work is a solved problem in traditional infrastructure. Cron has scheduled jobs for decades. However, running a cron scheduler for agent workloads still means owning a stack. That stack includes a scheduler process, compute for each run, failure handling, and monitoring.
Teams using managed sandbox infrastructure can still end up hand-managing a scheduler. The adjacent compute isn't theirs to manage at all.
The scheduler belongs inside the platform that runs the job. The scheduling primitive sits in the same configuration file as the execution primitive.
TL;DR
- Cron scheduling eliminates manual job triggers: Define a cron expression. The platform runs the job without anyone clicking a button.
- Built-in batch job scheduling: Scheduling is built into the batch jobs platform. No cron daemon, dedicated VM, or external orchestrator.
- Each scheduled run is a full batch job: Every trigger provisions isolated microVMs for task execution. The microVMs terminate afterward.
- Failure handling is per-run, not per-schedule: A failed run leaves future runs untouched. The next scheduled execution starts fresh.
- Cron expressions handle most recurring patterns: Nightly, hourly, weekly, and custom intervals all fit standard five-field syntax.
Why agent workloads need managed scheduling
Scheduling looks trivial until you count what a reliable scheduler needs.
Self-hosted cron adds infrastructure you have to run
Running cron yourself means keeping a daemon alive and maintaining its VM. You also need scheduler monitoring and alerts for missing triggers. Cron's failure domain is one machine. In a large datacenter, losing that host stops every scheduled job.
That's a 1/1000th failure with total blast radius. Slack ran its cron jobs on a single server for a decade. Maintenance pain then forced a rewrite into a distributed scheduler.
Redundancy trades one problem for another. Two schedulers without coordination fire every job twice. Google used multiple replicas, Paxos consensus, and leader election. Few agent teams want that machinery for a nightly report.
A self-hosted scheduler also splits the stack for teams using managed sandbox infrastructure. Compute is managed, but scheduling is not. A silent scheduler failure resembles an execution failure. You debug the wrong layer first.
Managed scheduling collapses two problems into one
A managed cron service combines when to run with where to run by keeping the cron expression and job code in one file; that file also holds the input parameters. The platform evaluates the schedule, provisions compute, runs tasks, and cleans up. Scheduling inherits the platform's reliability guarantees. There's no separate service level agreement (SLA) or second monitoring stack.
Kubernetes CronJob, AWS EventBridge Scheduler, and Google Cloud Scheduler use this model for their compute targets. Blaxel Batch Jobs apply it to tasks running in a hardware-isolated microVM. This makes each scheduled run standalone.
The platform spawns one microVM per task. It executes the tasks, writes outputs, and tears the microVMs down. Nothing runs between executions, so state cannot accumulate and drift across nightly runs.
Cron jobs are Included under Batch Jobs. The schedule adds nothing to the bill. You pay for active compute during each run.
Cron expression syntax for agent scheduling
Before configuring a schedule, it helps to understand the five-field syntax that defines when your job runs.
The five-field cron format
A cron expression has five space-separated fields: minute hour day-of-month month day-of-week. Standard crontab(5) ranges are:
- minute 0–59
- hour 0–23
- day-of-month 1–31
- month 1–12
- day-of-week 0–7, where 0 or 7 means Sunday
Each field accepts a wildcard (*), range (8-11), list (1,15), or step value (*/15). Step values are widely supported extensions rather than part of the POSIX (Portable Operating System Interface) standard.
| Expression | Meaning |
|---|---|
0 2 * * * | Every night at 2:00 AM |
0 * * * * | Every hour on the hour |
0 9 * * 1 | Every Monday at 9:00 AM |
*/15 * * * * | Every 15 minutes |
0 0 1 * * | First day of every month at midnight |
Read an expression from left to right before deployment: first confirm the minute and hour, then check whether the calendar fields limit the run to particular dates, months, or weekdays. Keep a plain-language interpretation beside the expression during review. This makes accidental daily or hourly schedules easier to catch before activation.
Cron expressions use the scheduler's configured time zone. Some platforms use UTC, while others use host-local time. Confirm the zone before scheduling around local business hours.
Choosing the right schedule for agent workloads
Match cadence to data freshness. A nightly report agent runs daily. A monitoring agent might run every five or 15 minutes. Leave room for the worst-case runtime plus verification time. A job that occasionally runs longer than its scheduling interval breaks its cadence.
Avoid the top-of-the-hour pile-up. If every job uses 0 * * * *, provisioning spikes at :00 each hour. This thundering herd often occurs because teams choose midnight.
Stagger schedules several minutes apart with 7 * * * * or 23 * * * *. AWS EventBridge Scheduler offers a flexible time window for the same purpose. It disperses invocations within a window you set.
External rate limits create another constraint. OpenAI commonly expresses API limits using metrics such as requests per minute (RPM) and tokens per minute (TPM), among others. A job fanning out hundreds of tasks can hit 429 responses against such an API. Set task concurrency and stagger launches so aggregate request and token rates stay below provider limits.
How to configure a recurring batch job with cron
The following four steps take a job from deployed code to a monitored, self-running schedule.
1. Deploy the job code
- Write and deploy the job function before creating a schedule.
- Blaxel Batch Jobs support Python and TypeScript.
- Scaffold with
bl new job, then runbl deployfrom the project directory. - Follow the CLI path in Blaxel's job deployment guide, or use
bl push. - The
bl pushcommand publishes the image without creating or updating a deployment. - The function receives a list of tasks as input parameters and produces output.
- Design it to be idempotent from the first commit.
- Never use the Python datetime
now()function inside a task. - Derive the data window from the schedule boundary instead.
- Write output to a deterministic location keyed by that window.
- A rerun then overwrites output instead of creating duplicates.
Verify the job with a manual trigger before attaching a schedule:
bl run job <JOB-NAME> --data '{"tasks": [{"name": "John"}, {"name": "Jane"}]}'
A job that succeeds manually but fails under cron likely has a scheduling or trigger-configuration problem.
2. Define the cron schedule and input parameters
- Add a
[[triggers]]block withtype = "cron"inblaxel.toml. - This trigger type is valid only for job resources in the deployment reference.
- The
taskslist under[triggers.configuration]defines each scheduled execution's input.
[[triggers]]
id = "cron-trigger"
type = "cron"
[triggers.configuration]
schedule = "0 * * * *"
tasks = [
{ p_id = "1234", arg2 = "foo" },
{ p_id = "5678", arg2 = "bar" }
]
- Replace
0 * * * *with0 2 * * *for the nightly report. - Without a
taskslist in the trigger configuration, each trigger runs one task without arguments. - The documented
tasksvalues are static. - Compute changing date windows inside the job from the scheduled execution boundary.
- Round the execution timestamp down to that boundary.
- A delayed retry then processes the original scheduled window.
- Redeploy with
bl deployto activate the schedule. - Blaxel defines an execution as a run at a timestamp, with tasks running in parallel.
- The trigger block schedules those executions.
- Each firing therefore creates a new execution.
3. Verify the first scheduled run
- Wait for the first trigger and verify the execution:
- It started at the expected minute with every task reaching
succeeded. - The output reached its destination.
- It started at the expected minute with every task reaching
- Blaxel documents job execution states as
pending,running,completed, andfailed; its API examples also showqueuedwhen executions wait for capacity:queuedpendingrunningcompletedfailed
- Per-execution counts cover successful, failed, and retried tasks.
- Pull logs with
bl logs job my-job my-execution-id. - You can add a task ID to isolate one task's output.
A first run may fail because the schedule fired at the expected UTC minute but the wrong local hour. Other common causes are:
- The configured
taskslist omits a parameter that the function expects. - The job's credentials cannot write to the output destination.
A successful bl run job --data test validates the manually supplied data shape. It does not validate the trigger's configured tasks list. Compare both inputs when investigating a missing parameter. Don't retire the manual process until monitoring exists. The previous operator remains your only alert until step four is complete.
4. Configure alerting and failure handling
A silently failing scheduled job is worse than no scheduled job. Someone may still trust its output. Alert on execution failures. Also watch for:
- missed triggers
- runs exceeding their usual duration
Missed triggers require a heartbeat check because nothing errors when nothing starts. These tools implement that pattern:
- Prometheus
absent_over_time() - Google Cloud Monitoring absence conditions
- CloudWatch's "treat missing data as breaching" setting
Track the timestamp of the last successful run. Alert when it becomes older than one schedule interval.
Then set the failure policy:
- Blaxel's
maxRetriessets automatic retries before a task is marked as failed. - When workspace capacity is full, new executions queue and retry until capacity becomes available.
- Set
allowQueue: falseto receive a429error instead. - Prefer skipping to double-launching when duplicate work is more dangerous than a missed launch.
- "Recovering from a skipped launch is more tenable than recovering from a double launch."
- A cache warmer tolerates a skip.
- A monthly billing job does not, so give it retries and an alert.
- Managed schedulers treat each trigger independently.
Operational patterns for scheduled agent jobs
Once a schedule is live, day-to-day reliability depends on how you handle time zones, overlapping runs, and deployments.
Timezone handling for globally distributed teams
Schedulers disagree on their default time zones. Google Cloud Scheduler defaults to UTC. Kubernetes CronJob defaults to the kube-controller-manager process time zone. You can override it with .spec.timeZone. The same expression can therefore fire at different wall-clock times.
Blaxel's TriggerConfiguration schema includes schedule and tasks. Verify the evaluation zone empirically before relying on local time. Deploy a throwaway job on a frequent schedule and compare execution timestamps against the clock.
Daylight saving time (DST) makes "two am in New York" move relative to UTC. AWS EventBridge Scheduler drops schedules in the skipped spring hour. The repeated autumn hour runs only once. Account for these transitions when a report must align with a local business hour.
Schedule in UTC to avoid DST ambiguity. Document the local equivalent and accept the twice-yearly one-hour shift. Serve regions with separate schedules rather than one global schedule. A 02:00 UTC report means different business periods in London and Sydney.
Preventing overlapping runs and managing versions
A job schedule overlaps when execution takes longer than its interval. The second execution starts while the first is still writing. Managed schedulers can use concurrency policies. Kubernetes CronJob offers Allow, Forbid, and Replace. Under Forbid, long-running jobs may cause scheduled times to be skipped.
Google Cloud Scheduler "will never allow two simultaneously outstanding executions." It delays the next start until the previous execution ends.
Blaxel controls within-execution parallelism through maxConcurrentTasks in the [runtime] block. Enforce cross-execution boundaries inside the job when overlap would damage output. Write a lock keyed by the schedule window. Alternatively, check whether the previous window's output already exists. Track execution duration using timestamps and logs.
Widen the interval when runs regularly approach it. You can also divide the work among more tasks. Fan-out execution supports that split.
Scheduled jobs run unattended. A bad deployment can therefore run at the next trigger without anyone watching. Treat deployments like production service releases. Ship during a low-activity window and watch the first scheduled run. Keep a rollback ready.
When failures occur, roll back first and diagnose afterward. Mark bad pipeline output, then reprocess it.
Blaxel deploys job revisions with a blue-green strategy. It stores the last five revisions for Batch Jobs on Mark 3 infrastructure. Configuration updates affect new executions only. Running executions finish with their original configuration. The RevisionConfiguration object exposes active, canary, and canaryPercent.
A critical job can first send a traffic slice to a new revision.
Schedule agent tasks without running a scheduler
This pattern fits coding agents and scheduled evaluation suites, including nightly reports. Blaxel also provides Sandboxes for interactive code execution. Jobs, executions, and tasks are detailed in the Batch Jobs overview.
Start building at app.blaxel.ai or talk to the team at blaxel.ai/contact.
FAQ
What is a cron job as a service?
It is a managed way to run recurring work without operating a scheduler. Use it when an agent must execute on a predictable cadence but should not remain active between runs.
Can I schedule different agent tasks at different frequencies?
Yes. Create a separate batch job and cron configuration for each cadence so changes to one workload do not alter another.
What happens if a scheduled job fails?
The failed execution is recorded, while future scheduled runs remain unaffected and start fresh.
Related articles

Guides
How to share files between AI agents working on the same task
Learn how to share files between AI agents using shared filesystems, naming conventions, and concurrency controls for reliable multi-agent pipelines.
15 min read

Guides
Persistent storage for AI agents: choosing between in-memory, volumes, and shared filesystems
Compare AI agent storage options: in-memory tmpfs, persistent volumes, and shared filesystems. Use a three-question test to match each data type to the right one.
13 min read

Guides
Sandbox snapshots: hibernate and resume agents fast
Learn how sandbox snapshots capture filesystem, memory, and process state so AI agents resume in milliseconds instead of rebuilding from scratch every session.
11 min read