Durable Execution Pricing Models in 2026: Metered, Flat, and What Agentic Workloads Do to Each
Durable execution pricing falls into usage and capacity models. See what each meter counts and how agentic workloads affect both.
Haziqa Sajid
Technical Writer
Your agent calls a tool. The call fails, the runtime retries, and the fourth attempt goes through. One piece of useful work, four attempts, three of them wasted.
What those three wasted attempts add to your durable execution cost depends entirely on which platform sits underneath them. Temporal's cost guidance is explicit that every Activity retry counts as one Action, so you're paying for all four. Cloudflare draws the opposite line. Its pricing reference says step count leaves out rollback handlers and retries, so those three failures never reach that meter at all.
Same failure, same agent, two different answers about who pays for it.
Both of those are orchestration bills, the platform's charge for running your workflow. Neither is what your model provider charges for the tokens that tool call burned, which is a separate line and one we've covered before. This article is about the orchestration side. What's awkward about it is that the bill moves on how your agent behaves, and your model vendor's price list has nothing to do with it.
Durable execution pricing comes in two shapes. Either you pay for units of work, which is what Temporal and Cloudflare are doing and what the industry usually calls usage-based pricing, or you buy capacity up front and pay for the size whatever runs inside it. Neither label tells you a thing about what you'll actually be charged for. So the rest of this goes to the counting rules, what each model bills, what it quietly leaves out, and what an agentic workload does to both.
What the metered model bills, and what it quietly excludes
Metered pricing, often sold as per-action pricing, charges you for units of work the runtime performs, and every durable execution platform in this group works that way. The difference is what each one treats as a unit, and every vendor decides that for itself. That decision shapes your bill more than the headline rate does.
How Temporal counts an Action
Temporal Cloud calls its unit an Action, which the documentation describes as tracking "billable operations within the Temporal Cloud Service." That covers a great deal more than it sounds like.
| Operation | Actions charged |
|---|---|
| Workflow start | 1 |
| Activity start | 1 |
| Activity retry | 1 per retry |
| Timer, signal, query, or update | 1 each |
| Heartbeat that reaches the server | 1 |
| Child workflow start | 2 |
| Scheduled run | 3 |
Nothing on that list is exotic. Every line is something an ordinary workflow does on an ordinary day, which is why the count climbs faster than teams expect.
Temporal Cloud pricing runs on a sliding scale. As of September 2026, Actions cost $50 per million for the first five million, dropping to $25 per million between one hundred and two hundred million, and active storage adds $0.042 per gigabyte-hour.
Those rates apply above what your plan already includes. Temporal's Essentials plan comes with a million Actions and a gigabyte of active storage, and the plan itself is charged at the greater of $100 a month or five percent of what you consume.
How other platforms count
Cloudflare Workflows pricing runs on four meters at once. Steps come with 500,000 included per month on Cloudflare's paid Workers plan and $0.80 for each additional 100,000, alongside separate meters for requests, CPU time, and storage. AWS Step Functions bills state transitions, at $0.000025 each in US East (N. Virginia) for Standard workflows, one of its two workflow types.
DBOS isn't metered in quite the same way. It sells a $99 a month subscription that includes a million checkpoints, defined as workflows, steps, and transactions, with more at $50 per million.
None of these words are interchangeable. A workload that costs one Action on Temporal doesn't cost one step on Cloudflare, and no conversion exists between them.
What the meters exclude
The bigger differences are in what each meter leaves out, and retries are the clearest case. Temporal's cost guidance states that "Each Activity retry counts as one Action." AWS says the same, that a retry in a step with retry error handling is charged as another state transition. Cloudflare goes the other way, stating that "Step count does not include rollback handlers or retries."
Two platforms bill your failures and one doesn't. The same failed tool call reaches your invoice or never appears on it, depending only on where you happen to be running.
| Platform | Unit | Retries billed | Idle billed |
|---|---|---|---|
| Temporal Cloud | Action | Yes | No Action meter, storage accrues, timer costs 1 |
| AWS Step Functions | State transition | Yes | Not stated |
| Cloudflare Workflows | Step | No | Yes on steps, no on CPU time |
| DBOS | Checkpoint | Not stated | Not stated |
Idle time is treated differently again. A sleeping workflow still costs a step on Cloudflare, since the changelog that introduced step billing defines a step as including "sleeping or waiting for events," yet that same workflow incurs no CPU time. One pause, charged on one meter and not the other.
Where the counting rules live
None of this is visible from the words "usage-based pricing," and the rules aren't always where you'd expect to find them. Cloudflare's pricing reference tells you a step is the number of steps your workflows execute, which is circular. The working definition sits in a changelog entry from July 2026, published a month before step billing was set to start no earlier than August 10.
What the flat model bills instead
Flat pricing charges you for capacity you reserve, not for work the runtime performs. Execution volume stops moving the bill. What you buy instead is a size, and you have to pick it before the month starts.
Microsoft prices one product two ways
Microsoft's Durable Task Scheduler is the clearest case in durable execution. It sells one product in two versions, which Microsoft calls stock keeping units (SKUs), and publishes a spend formula for each.
| Dimension | Consumption | Dedicated |
|---|---|---|
| You are billed for | Actions dispatched | Capacity Units provisioned |
| Monthly spend | (monthly actions ÷ 1,000,000) × regional price per million | provisioned CUs × regional monthly price per CU |
| Commitment | None | Provisioned in advance |
Microsoft defines an action as a message the scheduler sends to your application to trigger an orchestrator, activity, or entity function. On the Dedicated side you buy Capacity Units (CUs) up front at a fixed monthly cost per CU set regionally, and the documentation is blunt that this is "Not 'per action' billing."
One caution on vocabulary. Microsoft's action and Temporal's Action share a name and not much else, because Temporal counts retries, timers, and heartbeats, while Microsoft counts messages dispatched to your application. Neither number tells you anything about the other.
What each model asks you to forecast
Flat pricing doesn't remove the forecast. It changes which number you have to get right.
Consumption asks for a number your workload produces. You can't know in advance how many messages your orchestrations will dispatch. Dedicated looks like the way out of that, because you pick the CU count yourself. Then you read Microsoft's own sizing method and find the same forecast sitting underneath it.
The documentation works CUs out from projected volume. Take your monthly actions, divide by 2,628,000 for actions per second, then divide again by the 2,000 actions per second each CU supports. One worked example runs 20 million orchestrations at 12 actions each, arrives at roughly 91 actions per second, and fits comfortably inside a single CU.
So the CU count isn't really a free choice. It's a volume forecast with an extra division step, and the quantity you're estimating is the one Consumption would have billed you for directly. What changes is the penalty for getting it wrong. Underestimate on Consumption and the meter charges you more, while on Dedicated you meet a ceiling instead, since a deployment supports at most three CUs and Microsoft doesn't say what happens when you reach them.
How Diagrid Catalyst prices concurrency
Diagrid Catalyst runs concurrency-based pricing, the same model with a different unit. The two tiers in its pricing calculator, Dedicated Cloud and BYOC, short for Bring Your Own Cloud, are both sized on concurrent workflows. Diagrid counts that as the number of workflow instances running at a given time, with child workflows counted separately from their parents and repeat instances of the same workflow counted separately as well.
The bands run 0 to 200, 201 to 500, 501 to 1,000, and 1,001 to 2,000 concurrent workflows. Dedicated Cloud starts at $1,499 a month, or $1,199 billed annually, and BYOC is sized the same way at a higher starting price. Workflows and activities aren't metered at all, so the pricing page tells you to "size for concurrency only," never for execution volume.
What agentic workloads do to metered pricing
Temporal publishes its own account of which workloads run up large Action counts, in guidance written to help customers spend less. That makes it the most useful public account there is of agent orchestration cost under a metered model. The platform wrote it, and it has no reason to overstate the problem.
Temporal's own list of cost drivers
The guidance names five patterns behind a high Action count. Many activities per workflow. Frequent signals, queries, or updates. Long-running activities with heartbeats. High retry rates on activities. And extensive query usage. It also spells out the arithmetic behind two of them. Each Activity retry counts as one Action, the guidance says, and so does each heartbeat.
Why agent workloads hit those patterns
Four of those five describe an ordinary agent run. Nothing misconfigured, nothing abused.
One user request turns into a run of reasoning steps, tool calls, and handoffs to sub-agents. That's pattern one, and your team hasn't done anything wrong yet. Those calls then fail in production, where failure is "a given rather than an exception", and each retry adds another Action. They run long enough to need heartbeats. Agents that pause for a human approval take signals and updates to come back.
The fifth pattern, extensive query usage, is about how often you inspect workflow state from outside. That one's an observability habit, something your team chose rather than something the agent does on its own.
The number nobody can forecast
Underneath the first four sits a harder problem. An agent decides its own step count. The same request might take four tool calls or fourteen, depending on what the model concludes along the way. You can't know the unit count before the run, which means you can't know the bill. Diagrid describes it as cost that "swings with every loop you can't predict".
None of this makes metered pricing wrong. It ties your bill to how much the agent reasons, and that's the variable your team has least control over.
What agentic workloads do to flat pricing
Flat pricing takes that variance out of the bill. The step count that made a metered forecast impossible doesn't touch this invoice at all. An agent that takes fourteen tool calls where you expected four costs exactly what one that takes four costs, which is the point of the model.
What it asks in return is a single number, decided before the month starts.
Fan-out sets that number, not you
The number is harder to pick than it looks. The agent picks it for you. Diagrid's documentation lists fanning work out across application instances as one of the reasons to reach for child workflows, alongside composing larger orchestrations and isolating retry boundaries.
Catalyst counts workflow instances, as the previous section covered, so when an agent fans out through child workflows, five sub-agents running in parallel occupy six slots instead of one. Nothing in your tier selection decided that. The model did, mid-run.
The same behavior costs you on the other model too. Temporal prices a child workflow start at two Actions against an activity's one, so fanning out runs up a higher unit count there for the same reason it fills more slots here. Two different currencies, one behavior driving both, and it isn't a behavior you scheduled.
You buy for the peak and pay through the quiet hours
Bands are bought in advance and charged monthly. Your size gets set by your busiest moment, not your typical one. A workload that hits four hundred concurrent workflows during a nightly batch and sits at sixty the rest of the day still pays for the band covering four hundred, every day of the month.
This is where metered pricing wins outright. Entry floors elsewhere are low, with Temporal's Essentials plan starting at a $100 monthly minimum, while a capacity band gets paid for whether you fill it or not.
The question no pricing page answers
Diagrid's documentation says most long-running workflows spend nearly all their wall-clock time durably suspended on a timer or an event, and it classifies an instance in that condition as running.
That describes runtime state, not billing. A scheduler's status label tells you what the runtime is doing with an instance, and says nothing about what a meter counts. It does not tell you what the platform charges for. Still, the distinction matters. If approval-heavy agents spend most of their hours suspended, then whether a parked instance holds a slot is the difference between sizing for the work in flight and sizing for every workflow you've opened and not yet closed. For a human-in-the-loop process those two numbers may be nowhere near each other.
Ask any capacity-priced vendor directly, since that answer isn't on a pricing page.
Capacity pricing suits volume that's high and steady. It works against you when volume is small or shows up in bursts.
The bottom line
Neither model of durable execution pricing is right in general. Metered pricing bills you for what your workload did. Capacity pricing bills you for what you set aside. They price different risks, and which risk you'd rather carry depends on your own traffic.
Agents lean on both. Fan out through child workflows and you fill more slots on one model and accrue more Actions on the other. One behavior, two currencies, neither of them scheduled by you.
The real difference is where the unknown sits. Metered pricing puts it in your unit count. An agent decides how many calls a request takes, so you can't price a run before it happens. Capacity pricing moves that unknown into your sizing, where the same agent decides how wide it fans out. That holds whether you're buying Capacity Units or paying for concurrency bands.
The number you can't pin down travels from one side of the invoice to the other. It doesn't disappear.
The other thing to carry out of this is how much of workflow engine pricing nobody writes down. Cloudflare defines a step in a changelog and not on the page that charges for steps. Microsoft names a three CU ceiling and stays quiet about what happens when you reach it. Whether a parked workflow holds a concurrency slot appears on no pricing page at all. The rate is the easiest number to find and the least useful on its own.
So, three things before you sign anything. Read the counting rules for the platform you're actually considering instead of the category it belongs to, and expect them to live somewhere other than the pricing page. Model your own retry rate against those rules, since that one number decides whether your failures are free or billable. Then ask any capacity-priced vendor how a workflow waiting on a timer or an approval gets counted, and get the answer in writing.
If you're chasing the token side of the bill instead of the orchestration side, an earlier piece here covers what a failed run costs in tokens you already paid for. A separate one works through the wider economics of running agents at scale, including cost drivers that sit outside the orchestration meter entirely.


