Agentic AI Pricing Models Compared
Seven units of billing, three flavours of predictability, one set of worst-case scenarios.
Predictability at a glance
Model-by-model heatmap
Green = predictable enough for a board paper. Amber = needs sensitivity analysis. Red = budget must include a 25%+ contingency.
| Feature | Predictable |
|---|---|
| Per-conversation (Agentforce) | ✓ |
| Per-resolution (Sierra, Decagon) | ◐ |
| Per-message (Copilot Studio) | ◐ |
| Per-seat (Devin, Claude Code, Lindy, Artisan) | ✓ |
| Per-task (Devin Teams overage) | ✗ |
| Per-credit (Manus, LlamaCloud) | ◐ |
| Per-node-hour (LangGraph Platform) | ✗ |
Worst-case scenarios
Per-conversation
Long conversations with many actions inflate per-action cost.
Per-resolution
Resolution definition is vendor-defined; bad definition wins for vendor.
Per-message
Credit-per-action varies; tenant graph lookups consume more than agent messages.
Per-seat
Heavy individual users still capped; light users overpay.
Per-task
Task complexity is unknown until ACU is consumed.
Per-credit
Credit-per-task varies across task types; unused credits expire.
Per-node-hour
LCU consumption is hard to forecast for tool-heavy graphs.
Decision aid
For a forecastable budget, pick per-seat or per-conversation. For outcome alignment, pick per-resolution (Sierra) and negotiate resolution definition tightly. For developer-built systems, pick per-token (Vertex, OpenAI) plus a clear separate hosting line. Avoid per-task and per-node-hour without a working dev-deployment cost model.
Continue with token cost vs platform cost to size the second line on the bill.