Skip to content

Free tool · no email needed

AI agent cost calculator

Work out the LLM API cost and infrastructure of running an AI agent, support chatbot or RAG app per month. Pick a workload, model and volume to see a chatbot cost per month, per conversation and per 1,000 tasks, using vendor list prices checked October 2026.

1What does the agent do?
2Which model?

OpenAI

Anthropic

Google

List prices checked October 2026 from each vendor’s pricing page; see the table below the calculator.

3How many conversations a month?

One conversation is about 21,000 input and 1,500 output tokens across 6 LLM calls.

4Cost controls
Per month · 10,000 conversations$65 to $173Details
How it works

How the calculator prices an agent.

Model and infrastructure prices checked on 9 October 2026 against each vendor's pricing page; formula and planning figures reviewed by Zain Khalid Malik, CTO.

Each workload carries a token profile from agents we have shipped: input tokens per LLM call (system prompt, tool definitions, retrieved context and conversation history), output tokens per call, and how many calls one unit of work takes. A support conversation is about six turns of 3,500 input and 250 output tokens; an agent that takes actions runs about eight steps of 8,000 input tokens each. Tokens are multiplied by the vendor’s list price per million and by your monthly volume.

Prompt caching bills the repeated part of the input at the cache-read rate: the published cached-input price where the vendor lists one (5 percent of input on Claude Opus and Sonnet 5.5), otherwise 10 percent of the input price. The batch toggle takes 50 percent off the model cost and is only enabled for document extraction and document Q&A, the two workloads that can run offline.

Model cost is shown as a range of plus or minus 20 percent because conversation length, retries and context size vary from user to user. Infrastructure uses our planning bands for hosting and a vector database by volume, plus 5 to 10 percent of model spend for evaluation and tracing. For a phone agent, telephony, speech-to-text and text-to-speech are added per minute from the vendor pages listed under the voice table, with the LLM priced from the same token maths.

The calculator covers running cost only. Read what it costs to run an app for the hosting side and AI agent development cost for the build.

Typical monthly running costs

Four example scenarios computed with the calculator
Support chatbotClaude Haiku 5.5 · 10,000 conversations · $0.0017 per conversation$65 to $173
Document Q&A (RAG)Claude Sonnet 5.5 · 20,000 questions · $0.0126 per question$261 to $482
Agent that takes actionsGPT-5.4 mini · 5,000 tasks · $0.0480 per task$232 to $397
Voice agent, self-hostedClaude Haiku 5.5 · 10,000 minutes · $0.0339 to $0.0471 a minute$369 to $552

Model cost plus hosting, vector database and evaluation, with prompt caching on. Open a scenario in the calculator with ?workload=support-chatbot&model=claude-haiku-5.5.

Rule of thumb: a support chatbot on a small model costs well under a cent per conversation ($0.0017 on Claude Haiku 5.5), so at 10,000 conversations a month the hosting and vector database cost more than the tokens. Model choice matters once volume or context grows.

Reference

Model prices (checked October 2026).

USD per 1 million tokens, from each vendor's public pricing page. Cached input is the vendor's cache-read price where published.

LLM API list prices per million tokens
ModelVendorTierInputCached inputOutput
GPT-5.5OpenAIfrontier$5.00about $0.500$30.00
GPT-5.4OpenAIfrontier$2.50about $0.250$15.00
GPT-5.4 miniOpenAImid$0.75about $0.0750$4.50
GPT-5.4 nanoOpenAIsmall$0.20about $0.0200$1.25
GPT-6 SolOpenAIfrontier$2.00about $0.200$10.00
GPT-6 LunaOpenAIsmall$0.10about $0.0100$0.50
Claude Opus 5.5Anthropicfrontier$4.00$0.200$20.00
Claude Sonnet 5.5Anthropicmid$2.00$0.100$10.00
Claude Haiku 5.5Anthropicsmall$0.10$0.0100$0.50
Gemini 3.1 ProGooglefrontier$2.00about $0.200$12.00
Gemini 3.8 FlashGooglemid$0.75about $0.0750$3.75
Gemini 3.5 Flash-LiteGooglesmall$0.30about $0.0300$2.50

Voice stack, per minute

Per-minute telephony, speech and platform prices for a phone agent
SetupLine itemPer minute
Self-hostedTelephony: Twilio US inboundOutbound is $0.014 per minute$0.0085
Speech-to-text: Deepgram Nova-3 streaming$0.0048 promotional, $0.0077 regular$0.0048 to $0.0077
Text-to-speech: ElevenLabs API to Deepgram Aura-2ElevenLabs API promotional $0.022 per 1K characters, about $0.02 a minute of speech; Deepgram Aura-2 $0.030 per 1K characters$0.02 to $0.03
VapiPlatform fee: VapiPay-as-you-go allows 4 concurrent calls; Core is $29 a month for 10$0.05
Telephony: Twilio US inboundOutbound is $0.014 per minute$0.0085
Speech-to-text: Deepgram Nova-3 streaming$0.0048 promotional, $0.0077 regular$0.0048 to $0.0077
Text-to-speech: ElevenLabs API to Deepgram Aura-2ElevenLabs API promotional $0.022 per 1K characters, about $0.02 a minute of speech; Deepgram Aura-2 $0.030 per 1K characters$0.02 to $0.03
Retell AIVoice engine: Retell (transcription and orchestration)$0.055
Telephony: RetellFree with your own SIP trunk$0.015
Text-to-speech: Retell, Cartesia or OpenAI voices to ElevenLabsElevenLabs v3 voices are up to $0.10 a minute$0.015 to $0.03

Per-minute list prices checked October 2026. Speech-to-speech models such as gpt-realtime-2.1 ($32 input / $64 output per 1M audio tokens) are not modelled; they usually land above $0.30 a minute. Compare the managed platforms in Vapi vs Retell vs Bland.

Infrastructure bands, per month

Hosting and vector database planning figures by monthly volume
VolumeHostingVector DB
Small (up to 25,000 a month)$30 to $80$20 to $70
Medium (up to 250,000 a month)$100 to $350$50 to $250
Large (over 250,000 a month)$400 to $1,500$200 to $1,000

Our planning figures from production agents, not vendor list prices. Small: about $50 to $150 a month for hosting and a vector database; medium: $150 to $600; large: $600 to $2,500. Evaluation and observability add 5 to 10 percent of model spend.

Build cost

Running cost is the small number. Here is the build.

Most agents cost far less to run than to build. Adding an AI feature to an existing product runs $8k to $60k and an AI-first MVP $25k to $60k with our AI agent development team. A support chatbot on your website is $10k to $18k; one that takes actions in your systems $18k to $35k; a multi-channel assistant with handover and analytics $35k to $60k. A phone agent pilot is $12k to $22k, a production voice agent $22k to $45k and an enterprise rollout $45k to $90k+.

Every build includes the cost controls above: model routing, prompt caching, token budgets and an evaluation set, so the number this calculator gives you holds up in production. Use the app development cost calculator for a feature-level build estimate, and see our pricing for rates.

FAQ

AI agent cost questions, answered.

How accurate is the AI agent cost calculator?

The model cost is arithmetic on vendor list prices and the token counts shown under each workload, with plus or minus 20 percent for conversation length and prompt variance. The infrastructure band is our planning figure from agents we run in production. Your own prompts, context size and retry rate can move the model cost by two or three times, so measure a week of real traffic before setting a budget.

What is excluded from the estimate?

Building the agent (see the build cost section), embedding your documents the first time, human review time, fine-tuning, vendor minimum plans and concurrency fees, phone numbers, and speech-to-speech models such as gpt-realtime. Enterprise agreements and committed-use discounts are also outside list prices.

Why do LLM API prices keep changing?

Vendors release a new generation every few months and reprice the old one. Prices here were checked on October 9, 2026 against the OpenAI, Anthropic and Google pricing pages linked in the model table. Design the agent so you can switch models behind a config flag rather than a rewrite.

How does prompt caching cut the cost?

The system prompt, tool definitions and earlier turns are resent on every call. When they are cached, that part of the input is billed at the cache-read rate: 5 percent of the input price on Claude Opus and Sonnet 5.5 and about one tenth on OpenAI and Gemini models. For a support chatbot with 60 percent of input repeated, that is a large share of the input bill.

When does the Batch API apply?

All three vendors take 50 percent off requests sent through their batch endpoints, with results returned within hours instead of seconds. It suits document extraction, nightly re-indexing, bulk classification and offline evaluation. It cannot be used for a live chat, phone call or an agent a person is waiting on, which is why the calculator disables it for those workloads.

How can I cut an agent's running cost by 50 to 80 percent?

Route easy requests to a small model and escalate only the hard ones; turn on prompt caching; trim the system prompt and the number of tool definitions sent per call; cap retrieved chunks and conversation history; batch anything that is not interactive; and set per-user and per-day token budgets. Together these usually take 50 to 80 percent off the model bill without a visible drop in quality.