Skip to content

AI and data · Voice agents

AI voice agent development services for phone lines that answer, book and follow up

We design and build voice agents that take real calls: inbound support and booking lines, outbound reminders and follow-ups, and voice assistants inside your app. Built on Vapi, Retell, ElevenLabs or your own Twilio and Deepgram stack, with the latency, barge-in and handoff work that decides whether callers stay on the line.

Book a call
Close-up of source code on an engineer screen while building a voice agent integration

Short answer

An AI voice agent answers or places phone calls by chaining telephony, speech-to-text, a language model with tools and text-to-speech. Innovation Insight is an AI voice agent development company building them for support, booking and outbound follow-up. A pilot on one call flow costs $12k to $22k over four to six weeks; a production agent with CRM, calendar and handoff integrations costs $22k to $45k. Running costs are usually $0.05 to $0.20 per minute.

Reviewed by Zain Khalid Malik, CTO & Co-founder · Updated

What we build

A voice agent is a chatbot with a much harder interface. Callers will not wait two seconds for a reply, they interrupt, they mumble, and they expect to be put through to a person when the agent is stuck. The agents that earn their keep are narrow: one or two call flows done well, connected to the systems that hold the answer, and measured on every call. That is what we build.

Agent typeWhat it doesTypical integrations
Inbound support lineAnswers account, order and policy questions, verifies the caller, creates tickets, transfers with a summaryZendesk, Intercom, HubSpot, Salesforce, your API
Booking and scheduling lineBooks, moves and cancels appointments, checks availability, sends confirmations by SMSGoogle Calendar, Calendly, Cal.com, clinic and salon systems
After-hours and overflowTakes calls your team cannot, captures the request, books a callback or escalates emergenciesYour phone system via SIP, Twilio, on-call rota
Outbound follow-upAppointment reminders, lead qualification callbacks, payment and renewal reminders, surveysCRM, dialler rules, consent records, calendar
In-app voice assistantLets users ask questions and complete tasks by voice inside a mobile or web appYour product API, retrieval over your content, analytics

How a voice AI agent works

Most production agents are a pipeline. Each stage adds delay, so the whole design is a latency budget.

  1. Telephony: the call arrives over Twilio, a SIP trunk or WebRTC from your app, and audio streams to the agent in small frames.
  2. Speech-to-text: a streaming model such as Deepgram Nova-3 turns audio into partial transcripts and signals when the caller has finished speaking (endpointing).
  3. Language model with tools: the transcript, the conversation so far and the call context go to a model that decides what to say and which tools to call, such as looking up an order or booking a slot.
  4. Text-to-speech: the reply streams to a voice model (ElevenLabs, Cartesia, Deepgram Aura or OpenAI) and the first audio plays before the sentence is finished.
  5. Orchestration: a layer that handles interruptions, silence, retries, recording, transcripts and the handoff to a human.

Speech-to-speech models such as OpenAI's gpt-realtime take audio in and produce audio out in one step. They sound more natural and cut a few hundred milliseconds, but cost several times more per minute and give you less control over what is said. We usually start with the pipeline and test a speech-to-speech model on the flows where tone matters most.

The latency budget is the first design decision. We aim for under about 800 ms from the caller finishing a sentence to the agent starting to speak. Above a second, callers talk over the agent and the conversation falls apart.
  • Barge-in: the agent stops talking when the caller interrupts, and discards the half-spoken reply instead of finishing it.
  • Endpointing: tuned per flow, so a caller reading out a long reference number is not cut off, and a short yes is not followed by two seconds of silence.
  • Handoff: warm transfer to a person with a written summary of the call so far, plus rules for when to transfer without being asked (anger, repeated misunderstanding, emergencies).
  • Guardrails: the agent only promises what its tools confirm, never quotes a price or a slot it has not checked, and reads back critical details.

Platform choice: Vapi, Retell, Bland, ElevenLabs or a self-hosted stack

Managed platforms give you a working agent in days and charge per minute. A self-hosted stack on Twilio, Deepgram, an LLM API and a TTS provider costs less per minute and gives you full control, but you own the orchestration code and the on-call. Vendor prices below were checked in October 2026 on the vendor's own pricing page and will change.

PlatformPrice per minute (checked October 2026)Plans and limitsFits
Vapi$0.05 platform fee plus STT, LLM, TTS and telephony at cost; their calculator puts 1,000 minutes at about $82 to $129Pay-as-you-go (4 concurrent calls), Core $29/month (10), Pro from $999/month (30); extra lines $10/month; HIPAA add-on $2,000/monthTeams that want to pick each model and see each cost
Retell AI$0.07 to $0.31 all-in: voice engine $0.055, telephony $0.015, TTS $0.015 to $0.10, LLM $0.016 to $0.08; speech-to-speech $0.07 to $0.3820 concurrent calls free, then $8/line/month; numbers $2/month; add-ons $0.005 to $0.01/minSupport and booking lines with a strong flow builder
Bland AIStart $0.14 (LLM, STT and TTS included); Build $0.12 plus $299/month; telephony at Twilio pass-through; transfers $0.04 to $0.05Start: 10 concurrent, 100 calls a day; Build: 50 concurrent; Enterprise customOutbound campaigns where a simple all-in rate is preferred
ElevenLabs Agents$0.08 within plan minutes, $0.16 beyondCreator $22, Pro $99, Scale $299, Business $990 a monthProducts where voice quality is the selling point
Self-hosted (Twilio + Deepgram + LLM + TTS)Roughly $0.05 to $0.10 all-in: Twilio $0.0085 in or $0.014 out, Deepgram Nova-3 $0.0048, a mini-class LLM, Aura-2 or ElevenLabs TTSNo platform fee; you run the orchestration server and pay each providerHigh volume, strict data control, or a product where voice is core

Our rule of thumb: a managed platform lands at $0.10 to $0.20 per minute all-in, a self-hosted stack at $0.05 to $0.10, and realtime speech-to-speech models at $0.30 or more. We compare the vendors in detail in Vapi vs Retell vs Bland, and you can model your own volume in the AI agent cost calculator.

What a voice agent costs to build

Build cost depends on the number of call flows, the systems the agent must read and write, and how much conversation testing the domain needs. Our tiers are below; the calculator's voice agent preset lands at $23k to $42k for a comparable scope (see the breakdown).

TierPriceWhat is includedTimeline
Pilot$12k to $22kOne call flow (for example bookings or FAQs), one phone number, a managed platform, a knowledge base from your content, call recording and transcripts, a 200-call evaluation set, handoff to a human4 to 6 weeks
Production$22k to $45kTwo to four flows, caller verification, CRM, calendar or helpdesk integrations with write access, SMS confirmations, a conversation QA dashboard, consent and recording controls, load testing on concurrency8 to 12 weeks
Enterprise$45k to $90k+Self-hosted or private-cloud stack, SIP integration with your contact centre, several languages, HIPAA or PCI controls, custom voices, red-team testing, SLAs and a support retainer3 to 5 months

After launch, running costs are the per-minute charges above plus a support retainer of $1.5k to $8k per month for prompt, model and integration upkeep. Prompts and evaluation sets are yours; if you want to change platform later, they move with you.

Running cost per minute

StackCost per minute2,000 minutes a month20,000 minutes a month
Self-hosted pipeline$0.05 to $0.10$100 to $200$1,000 to $2,000
Managed platform (Vapi, Retell, Bland, ElevenLabs)$0.10 to $0.20$200 to $400$2,000 to $4,000 plus plan fees
Speech-to-speech model$0.30 or more$600 or more$6,000 or more
  • Call recording consent: several US states and most of Europe require all parties to consent to recording. The agent announces recording and we store the consent decision with the call.
  • Outbound rules: automated and AI-generated voice calls to consumers fall under the TCPA and FCC robocall rules. Outbound agents only call numbers with recorded consent, respect do-not-call lists and calling hours, and identify themselves as automated.
  • Payments by phone: card numbers spoken to an agent put the whole call path in PCI DSS scope. We use pause-and-resume recording and DTMF capture to a tokenisation provider, so the card never reaches the transcript or the model. See PCI DSS compliant development.
  • Healthcare: appointment lines that handle protected health information need a BAA from every vendor in the chain. Vapi sells a HIPAA add-on for $2,000 a month (checked October 2026); other platforms offer BAAs on enterprise plans; a self-hosted stack lets you sign BAAs directly with Twilio, Deepgram and your model provider. Our HIPAA guide covers the rest.
  • Data retention: transcripts and recordings are personal data. We set retention periods, redact card and health details from transcripts, and keep model providers on zero-retention terms where available.

Evaluation and QA of conversations

A voice agent that sounds fine in a demo can fail on the third real accent, a noisy car or a caller who answers two questions at once. We treat conversation quality as an engineering problem with a test suite.

  • Scripted test calls: a set of 100 to 300 scenarios per flow, including interruptions, wrong answers, background noise and silence, replayed automatically against every prompt or model change.
  • Transcript review: every production call is scored by a model against a rubric (resolved, accurate, polite, transferred correctly), and a sample is reviewed by a person each week.
  • Tool-call accuracy: bookings and lookups are checked against the source system, so the agent is scored on what it did, not on what it said.
  • Latency and audio metrics: time to first word, interruptions handled, drop-outs and transfer times, tracked per call and per provider.
  • Business metrics: containment rate, transfer rate, booking completion and caller ratings, reported weekly and compared with the human baseline.

How a project runs

  1. Discovery (one to two weeks, $3k to $8k): listen to recorded calls, pick the flows, map the systems, agree the latency budget, platform and compliance needs.
  2. Prototype (two weeks): a working agent on a test number for one flow, reviewed with your team on real scenarios.
  3. Build and integrate (three to eight weeks): tools, verification, handoff, SMS, recording and the evaluation suite.
  4. Soft launch: a share of calls routed to the agent, with transcripts reviewed daily and prompts tuned.
  5. Scale and support: full traffic, weekly QA review, monthly reporting and a retainer for ongoing changes.

Voice and conversational AI work we have shipped

  • Chat Your History: an AI personal historian that interviews people by web chat, SMS and recorded Twilio phone calls, then turns the transcripts into memoir chapters with streaming GPT-4 generation, so older users can take part in whatever channel suits them.
  • Treasure Island AI: a React Native consultation app with speech-to-text input and text-to-speech responses, where an OpenAI counsellor persona keeps context across a session and stays in character.
  • GovDoc AI: a retrieval pipeline on AWS Bedrock that grounds answers in long documents, the same approach we use to give voice agents accurate knowledge bases.

If you are weighing a voice agent against a text assistant, our AI chatbot development and AI agent development pages cover the alternatives, and pricing lists every published rate.

Next step

Tell us what you're building and get a written estimate.

A senior engineer replies within one business day. NDA on request.

FAQ

Questions we get asked a lot.

How much does AI voice agent development cost?

A pilot on one call flow costs $12k to $22k, a production agent with integrations $22k to $45k, and an enterprise deployment with a self-hosted stack and compliance controls $45k to $90k+. Running costs are usually $0.05 to $0.20 a minute.

Which is better, Vapi or Retell?

Both are good managed platforms. Vapi suits teams that want to choose and price each model separately; Retell has a strong flow builder and simple all-in pricing. For high volume or strict data control, a self-hosted stack on Twilio and Deepgram costs less per minute. We compare them in our Vapi vs Retell vs Bland post.

How fast does a voice agent respond?

We design to a budget of under about 800 ms from the end of the caller's sentence to the first word of the reply, and measure it on real phone calls, not in the browser. Speech-to-speech models are faster but cost more per minute.

Can the agent transfer to a human?

Yes. Warm transfer with a written summary is part of every build, along with rules for transferring automatically when the caller is frustrated, the agent has misunderstood twice, or the request is outside its scope.

Can a voice agent take card payments?

Yes, but card numbers must never reach the transcript or the model. We capture them by keypad to a tokenisation provider with recording paused, which keeps PCI DSS scope small.

Does the agent work in languages other than English?

Yes. Speech-to-text and voice models cover most major languages, and the agent can switch language on the first turn. Accuracy varies by language, so we test each one against real calls before launch.