Comparison
AI Voice Agent Pricing in 2026: A Transparent Cost Comparison
AI voice agent pricing 2026: the real all-in per-minute costs, the hidden fee layers, and how MCPDial, Vapi and Retell actually compare.
Published 9 min read

The honest answer: most AI voice agent quotes in 2026 are not comparable, because the advertised per-minute rate is only one of five layers you pay for. As of October 2026, one platform advertises entry rates from €0.05 per minute (Famulor's 2026 comparison of 10 platforms) — a number that tells you almost nothing about your invoice, because speech recognition, the language model, speech synthesis, telephony and the platform fee all sit underneath it. Budget on all-in cost per resolved call, not on the headline rate.
Key takeaways
- Every AI phone call carries four cost layers below the headline rate: STT, LLM, TTS and telephony. A platform that quotes one all-in number is easier to forecast than one that bills each layer separately.
- Per-minute pricing rewards short calls; per-resolution pricing rewards correct ones. Pick the model that matches the call type, not the one with the lowest advertised number.
- The cheapest way to lose money is a contract with uncapped concurrency, overage tiers you have not read, and no stated data residency.
- ROI is a resolved-call calculation, not a minutes calculation. If you cannot measure resolution, you cannot measure ROI.
- If your call volume is a handful of calls a day, you probably do not need this yet.
What are the real costs of AI voice agents in 2026?
The real cost of an AI voice agent in 2026 is the sum of five layers — speech-to-text, language model, text-to-speech, telephony, and the platform or orchestration fee — multiplied by the minutes you actually use, plus any fixed monthly platform charge. Advertised per-minute rates typically cover only one or two of those layers.
That is why two quotes of "€0.10 per minute" can produce invoices that differ by a factor of three. The published rate is a marketing surface. The layers are the bill.
The useful discipline is to ask for the same call priced twice: once as a ten-second call that gets transferred to a human, and once as a four-minute call that books an appointment. The first exposes your minimum billing increment. The second exposes your real average. Most surprises live in that gap.
The 4 hidden cost layers of every AI phone call

Four layers sit under the headline rate: speech recognition, the language model, speech synthesis, and telephony. Each is metered differently, and each has a failure mode that costs you money without producing a second of useful conversation.
- Speech recognition. Billed per audio minute, not per word. Hold music, silence and background noise all bill.
- Language model. Billed per token, in and out. Long system prompts and verbose tool schemas are charged on every single turn, not once per call.
- Speech synthesis. Billed per character or per audio second. Reading a long confirmation back to a caller is not free.
- Telephony. Billed by whoever carries the call, usually per minute with a separate inbound or outbound rate, and often with a monthly number rental on top.
Two more costs hide in the seams. Concurrency — how many calls can run at once — is frequently a tier gate rather than a metered cost, so a spike can simply fail rather than bill. And tool calls into your CRM or calendar can bill at the platform level even when the conversation itself is short. If your agent books into a calendar, read how voice AI connects to your CRM and calendar before you sign anything, because that is where the per-call cost quietly doubles.
Per-minute vs per-resolution AI pricing: which model should you choose?

Per-minute pricing suits short, high-volume, predictable calls. Per-resolution pricing suits longer calls where the outcome matters more than the duration — bookings, claims, order changes. Per-call flat pricing sits between the two and is easiest to forecast, but it punishes you on long tail calls.
| Model | What you pay for | Best for | Main risk |
|---|---|---|---|
| Per minute | Talk time, often with a minimum increment | High-volume, short interactions | Long calls and silence bill the same as useful ones |
| Per call | Each connected call, flat | Predictable call types | Long calls subsidised by short ones |
| Per resolution | Only calls the agent completes | Bookings, order changes, claims | You must define "resolved" in writing |
The trap with per-resolution is definitional. "Resolved" has to mean something you can audit — an appointment in the calendar, a confirmation number, a ticket closed — not "the caller did not ask for a human." Get that sentence into the contract.
What are AI call center pricing benchmarks in 2026?
Published benchmarks for AI call centre pricing are unreliable, and I would not budget from them. The sources that publish platform comparisons differ on what they count as a minute, whether telephony is included, and whether failed calls bill at all — so the ranges they produce are not measuring the same thing.
What you can compare honestly is the shape of a quote. A benchmark you can act on has four numbers: the all-in rate per minute, the minimum billing increment, the monthly platform floor, and the concurrency ceiling. Ask four vendors for those four numbers and you will learn more than from any published range.
AI voice agent pricing compared: MCPDial vs Vapi vs Retell
MCPDial publishes one all-in per-minute rate covering streaming voice, telephony and tool calls, with CH/EU/US in-region processing. Vapi is the better choice when you want to assemble and tune every component yourself. Retell publishes a detailed cost breakdown and platform comparison, which makes them the easier starting point for research.
| Dimension | MCPDial | Vapi | Retell |
|---|---|---|---|
| Pricing model | All-in per minute, one line item (MCPDial) | Not published in my sources — verify directly | Publishes a full cost breakdown and platform comparison (Retell's 2026 breakdown) |
| Layers billed separately | No — STT, LLM, TTS and telephony are inside the rate (MCPDial) | Not published in my sources | Their breakdown separates platform, model and telephony layers (Retell) |
| Data residency | CH, EU and US in-region processing (MCPDial) | Not published in my sources | Not published in my sources |
| Compliance enforced on the call | AI disclosure on every call, consent gates, do-not-call and calling hours enforced by the runtime (MCPDial) | Not published in my sources | Not published in my sources |
| Best for | Teams that want a forecastable all-in rate and compliance handled by the runtime | Developers who want to configure every component themselves | Buyers who want to read a published breakdown before they talk to sales |
I have deliberately left cells blank rather than guess. A comparison table with invented numbers is worse than no table.
Is Twilio AI voice agent pricing the same as platform pricing?
No. Telephony is a separate layer from the agent platform, so a carrier's per-minute rate and a platform's per-minute rate are two different invoices for the same call. I am not quoting a carrier rate here because I cannot source one as of October 2026, and a stale telephony number is worse than none.
What matters is whether your platform resells telephony inside its rate or passes it through. Pass-through is cheaper on paper and harder to forecast, because carrier rates move. Bundled is duller and easier to defend in a budget meeting. For most teams putting an agent on a real phone line, bundled wins.
How do you calculate the true ROI of an AI voice agent?
ROI is resolved calls multiplied by value per resolution, minus all-in cost, divided by all-in cost. The all-in cost must include the monthly platform floor and telephony, not just the per-minute rate. If you cannot count resolutions, you are measuring activity, not return.
A workable formula:
- All-in cost per call = (all-in per-minute rate × average handle time including wrap-up) + (monthly platform fee ÷ calls per month).
- Value per resolution = bookings or tickets completed × your value per completion × your historical completion rate.
- ROI = (value − cost) ÷ cost.
The comparison that matters is against your current answer, not against zero. If you are weighing this against a human answering service, the honest arithmetic is in AI receptionist vs human answering service — and there are call types where the human still wins.
What hidden fees should you look for in AI voice agent contracts?
The fees that hurt are rarely hidden in the rate. They are hidden in the definitions: what counts as a billable minute, what counts as a resolution, and what happens when you exceed a tier. Read those three sentences before you read the price.
- Minimum billing increments. A 30-second minimum turns a hang-up into a paid call.
- Concurrency ceilings. The number that decides whether your busiest hour works.
- Overage tiers. Ask for the rate after the included volume, not before.
- Number porting and rental. Monthly, per number, and easy to overlook.
- Data residency surcharges. Processing in-region is a feature, and sometimes a line item. Ask which.
- Termination and minimum commitment. How many months are you buying?
What should you demand from a transparent AI voice agent provider?
Demand four things in writing: an all-in rate, a minimum billing increment, a definition of a resolved call, and a stated processing region. If a provider will not put those four on paper, the pricing is not transparent — it is just unpublished.
Compliance belongs on that list too. Consent capture, do-not-call handling, calling hours and AI disclosure are product features, not legal afterthoughts, and they should be enforced by the runtime rather than by your prompt. If you operate in Switzerland or the EU, the EU AI Act and GDPR guide for Swiss voice AI deployments covers what the disclosure and residency requirements actually look like in practice.
MCPDial's pricing explained: all-in per-minute costs for streaming voice
MCPDial bills one all-in rate per minute. Streaming voice, telephony, tool calls into your systems, AI disclosure on every call, and consent, do-not-call and calling-hour gates are inside that rate rather than stacked on top of it. Processing runs in-region in Switzerland, the EU or the US.
That structure exists because the alternative is unforecastable. When four layers bill separately, your finance team cannot approve a budget, and your first month's invoice becomes the real pricing page. One rate, one increment, one monthly floor is a duller pitch and a better contract.
It is not the right pick for everyone. If you want to swap components, tune the pipeline and own the orchestration, a developer-first platform such as Vapi is the better fit. If you want to read a detailed published breakdown before you talk to anyone, start with Retell's. If you want a forecastable all-in number and compliance handled by the runtime, see MCPDial's pricing.
FAQ
How much does an AI phone agent cost per minute in 2026?
Advertised entry rates start as low as €0.05 per minute (Famulor's 10-platform comparison), but that figure rarely reflects all-in cost. Once speech recognition, the language model, speech synthesis, telephony and the platform fee are counted, the effective rate is usually higher and varies with call length.
What are the hidden fees in AI voice agent contracts?
The common ones are minimum billing increments, concurrency ceilings that force a tier upgrade, overage rates above your included volume, per-number rental, and data residency surcharges. None of these are usually concealed — they are just defined in the terms rather than the pricing table.
Is per-minute or per-resolution pricing better?
Per-minute pricing is better for short, high-volume calls where duration is predictable. Per-resolution pricing is better for bookings, claims and order changes, where a long call that succeeds is worth more than a short one that fails. Whichever you choose, get the definition of a billable minute or a resolution in writing.
Can I compare AI voice agent pricing across platforms at all?
Only if the quotes share the same four numbers: all-in rate per minute, minimum billing increment, monthly platform floor, and concurrency ceiling. Published comparisons differ on whether telephony and failed calls are included, so their ranges are not measuring the same thing.
Do I need an AI voice agent yet?
If you handle a handful of calls a day, no. The fixed platform cost and the setup work will not pay back. This becomes worth doing when missed calls are a measurable revenue leak — after hours, during peak, or on a line that rings out more often than you would like to admit.