AI Economics - July 20, 2026 - 10 min read

The Unit Economics of Real-Time Voice AI

Voice-model prices are falling, but production cost remains variable. Lenders should budget the entire resolution path, including duration, tool use, retries, transfers, controls, and failed outcomes.

By Muthukkumaran K

Last reviewed July 20, 2026

Voice waveforms accumulating variable cost through metered chambers before reaching a balanced resolution outcome

Real-time voice AI is often sold with a simple comparison: an automated minute costs less than a human minute.

That may be true and still produce a poor business case.

A collections call is not valuable because it occupied a channel cheaply. It is valuable when it creates a valid resolution: payment, kept promise, verified dispute, hardship path, useful commitment, or correct escalation. The economic denominator should therefore be the completed outcome, not the minute.

The right metric is cost per valid resolution under control, not cost per automated conversation.

Unit prices are falling; exposure is becoming more complex

As of July 2026, provider price lists show a widening range between smaller real-time models and more capable reasoning models. OpenAI lists audio-token rates for its smaller real-time model below its larger model, while both can also incur text, reasoning, and tool-related usage. Google prices some speech-generation models by audio tokens and states that its Gemini TTS audio output uses 25 tokens per second.

These prices will change. The durable point is how they are metered.

Some components charge by audio duration, some by audio token, some by text token or character, some by model invocation, some by telephony minute, and some by concurrent session. Tool calls can trigger additional model and infrastructure usage. A human transfer can add contact-centre cost while preserving all automated cost already incurred.

Falling rates do not remove variability. They make previously uneconomic interactions feasible, which can increase total volume and create a larger uncontrolled bill if campaign design is weak.

The seven-part cost stack

An all-in model should separate at least seven components.

1. Telecom and connection

Outbound origination, carrier charges, number reputation, recording, SIP infrastructure, and failed connection attempts exist before the AI produces a useful word. Geography and carrier policy can materially change this layer.

2. Listening

The system pays to receive, buffer, recognise, or encode customer audio. Silence, background speech, hold time, repetition, and long explanations can all consume resources even when the final action is simple.

3. Reasoning and context

The agent needs policy instructions, account context, conversation history, and retrieved information. Longer conversations can repeatedly carry forward prior context. More capable reasoning may improve difficult turns but can increase both latency and output usage.

4. Speaking

Generated audio is often priced differently from input. A verbose agent is therefore costly twice: it consumes generation and extends telecom duration. Concise, understandable speech is an economic control.

5. Tools and data

Balance enquiry, identity checks, payment-link creation, offer lookup, promise capture, CRM update, and fraud or vulnerability screening call other services. Their cost may be small individually but becomes material at portfolio scale.

6. Control and observability

Recording, transcription, policy evaluation, redaction, audit logs, quality sampling, model monitoring, secure storage, incident response, and vendor assurance are not optional overhead for a regulated deployment. They are part of the product cost.

7. Exception and remediation

Human transfer, callback, complaint investigation, wrong-party correction, failed promise, duplicate contact, and reconciliation repair can dominate the economics of a weak design.

From call cost to resolution cost

The bill continues after inference when the outcome is incomplete or incorrect.

ConnectListen and reasonActControlResolve or remediate

Why conversation length is unpredictable

A scripted notification has a bounded duration. A real conversation does not.

The customer may ask for the amount again, change language, search for a reference number, challenge the account, explain a personal circumstance, lose network connectivity, interrupt the agent, or wait while a payment tool responds. Each event changes duration and may add context, generation, and tool cost.

The distribution matters more than the average. A portfolio can show a three-minute mean while a small tail of long, unresolved calls consumes a disproportionate share of cost and creates conduct risk.

Track at least:

  • median, 90th, and 99th percentile call duration;
  • customer-speaking, agent-speaking, and silence time;
  • model turns and repeated turns;
  • tool calls, retries, and timeouts;
  • context and audio usage per resolved and unresolved call;
  • transfer and callback rate;
  • cost by language, intent, segment, and outcome.

A finance-readable equation

For a defined segment, start with:

Expected automated cost per account

= connection attempts

  • connected duration cost

  • model input, output, and tool cost

  • control and storage allocation

  • probability-weighted human exception cost

  • expected remediation cost

Then calculate:

Cost per valid resolution

= total segment cost / incremental valid resolutions

Finally compare it with the next-best treatment, not with a theoretical fully loaded human average.

For an early-bucket customer likely to self-pay, the alternative may be SMS, WhatsApp, a payment link, or no immediate action. Voice can be more expensive than all four. For a customer who repeatedly fails a digital payment journey, a short voice interaction may be cheaper than field allocation and more effective than another message.

Native speech-to-speech versus a cascaded stack

There are two broad architectures.

Cascaded systems separate speech recognition, text reasoning, and speech generation. They make intermediate text visible, allow specialised language components, and can simplify policy testing and provider substitution. They also add latency and more integration points.

Native speech-to-speech systems can improve turn-taking, prosody, and interruption handling by processing audio more directly. They may reduce orchestration complexity but can make it harder to inspect exactly what was understood and why a spoken response occurred.

The choice is not only technical. It changes cost observability, auditability, language strategy, provider concentration, and failure containment.

Current full-duplex benchmarks evaluate latency, interruption, tool use, and realistic disfluency precisely because one architecture does not dominate every production condition. A lender may use native voice for a bounded inbound journey and retain a cascaded architecture where exact transcript and policy traceability matter more.

Eight practical cost controls

1. Select before dialling

Do not spend voice inference on customers who are likely to self-pay through a cheaper path, have already paid, are under active dispute, or should be suppressed.

2. Bound the job

Give the agent a limited objective and action set. Open-ended "resolve the account" prompts create long interactions and unpredictable behaviour.

3. Make the opening useful

After proportionate verification, move quickly to the purpose and available action. A collections pilot found that offers placed late in IVR flows were often never heard. Duration before value is both customer friction and cost.

4. Use tiered models

A smaller model may handle identification, due-amount explanation, callback scheduling, and payment-link delivery. Escalate to a more capable model or human only for defined complexity. Do not pay premium reasoning rates for every greeting.

5. Keep policy outside the conversation memory

Retrieve the smallest relevant policy and account context for the current state. Summarise prior turns into structured fields rather than carrying unnecessary transcript indefinitely.

6. Stop unproductive loops

After repeated recognition or tool failure, transfer, schedule a callback, or offer a digital route. The third apology rarely creates more value than the first two.

7. Design silence and timeout rules

Distinguish natural thinking pauses from dead air, but cap indefinite listening. Allow customers to resume through a safer channel instead of forcing completion in one session.

8. Reconcile quickly

Fresh payment and contact outcomes prevent duplicate calls. The cheapest inference is the call correctly suppressed because the customer has already acted.

Public case studies: useful evidence, incomplete economics

DoorDash's published voice-support case reports hundreds of thousands of calls per day, response latency of 2.5 seconds or less, and a 50% reduction in generative AI application development time. It also reports earlier IVR improvements in agent transfer and first-contact resolution.

Those are relevant production signals. They do not disclose a complete cost per resolution, the long-tail duration distribution, or collections conduct outcomes. Vendor and customer case studies should be treated as architecture evidence, not a substitute for an institution's own control-group economics.

The same discipline applies to model pricing. A current price page is an input to a pilot model, not a five-year forecast. Contracts should address price change, model retirement, regional availability, data use, audit access, service levels, and exit.

The BIS identifies third-party dependency, confidentiality, model risk, and reputation as core AI governance concerns. RBI's IT outsourcing directions similarly preserve the regulated entity's access, monitoring, audit, security, continuity, and exit responsibilities. Cost optimisation that creates provider lock-in or weakens evidence is not value optimisation.

The investment decision

Approve a voice pilot only when the team can state:

  • the segment and next-best treatment;
  • the valid resolution being purchased;
  • the expected duration and tail limits;
  • the model, telecom, tool, control, and exception cost;
  • the human handoff and remediation design;
  • the measurement control;
  • the stop, revise, and scale thresholds.

The purpose of cost modelling is not to predict every invoice line perfectly. It is to expose which operating assumptions must be true for value to survive production.

Primary sources

  1. GPT-Realtime-2.1 mini model and pricing OpenAI, accessed July 2026
  2. GPT-Realtime-2.1 model and pricing OpenAI, accessed July 2026
  3. Text-to-Speech pricing and audio-token metering Google Cloud, accessed July 2026
  4. Full-Duplex-Bench-v3 for realistic voice-agent tool use Research paper, 2026
  5. DoorDash generative AI contact centre case study AWS and DoorDash
  6. Governance of AI adoption in central banks Bank for International Settlements, 2025
  7. Master Direction on Outsourcing of Information Technology Services Reserve Bank of India, 2023

Regulatory references are an operating interpretation, not legal advice. Requirements should be confirmed for the institution, activity, and jurisdiction at deployment.

Related Insights