tuesday, october 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · working with ai

Putting an AI agent on your phone lines

Price a voice agent honestly, scope it to one call type, meet the disclosure rules that apply to you, and run a pilot that tells you whether to keep it.

Published 2026-09-04 · Updated 2026-09-04 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-04 · Rami
on this page · 0 / 0 checked

A phone line is the least forgiving channel you own. A caller cannot skim past a bad answer, cannot reread the sentence, and cannot open a second tab while the thing on the other end works out what to say. Putting a model on that line stopped being an engineering project some time ago. OpenAI’s Realtime API now accepts calls directly over SIP: you point a trunk from a telephony provider at sip:$PROJECT_ID@sip.api.openai.com;transport=tls, a realtime.call.incoming webhook fires when someone dials, and your code either accepts the call or rejects it [5].

Because the trial is cheap, the trial is not the decision. This guide is for someone with a real line and repeat call types behind it: bookings, reschedules, order status, after-hours coverage that currently goes to voicemail. It is not for a business whose calls are mostly upset people, regulated advice, or one-off negotiations where the right answer depends on the history you have with that specific caller. On those calls a fluent voice being confidently wrong is the expensive outcome, and no amount of pilot discipline fixes it.

What a phone agent costs to run

Two vendors price this two different ways, and the difference matters more than the headline number.

Google publishes per-minute audio rates. Its gemini-3.1-flash-live-preview model costs $0.005 per minute of audio in and $0.018 per minute of audio out [1]. Input runs for the whole call, since the model is listening the entire time; output only runs while the agent is speaking. On a 5-minute call where the agent talks for 2 of those minutes, that is 5 × $0.005 plus 2 × $0.018, or about $0.061 of model time. Even in the impossible case where the agent talks non-stop for the full call, you are at $0.023 a minute.

OpenAI prices the same job by token. gpt-realtime-2.1 is $32 per 1M audio input tokens and $64 per 1M audio output tokens; the smaller gpt-realtime-2.1-mini is $10 and $20 [2]. Cached audio input is much cheaper, at $0.40 and $0.30 per 1M tokens respectively [2]. Token pricing is fine once you have data, but it makes the per-call number hard to predict in advance, which is an argument for running the pilot on a metered account and reading the invoice rather than trusting your own estimate.

Then telephony, which people forget. On Twilio, a local US number is $1.15 a month, inbound minutes on it are $0.0085, and outbound local minutes are $0.0140 [3]. Add that to the model and a 5-minute inbound call lands around 10 cents.

The packaged alternative is worth pricing next to it. ElevenLabs lists $0.080 a minute for agent calls, the same rate on every tier, with the plans differing mainly in included minutes and how many calls can run at once: 4 concurrent calls on the free plan, 40 on the $990-a-month Business plan [4]. Roughly 8 cents against roughly 2 is not a trick. The higher price buys the pieces you would otherwise assemble and keep working, and the concurrency ceiling is the number to check before a Monday morning rush finds it for you.

The calls that survive contact with a real caller

The call types that work share a shape: high volume, narrow vocabulary, and a definition of done that a computer can check.

Appointment handling is the cleanest of them. Booking, rescheduling and cancelling have bounded language, an obvious success test, and a real system behind them to write to. After-hours coverage is the second, and it is the easiest to justify, because the thing you are replacing is not a person. It is a voicemail box. An agent that answers at 2am, takes structured details and books a callback for the morning beats voicemail without needing to be impressive.

Order status, opening hours and lookups against documents you control also clear the bar. What does not clear it: anything where the caller is already angry, anything with regulatory language attached, anything where the correct answer depends on judgment about the relationship. Those calls are not a tuning problem. They are the reason you keep a human path.

There is a second cut to make, on what the agent can do rather than what it can say. An agent that can only read your FAQ can embarrass you. An agent that can move bookings, issue refunds or send emails on your letterhead can cost you money, and it will do so at the speed of a phone system. Scope write access to the minimum the pilot needs and add capabilities one at a time.

Outbound is the settled part. In a Declaratory Ruling adopted on 2 February 2024 in CG Docket No. 23-362, the FCC held that voices generated by AI are “artificial” under the Telephone Consumer Protection Act, on the reasoning that “a person is not speaking them”, and that “callers must obtain prior express consent from the called party before making a call that utilizes artificial or prerecorded voice simulated or generated through AI technology” [7]. If your plan involves the agent dialling out to customers, that consent is the project, and the agent is a detail inside it.

Answering your own line is a different matter, because the consent obligation in that ruling attaches to making a call [7]. This is why so many sensible first deployments are inbound only: the caller chose to ring you.

Disclosure is separate again from consent, and it does not follow the same border. Article 50(1) of the EU AI Act requires that AI systems “intended to interact directly with natural persons are designed and developed in such a way that the natural persons concerned are informed that they are interacting with an AI system”, and that obligation has applied since 2 August 2026 [8]. Note where it sits: on the provider of the system. If you are buying rather than building, the useful move is to make the vendor show you how their product satisfies it, in writing, before you sign.

Even where no rule reaches you, decide the disclosure question once and on purpose rather than discovering your position from a review. A caller who works out halfway through that they have been talking to software tends to be angrier than one who was told at hello.

Where the audio goes after the caller hangs up

Calls carry names, addresses, appointment reasons, sometimes card fragments read aloud. So the retention question is not paperwork.

Read the vendor’s actual policy rather than the marketing page. OpenAI states that “data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)”, and that abuse monitoring logs “are generated for all API feature usage and retained for up to 30 days, unless longer retention is required by law” [6]. Where you qualify for Zero Data Retention, that control excludes customer content from those logs and forces the store parameter to be treated as false even if your code asks otherwise [6].

That is one vendor’s default, and it is the smallest part of your exposure. Your own stack keeps far more: the telephony provider’s recordings, transcripts wherever you built the agent, the summary you wrote into your CRM at the end of the call. Write down the retention window for each of those four places, because the one you forget is the one that shows up in a subject access request. And check the recording-consent rule that applies where your callers are before you switch recording on by default.

The per-minute price is not the cost

The invoice is the easy half of the arithmetic and the half everyone runs. The other half is what a mishandled call costs you, multiplied by how often it happens.

Both of those are your numbers, not a vendor’s. A lost booking is worth what a booking is worth to you. A caller who never rings again takes their lifetime value with them. Set that against 10 cents a call and the per-minute price stops being the thing you are deciding on. The handoff rule becomes the part worth spending a week on, and narrow scope starts to pay, because a bounded call type is the only way to get the mishandling rate low enough that the infrastructure line matters at all.

A pilot that tells you something

Route one call type, not one line. After-hours booking is the standard choice. Everything else stays exactly where it is, and you give the pilot two weeks.

Then listen, properly. Not to the calls a dashboard flags as successful, which is a sampling method designed to make you feel good. Pull 20 recordings at random and hear them end to end. You are listening for three shapes: the agent inventing an answer your documents do not contain, the agent failing to notice it should hand off, and the caller repeating themselves with the pitch rising. Vendor benchmark numbers measure none of these, because they were not measured on your callers, your accents or your documents.

Measure four boring things. Completion rate on the narrow task. Handoff rate. Callback rate within 48 hours, which is the honest one, because a caller who rings back was not actually helped. And explicit complaints, counted by hand. If those hold for two weeks, widen by one call type and repeat.

checklist
Before you route a real call to an agent
0 of 8 · saved in this browser only

Two weeks of that produces the one figure no vendor page can give you: how often your own callers get mishandled on your own line, on your own call type. That is the input worth having before you widen anything.

calculator
True cost of one answered call
— $ / call

Run time plus the expected cost of getting it wrong. The defaults are placeholders, not findings; replace all four with your own. Computed in the page; nothing is sent anywhere.

Run it once before the pilot with a mishandling rate you would defend out loud, and again afterwards with the rate you actually measured. The distance between those two runs is what the two weeks bought you.

What still goes wrong

Prices move on a schedule you do not control. Google’s own pricing page already lists a scheduled increase, with Gemini 3.8 Flash text input at $0.75 per 1M tokens through 31 December 2026 and $1.50 from 1 January 2027 [1]. Model names churn faster than that. Anything you build should survive swapping the model underneath it, and any budget you approve should be re-checked against the pricing page every quarter rather than the day you signed.

The failures that hurt in production are rarely the ones a demo shows. A transfer that fires correctly into a queue nobody staffs is worse than no transfer, because the caller heard a promise. Accents, hold music bleeding in, two people talking in the room and a caller who interrupts mid-sentence all degrade a system that sounded fine in a quiet office. Concurrency limits bite on exactly the morning you would least like to lose calls [4]. And token-based billing means a pilot with unusually chatty callers can produce an invoice you would not have approved in advance [2].

The honest limit is that none of this is knowable from a vendor page, including the ones cited here. Prices, SIP endpoints and retention defaults you can read. Whether your callers will accept an agent on your particular line, for your particular call type, you can only find out by putting it on a small slice of real calls and listening to all of them.

sources
  1. 01Google — Gemini API pricingai.google.dev
  2. 02OpenAI — API pricingdevelopers.openai.com
  3. 03Twilio — Voice pricing, United Statestwilio.com
  4. 04ElevenLabs — Agents pricingelevenlabs.io
  5. 05OpenAI — Realtime API over SIPdevelopers.openai.com
  6. 06OpenAI — Your datadevelopers.openai.com
  7. 07FCC — Declaratory Ruling FCC 24-17, CG Docket No. 23-362docs.fcc.gov
  8. 08EU AI Act — Article 50, transparency obligationsartificialintelligenceact.eu
next guide
Why a smaller prompt can cost you more
9 min · verified 2026-09-04
related guides