Cost Per Output

OpenAI Realtime API pricing per minute

OpenAI bills the Realtime API in tokens, not minutes: gpt-realtime-2 charges $32 per million audio tokens in and $64 out, the mini models $10 and $20. By OpenAI's own count of 600 audio tokens a minute for the caller and 1,200 for the agent, a minute of the agent speaking costs $0.0768 in output and a minute of the caller $0.0192 in input on gpt-realtime-2, each counted once; every turn re-sends the conversation, so calls cost more. GPT-Live, OpenAI's newer voice model, costs $0.05 a minute, with the backend model billed separately.

gpt-realtime-2, half the call$0.048/minAt least: each second of audio counted once
gpt-realtime-mini, half the call$0.015/minAt least: each second of audio counted once
gpt-live-1 voice session$0.05/minBilled per second; backend model extra

Realtime models: tokens and minutes

ModelAudio in / cached / out per 1MText in / out per 1MAgent speaking, a minuteCaller speaking, a minuteHalf each
gpt-realtime-2.1$32 / $0.40 / $64$4 / $24$0.0768$0.0192$0.048
gpt-realtime-2.1-mini$10 / $0.30 / $20$0.60 / $2.40$0.024$0.006$0.015
gpt-realtime-2$32 / $0.40 / $64$4 / $24$0.0768$0.0192$0.048
gpt-realtime-1.5$32 / $0.40 / $64$4 / $16$0.0768$0.0192$0.048
gpt-realtime-mini$10 / $0.30 / $20$0.60 / $2.40$0.024$0.006$0.015
gpt-realtime$32 / $0.40 / $64$4 / $16$0.0768$0.0192$0.048

Per minute from OpenAI's statement that user audio is 1 token per 100 ms (600 a minute) and the assistant's 1 token per 50 ms (1,200 a minute), each token counted once. Instructions, text, tool calls and the conversation re-sent on later turns come on top.

Why a call costs more than the minutes of audio

So the figures above are a floor, not an estimate of a whole call.

Price a call

Set how much of each call the agent speaks; the model's audio tokens are priced at OpenAI's rates, counted once. GPT-Live is in the menu too, with its backend model as your own per-minute figure.

OpenAI Realtime API: your stack

1,000 minutes a month on OpenAI Realtime API: at least $56.50, or $0.0565 a minute on average. This counts each second of audio once; real calls cost more, because every turn re-sends the conversation.

PartPer minutePer month
Model
gpt-realtime-2
at least $0.048$48.00
Phone line
Twilio local number, inbound
$0.0085$8.50
Total$0.0565$56.50

GPT-Live: priced by the minute

GPT-Live separates the voice conversation from the backend that reasons and runs tools. The voice session costs $0.05 a minute, $50.00 for 1,000 minutes, and the backend is billed at its model's token prices.

What Retell AI charges for the same models

Retell AI resells OpenAI's realtime models per minute of call, the only per-minute price for them among the platforms we track:

Model on RetellPer minute1,000 minutes
GPT Realtime 1.5$0.345$345
GPT Realtime$0.345$345
GPT Realtime mini$0.07$70.00
GPT Realtime 2.1$0.38$380
GPT Realtime 2$0.38$380
GPT Realtime 2.1 mini$0.07$70.00

More on Retell's prices: Retell AI pricing.

The phone line

OpenAI charges nothing for bandwidth or connections; the phone line is your carrier's. Twilio's US list price for an inbound call on a local number is $0.0085 a minute.

Against voice agent platforms

PlatformCheapest picksDearest picks1,000 min a month10,000 min a month
Vapi$0.0681
AssemblyAI, Google and Cartesia
$0.1313
OpenAI and ElevenLabs
$68.10$681
Retell AI$0.0716
GPT 5 nano and Retell Platform Voices
$0.795
GPT 6 Astra, fast tier and ElevenLabs v3
$71.60$716
Deepgram Voice Agent API$0.075
the Standard tier
$0.163
the Advanced tier
$75.00$750
Bland$0.12
the Build plan
$0.14
the Start plan
$140
Start
$1,400
Start

Per minute of conversation before the phone line, at each vendor's own prices: the platform fee with the cheapest, then the dearest, model and voice it lists, no add-ons. Monthly figures use the cheapest picks on the cheapest plan for that volume.

PlatformPer minuteBilled on top
ElevenLabs Agents$0.08Language model
OpenAI GPT-Live$0.05Backend model and tools
OpenAI Realtime APITokensEverything: audio and text tokens, with the conversation re-sent each turn
Twilio ConversationRelay$0.07Your language model

Every platform side by side, with a calculator: AI voice agent cost per minute. OpenAI's transcription models on their own: Whisper API pricing; its text-to-speech: OpenAI text-to-speech pricing.

Quick answers

How much does the OpenAI Realtime API cost per minute?

At least $0.048 a minute on gpt-realtime-2 when the agent speaks half the time, counting each second of audio once at OpenAI's token prices. Real calls cost more, because each turn sends the whole conversation back in as input (cached input is $0.40 per million audio tokens). Retell AI charges $0.38 a minute for GPT Realtime 2.

How much does gpt-realtime-mini cost per minute?

At least $0.015 a minute when the agent speaks half the time: $0.024 for a minute of the agent's audio and $0.006 for a minute of the caller's, each counted once. Retell AI charges $0.07 a minute for it.

What does GPT-Live cost?

$0.05 a minute of voice session, billed per second with no rounding up. The backend model and tools that do the reasoning are billed separately at their token prices, and the session bills while both sides are silent.

Does the Realtime API charge for silence?

Not by itself: with voice activity detection on, empty audio doesn't count as input tokens. GPT-Live is different: its session bills for as long as it is open.

Is there a per-minute price for the Realtime API?

Not for the speech-to-speech models, which OpenAI prices in tokens; GPT-Live is priced per minute. Platforms that resell the realtime models set their own: GPT Realtime 1.5 $0.345, GPT Realtime $0.345, GPT Realtime mini $0.07, GPT Realtime 2.1 $0.38, GPT Realtime 2 $0.38 and GPT Realtime 2.1 mini $0.07 a minute on Retell AI.

Sources

Read on 7 October 2026:

Show the quoted text
Realtime and audio generation models Prices per 1M tokens unless noted. ### Grouped Pricing Table data | Model | Modality | Input | Cached input | Output / cost |
developers.openai.com/api/docs/pricing.md
| gpt-realtime-2.1 | Audio | $32.00 | $0.40 | $64.00 | | gpt-realtime-2.1 | Text | $4.00 | $0.40 | $24.00 | | gpt-realtime-2.1 | Image | $5.00 | $0.50 | - | | gpt-realtime-2.1-mini | Audio | $10.00 | $0.30 | $20.00 | | gpt-realtime-2.1-mini | Text | $0.60 | $0.06 | $2.40 | | gpt-realtime-2.1-mini | Image | $0.80 | $0.08 | - | | gpt-realtime-2 | Audio | $32.00 | $0.40 | $64.00 | | gpt-realtime-2 | Text | $4.00 | $0.40 | $24.00 | | gpt-realtime-2 | Image | $5.00 | $0.50 | - | | gpt-realtime-1.5 | Audio | $32.00 | $0.40 | $64.00 | | gpt-realtime-1.5 | Text | $4.00 | $0.40 | $16.00 | | gpt-realtime-1.5 | Image | $5.00 | $0.50 | - | | gpt-realtime-mini | Audio | $10.00 | $0.30 | $20.00 | | gpt-realtime-mini | Text | $0.60 | $0.06 | $2.40 | | gpt-realtime-mini | Image | $0.80 | $0.08 | - | | gpt-realtime | Audio | $32.00 | $0.40 | $64.00 | | gpt-realtime | Text | $4.00 | $0.40 | $16.00 | | gpt-realtime | Image | $5.00 | $0.50 | - |
developers.openai.com/api/docs/pricing.md. The realtime models' rows of the table: model, modality, then input, cached input and output prices per million tokens.
Audio tokens in user messages are 1 token per 100 ms of audio, while audio tokens in assistant messages are 1 token per 50ms of audio.
developers.openai.com/api/docs/guides/realtime-costs.md
The entire conversation is sent to the model for each Response. The output from a turn will be added as Items to the server Conversation and become the input to subsequent turns, thus turns later in the session will be more expensive.
developers.openai.com/api/docs/guides/realtime-costs.md
A Response can be created manually or automatically if voice activity detection (VAD) is turned on. VAD will effectively filter out empty input audio, so empty audio doesn't count as input tokens unless the client manually adds it as conversation input.
developers.openai.com/api/docs/guides/realtime-costs.md
Given the complexity in Realtime API token usage it can be difficult to estimate your costs ahead of time. A good approach is to use the Realtime Playground with your intended prompts and functions, and measure the token usage over a sample session.
developers.openai.com/api/docs/guides/realtime-costs.md
There is no cost currently for network bandwidth or connections.
developers.openai.com/api/docs/guides/realtime-costs.md
Aside from conversational Responses, the Realtime API bills for input transcriptions, if enabled.
developers.openai.com/api/docs/guides/realtime-costs.md
Number type To make calls To receive calls Refer from Twilio Local calls $0.0140 / min $0.0085 / min Toll-free calls $0.0140 / min $0.0220 / min Browser/app $0.0040 / min $0.0040 / min SIP interface $0.0040 / min $0.0040 / min
twilio.com/en-us/voice/pricing/us
[GPT-Live 1](https://developers.openai.com/api/docs/models/gpt-live-1) voice sessions are billed per second, without rounding up to a whole minute. Backend model and tool usage is charged separately. ### Pricing Table data | Model | Price per minute | | --- | --- | | gpt-live-1 | $0.05 |
developers.openai.com/api/docs/pricing.md
Voice sessions cost $0.05 per minute, billed per second. Backend model and tool usage is billed separately.
developers.openai.com/api/docs/models/gpt-live-1.md
Active session time includes time when the user speaks, the assistant speaks, both are silent, or the backend is working.
developers.openai.com/api/docs/guides/realtime-costs.md
A `POST /v1/live/sessions` request to create a WebRTC session bills 15 seconds of voice duration while the session initializes. That amount is credited against duration charges once the session starts running.
developers.openai.com/api/docs/guides/realtime-costs.md
Speech-to-speech LLM Notes Pricing GPT Realtime 1.5 $0.345/minute GPT Realtime $0.345/minute GPT Realtime mini $0.07/minute GPT Realtime 2.1 $0.38/minute GPT Realtime 2 $0.38/minute GPT Realtime 2.1 mini $0.07/minute
retellai.com/pricing

Prices on this page were checked against the official pricing pages on .