OpenAI Realtime API pricing per minute
OpenAI bills the Realtime API in tokens, not minutes: gpt-realtime-2 charges $32 per million audio tokens in and $64 out, the mini models $10 and $20. By OpenAI's own count of 600 audio tokens a minute for the caller and 1,200 for the agent, a minute of the agent speaking costs $0.0768 in output and a minute of the caller $0.0192 in input on gpt-realtime-2, each counted once; every turn re-sends the conversation, so calls cost more. GPT-Live, OpenAI's newer voice model, costs $0.05 a minute, with the backend model billed separately.
Realtime models: tokens and minutes
| Model | Audio in / cached / out per 1M | Text in / out per 1M | Agent speaking, a minute | Caller speaking, a minute | Half each |
|---|---|---|---|---|---|
| gpt-realtime-2.1 | $32 / $0.40 / $64 | $4 / $24 | $0.0768 | $0.0192 | $0.048 |
| gpt-realtime-2.1-mini | $10 / $0.30 / $20 | $0.60 / $2.40 | $0.024 | $0.006 | $0.015 |
| gpt-realtime-2 | $32 / $0.40 / $64 | $4 / $24 | $0.0768 | $0.0192 | $0.048 |
| gpt-realtime-1.5 | $32 / $0.40 / $64 | $4 / $16 | $0.0768 | $0.0192 | $0.048 |
| gpt-realtime-mini | $10 / $0.30 / $20 | $0.60 / $2.40 | $0.024 | $0.006 | $0.015 |
| gpt-realtime | $32 / $0.40 / $64 | $4 / $16 | $0.0768 | $0.0192 | $0.048 |
Per minute from OpenAI's statement that user audio is 1 token per 100 ms (600 a minute) and the assistant's 1 token per 50 ms (1,200 a minute), each token counted once. Instructions, text, tool calls and the conversation re-sent on later turns come on top.
Why a call costs more than the minutes of audio
- Caller audio counts 1 token per 100 ms (600 a minute) and the agent's audio 1 token per 50 ms (1,200 a minute).
- The whole conversation is sent to the model again for each response, so later turns cost more; cached input is billed at the lower cached price.
- Input transcription, if you turn it on, is billed separately at the transcription model's price.
- There is no charge for bandwidth or connections.
- OpenAI suggests measuring a sample session's tokens in the Realtime Playground, because costs are hard to estimate ahead of time.
- With voice activity detection on, empty audio doesn't count as input tokens, so silence costs nothing by itself.
So the figures above are a floor, not an estimate of a whole call.
Price a call
Set how much of each call the agent speaks; the model's audio tokens are priced at OpenAI's rates, counted once. GPT-Live is in the menu too, with its backend model as your own per-minute figure.
1,000 minutes a month on OpenAI Realtime API: at least $56.50, or $0.0565 a minute on average. This counts each second of audio once; real calls cost more, because every turn re-sends the conversation.
| Part | Per minute | Per month |
|---|---|---|
| Model gpt-realtime-2 | at least $0.048 | $48.00 |
| Phone line Twilio local number, inbound | $0.0085 | $8.50 |
| Total | $0.0565 | $56.50 |
GPT-Live: priced by the minute
GPT-Live separates the voice conversation from the backend that reasons and runs tools. The voice session costs $0.05 a minute, $50.00 for 1,000 minutes, and the backend is billed at its model's token prices.
- GPT-Live voice sessions cost $0.05 a minute, billed per second without rounding up to a whole minute.
- The session bills while it is open: when either side speaks, when both are silent, and while the backend works.
- Creating a WebRTC session bills 15 seconds up front, credited against the session once it runs.
What Retell AI charges for the same models
Retell AI resells OpenAI's realtime models per minute of call, the only per-minute price for them among the platforms we track:
| Model on Retell | Per minute | 1,000 minutes |
|---|---|---|
| GPT Realtime 1.5 | $0.345 | $345 |
| GPT Realtime | $0.345 | $345 |
| GPT Realtime mini | $0.07 | $70.00 |
| GPT Realtime 2.1 | $0.38 | $380 |
| GPT Realtime 2 | $0.38 | $380 |
| GPT Realtime 2.1 mini | $0.07 | $70.00 |
More on Retell's prices: Retell AI pricing.
The phone line
OpenAI charges nothing for bandwidth or connections; the phone line is your carrier's. Twilio's US list price for an inbound call on a local number is $0.0085 a minute.
Against voice agent platforms
| Platform | Cheapest picks | Dearest picks | 1,000 min a month | 10,000 min a month |
|---|---|---|---|---|
| Vapi | $0.0681 AssemblyAI, Google and Cartesia | $0.1313 OpenAI and ElevenLabs | $68.10 | $681 |
| Retell AI | $0.0716 GPT 5 nano and Retell Platform Voices | $0.795 GPT 6 Astra, fast tier and ElevenLabs v3 | $71.60 | $716 |
| Deepgram Voice Agent API | $0.075 the Standard tier | $0.163 the Advanced tier | $75.00 | $750 |
| Bland | $0.12 the Build plan | $0.14 the Start plan | $140 Start | $1,400 Start |
Per minute of conversation before the phone line, at each vendor's own prices: the platform fee with the cheapest, then the dearest, model and voice it lists, no add-ons. Monthly figures use the cheapest picks on the cheapest plan for that volume.
| Platform | Per minute | Billed on top |
|---|---|---|
| ElevenLabs Agents | $0.08 | Language model |
| OpenAI GPT-Live | $0.05 | Backend model and tools |
| OpenAI Realtime API | Tokens | Everything: audio and text tokens, with the conversation re-sent each turn |
| Twilio ConversationRelay | $0.07 | Your language model |
Every platform side by side, with a calculator: AI voice agent cost per minute. OpenAI's transcription models on their own: Whisper API pricing; its text-to-speech: OpenAI text-to-speech pricing.
Quick answers
How much does the OpenAI Realtime API cost per minute?
At least $0.048 a minute on gpt-realtime-2 when the agent speaks half the time, counting each second of audio once at OpenAI's token prices. Real calls cost more, because each turn sends the whole conversation back in as input (cached input is $0.40 per million audio tokens). Retell AI charges $0.38 a minute for GPT Realtime 2.
How much does gpt-realtime-mini cost per minute?
At least $0.015 a minute when the agent speaks half the time: $0.024 for a minute of the agent's audio and $0.006 for a minute of the caller's, each counted once. Retell AI charges $0.07 a minute for it.
What does GPT-Live cost?
$0.05 a minute of voice session, billed per second with no rounding up. The backend model and tools that do the reasoning are billed separately at their token prices, and the session bills while both sides are silent.
Does the Realtime API charge for silence?
Not by itself: with voice activity detection on, empty audio doesn't count as input tokens. GPT-Live is different: its session bills for as long as it is open.
Is there a per-minute price for the Realtime API?
Not for the speech-to-speech models, which OpenAI prices in tokens; GPT-Live is priced per minute. Platforms that resell the realtime models set their own: GPT Realtime 1.5 $0.345, GPT Realtime $0.345, GPT Realtime mini $0.07, GPT Realtime 2.1 $0.38, GPT Realtime 2 $0.38 and GPT Realtime 2.1 mini $0.07 a minute on Retell AI.
Sources
Read on 7 October 2026:
- developers.openai.com/api/docs/pricing.md
- developers.openai.com/api/docs/guides/realtime-costs.md
- twilio.com/en-us/voice/pricing/us
- developers.openai.com/api/docs/models/gpt-live-1.md
- retellai.com/pricing
Show the quoted text
Realtime and audio generation models Prices per 1M tokens unless noted. ### Grouped Pricing Table data | Model | Modality | Input | Cached input | Output / cost |
developers.openai.com/api/docs/pricing.md
| gpt-realtime-2.1 | Audio | $32.00 | $0.40 | $64.00 | | gpt-realtime-2.1 | Text | $4.00 | $0.40 | $24.00 | | gpt-realtime-2.1 | Image | $5.00 | $0.50 | - | | gpt-realtime-2.1-mini | Audio | $10.00 | $0.30 | $20.00 | | gpt-realtime-2.1-mini | Text | $0.60 | $0.06 | $2.40 | | gpt-realtime-2.1-mini | Image | $0.80 | $0.08 | - | | gpt-realtime-2 | Audio | $32.00 | $0.40 | $64.00 | | gpt-realtime-2 | Text | $4.00 | $0.40 | $24.00 | | gpt-realtime-2 | Image | $5.00 | $0.50 | - | | gpt-realtime-1.5 | Audio | $32.00 | $0.40 | $64.00 | | gpt-realtime-1.5 | Text | $4.00 | $0.40 | $16.00 | | gpt-realtime-1.5 | Image | $5.00 | $0.50 | - | | gpt-realtime-mini | Audio | $10.00 | $0.30 | $20.00 | | gpt-realtime-mini | Text | $0.60 | $0.06 | $2.40 | | gpt-realtime-mini | Image | $0.80 | $0.08 | - | | gpt-realtime | Audio | $32.00 | $0.40 | $64.00 | | gpt-realtime | Text | $4.00 | $0.40 | $16.00 | | gpt-realtime | Image | $5.00 | $0.50 | - |
developers.openai.com/api/docs/pricing.md. The realtime models' rows of the table: model, modality, then input, cached input and output prices per million tokens.
Audio tokens in user messages are 1 token per 100 ms of audio, while audio tokens in assistant messages are 1 token per 50ms of audio.
developers.openai.com/api/docs/guides/realtime-costs.md
The entire conversation is sent to the model for each Response. The output from a turn will be added as Items to the server Conversation and become the input to subsequent turns, thus turns later in the session will be more expensive.
developers.openai.com/api/docs/guides/realtime-costs.md
A Response can be created manually or automatically if voice activity detection (VAD) is turned on. VAD will effectively filter out empty input audio, so empty audio doesn't count as input tokens unless the client manually adds it as conversation input.
developers.openai.com/api/docs/guides/realtime-costs.md
Given the complexity in Realtime API token usage it can be difficult to estimate your costs ahead of time. A good approach is to use the Realtime Playground with your intended prompts and functions, and measure the token usage over a sample session.
developers.openai.com/api/docs/guides/realtime-costs.md
There is no cost currently for network bandwidth or connections.
developers.openai.com/api/docs/guides/realtime-costs.md
Aside from conversational Responses, the Realtime API bills for input transcriptions, if enabled.
developers.openai.com/api/docs/guides/realtime-costs.md
Number type To make calls To receive calls Refer from Twilio Local calls $0.0140 / min $0.0085 / min Toll-free calls $0.0140 / min $0.0220 / min Browser/app $0.0040 / min $0.0040 / min SIP interface $0.0040 / min $0.0040 / min
twilio.com/en-us/voice/pricing/us
[GPT-Live 1](https://developers.openai.com/api/docs/models/gpt-live-1) voice sessions are billed per second, without rounding up to a whole minute. Backend model and tool usage is charged separately. ### Pricing Table data | Model | Price per minute | | --- | --- | | gpt-live-1 | $0.05 |
developers.openai.com/api/docs/pricing.md
Voice sessions cost $0.05 per minute, billed per second. Backend model and tool usage is billed separately.
developers.openai.com/api/docs/models/gpt-live-1.md
Active session time includes time when the user speaks, the assistant speaks, both are silent, or the backend is working.
developers.openai.com/api/docs/guides/realtime-costs.md
A `POST /v1/live/sessions` request to create a WebRTC session bills 15 seconds of voice duration while the session initializes. That amount is credited against duration charges once the session starts running.
developers.openai.com/api/docs/guides/realtime-costs.md
Speech-to-speech LLM Notes Pricing GPT Realtime 1.5 $0.345/minute GPT Realtime $0.345/minute GPT Realtime mini $0.07/minute GPT Realtime 2.1 $0.38/minute GPT Realtime 2 $0.38/minute GPT Realtime 2.1 mini $0.07/minute
retellai.com/pricing
Prices on this page were checked against the official pricing pages on .