What AI transcription costs per hour
An hour of machine transcription costs between $0.10 and $0.61 at published self-serve API rates — a 6.1x spread, and the tightest of any lane on this site. Which means the rate card is rarely what decides your bill. Three things move it further than the vendor does: features billed separately by the hour, a streaming premium that runs anywhere from 0% to 278% at different vendors, and a per-request rounding rule that can bill a three-second clip as fifteen seconds.
One hour of audio, batch rates
Transcription is the one voice job where every serious vendor publishes a dollar rate in the same unit, so this table needs no conversion assumptions. These are asynchronous or pre-recorded rates — you upload a file and collect a transcript.
| Vendor and model | $ per hour of audio | Notes |
|---|---|---|
| Inworld STT, paid plans | $0.10 | $0.15 on demand |
| AssemblyAI Universal-2 | $0.15 | Previous-generation model |
| Gladia Growth, async | $0.20 | "as low as", volume commitment |
| AssemblyAI Universal-3.5 Pro | $0.21 | Base rate, add-ons extra |
| Deepgram Nova-3 monolingual, Growth | $0.216 | $0.0036/min, from $4,000/yr |
| ElevenLabs Scribe v2 | $0.22 | Diarization included |
| Deepgram Nova-3 monolingual, pay as you go | $0.258 | $0.0043/min |
OpenAI gpt-transcribe | $0.27 | $0.0045/min |
| Deepgram Whisper Large | $0.288 | No volume discount |
| Amazon Transcribe, batch | $0.36 | $0.006/min, per-second billing |
| Cartesia Ink-2, Scale credits | $0.40 | $0.42 Startup, $0.54 Pro |
| Gladia Starter, async | $0.61 | Diarization and language detection included |
| ElevenLabs subscription credits, Pro | $3.27 | Not an API rate — see below |
| Amazon Transcribe Medical, batch | $4.50 | Separate product, $0.075/min |
The Cartesia figures are CartSignal arithmetic rather than a published rate, and the derivation is worth stating because Cartesia sells transcription in credits: its pricing page gives the Free tier 20,000 credits for "~1h 51m" of Ink-2 and the Scale tier 8,000,000 for "~740h 44m", and both anchors independently imply about 10,800 credits per audio hour. Dividing each plan's price by its credits then gives $0.54 an hour on Pro at $5 a month, $0.42 on Startup at $49 and $0.40 on Scale at $299 — so Cartesia's ladder does discount with size, unlike several others this site has costed.
Two rows do not belong to the same market as the rest. The ElevenLabs subscription-credit route is what you pay if you transcribe inside a studio plan instead of through the API, and at $3.27 an hour on Pro it is about 15x the same company's own $0.22 API rate for the same model. Amazon Transcribe Medical is a separate product with its own compliance posture, and it is on the table only to show the size of the gap.
The headline number is the spread: 6.1x from top to bottom of the ordinary field. For comparison, this site has measured narration at roughly 30x, dubbing at 12.8x, avatar video at 12x and generated video at 12x. Transcription is the commodity lane. If you are choosing a vendor here, price is the least informative axis available to you.
Streaming is not reliably the expensive half
Every roundup states that real-time transcription costs more than batch. It usually does — but the premium is set so differently by each vendor that the rule is close to useless, and at one vendor it is simply false.
| Vendor | Batch | Streaming | Premium |
|---|---|---|---|
| AssemblyAI (Universal-2 → Universal-Streaming) | $0.15 | $0.15 | 0% |
| Deepgram Nova-3 monolingual, pay as you go | $0.258 | $0.288 | +11.6% |
| Gladia Starter | $0.61 | $0.75 | +23.0% |
| Amazon Transcribe | $0.36 | $0.60 | +66.7% |
| ElevenLabs Scribe v2 → v2 Realtime | $0.22 | $0.39 | +77.3% |
| AssemblyAI Universal-3.5 Pro → Pro Realtime | $0.21 | $0.45 | +114.3% |
OpenAI gpt-transcribe → gpt-live-transcribe | $0.27 | $1.02 | +278% |
AssemblyAI is the finding. Its pricing page lists Universal-Streaming English and Universal-Streaming Multilingual at $0.15 an hour each — the identical rate it charges for Universal-2 on pre-recorded files. Live transcription there costs exactly what batch costs, and it undercuts most of the batch column above. The same vendor then charges $0.45 for Universal-3.5 Pro Realtime, so AssemblyAI holds both the cheapest and nearly the dearest streaming row on the page depending on which model you pick.
The practical consequence is that "we need live captions" does not tell you your budget. It multiplies by 1.0 at one vendor and by 3.78 at another, and the difference is not explained by anything you can see in the transcript.
The base rate is a starting bid
AssemblyAI publishes the most detailed feature rate card in this field — more than a dozen add-ons, each quoted per hour, each stacking additively on the model rate. That transparency is genuinely useful, and it also makes visible something every vendor does to some degree: the advertised rate buys words on a page and very little else.
| Add-on | Per hour | Add-on | Per hour |
|---|---|---|---|
| Speaker diarization (async) | +$0.02 | Entity detection | +$0.08 |
| Speaker diarization (streaming) | +$0.12 | Auto chapters | +$0.08 |
| Speaker identification | +$0.02 | Topic detection (IAB) | +$0.15 |
| Sentiment analysis | +$0.02 | Content moderation | +$0.15 |
| Summarization | +$0.03 | Medical mode | +$0.15 |
| Keyterms prompting | +$0.05 | Profanity filtering | +$0.01 |
| Translation | +$0.06 | PII text redaction | +$0.08 |
Priced as real jobs on the $0.21 Universal-3.5 Pro base, the arithmetic is ours:
- Meeting notes — diarization, summarization, sentiment, entity detection: $0.36 an hour, +71.4%. That lands it exactly on Amazon Transcribe's plain batch rate.
- Podcast production — diarization, auto chapters, summarization, topic detection: $0.49, +133%.
- Compliance — diarization, both PII redactions, content moderation, profanity filtering: $0.52, +148%.
- Everything non-medical switched on: $1.06 an hour, about 5x the base rate.
Set that against the bundlers. Gladia charges $0.61 an hour on Starter with automatic language detection and speaker diarization already in the rate. ElevenLabs includes "Speaker diarization, up to 32 speakers" in Scribe v2's $0.22 and charges separately only for keyterm prompting and entity detection — its speech-to-text documentation says so explicitly. So à la carte wins until you want roughly four features at once, and then it stops winning: AssemblyAI has $0.40 of headroom under Gladia's bundled rate, and three of the configurations above spend most of it.
One detail worth noticing because it suggests a settled market price: keyterm prompting costs $0.05 an hour at both AssemblyAI and ElevenLabs, and entity detection is $0.08 at one and $0.07 at the other. Two vendors with entirely different base rates landed on nearly the same number for the same feature.
The word "medical" costs 0%, 71% and 1,150%
Three vendors sell a medical transcription option. Nothing else on this page varies as widely for a feature with the same name.
| Vendor | Standard | Medical | Premium |
|---|---|---|---|
| ElevenLabs Scribe v2 → Scribe v2 Medical | $0.22 | $0.22 | 0% |
| AssemblyAI Universal-3.5 Pro → medical mode | $0.21 | $0.36 | +71.4% |
| Amazon Transcribe → Transcribe Medical (batch) | $0.36 | $4.50 | +1,150% |
| Amazon Transcribe → Transcribe Medical (streaming) | $0.60 | $6.00 | +900% |
ElevenLabs lists Scribe v2 and Scribe v2 Medical on the same $0.22 row of its API pricing page. Amazon sells Transcribe Medical as a distinct service at 12.5x its own standard batch rate. CartSignal takes no position on whether these products are equivalent, and they almost certainly are not — a regulated clinical workflow involves accuracy validation, data handling and vocabulary coverage that no pricing page addresses. But if you have been quoted 12.5x by one vendor, it is worth knowing another charges nothing.
For short audio, rounding outweighs the rate
This is the multiplier nobody puts in a comparison table, because it lives in the billing terms rather than the price list. Google Cloud Speech-to-Text documentation states it plainly:
"Each request is rounded up to the nearest increment of 15 seconds. For example, if you make three separate requests, each containing 7 seconds of audio, you are billed $0.018 USD for 45 seconds (3 × 15 seconds) of audio. Fractions of seconds are included when rounding up to the nearest increment of 15 seconds. That is, 15.14 seconds are rounded up and billed as 30 seconds."
Amazon documents the opposite rule on its Transcribe pricing page: "Usage is billed in one-second increments, with no minimum applied." Neither AssemblyAI, Deepgram nor ElevenLabs publishes a billing increment at all on the pages checked, which is its own kind of answer — assume per-second and verify on your first invoice.
What 15-second rounding does is entirely independent of the rate, so the multiplier below applies whatever a vendor charges:
| Audio in one request | Billed as | You pay |
|---|---|---|
| 3 seconds | 15 seconds | 5.00x |
| 7 seconds | 15 seconds | 2.14x |
| 10 seconds | 15 seconds | 1.50x |
| 16 seconds | 30 seconds | 1.88x |
| 65 seconds | 75 seconds | 1.15x |
| 5 minutes or more | Same | 1.00x |
So the rule punishes exactly one shape of workload: a voice application sending many short utterances. A command-and-control interface averaging three-second turns pays five times for its audio under a 15-second increment, which swamps every price difference on this page — the entire competitive field only spans 6.1x. A podcast, a lecture or a support call is unaffected, because the rounding disappears into a long file. If you are building a voice agent and your candidate vendor publishes a rounding rule, price it on billed seconds and not on audio seconds. Amazon adds a second helpful line for call recordings: "For a two-channel conversation, you only pay for the total audio duration and won't be charged separately for each channel."
Free allowances are not comparable, and the biggest one is enormous
"Has a free tier" orders this field completely differently from the paid rates, and the units are all different — dollars of credit, hours per month, minutes for a fixed window.
| Vendor | Free allowance | Roughly, in hours of batch audio |
|---|---|---|
| Deepgram | $200 credit, "No expiration" | ~775 hours, once |
| AssemblyAI | "$50 in free credits, no credit card required" | 238–333 hours, once |
| Gladia | "50€ in free credits" | ~80 hours at Starter rates, once |
| Microsoft Azure (F0 tier) | "5 audio hours free per month" | 5 hours, every month |
| Amazon Transcribe | 60 minutes/month for 12 months | 12 hours total |
| Cartesia Free | 20,000 credits/month | ~1h 51m, every month |
Deepgram's is the outlier by a wide margin: $200 at its $0.0043-a-minute pre-recorded rate is roughly 775 hours of audio, and its pricing page attaches no expiry to it. For most individual projects that is not an evaluation allowance, it is the entire project. Azure's is the only recurring one large enough to matter — 5 hours every month, 60 hours a year, which will quietly cover a weekly podcast forever.
The conversions in the right-hand column are ours, and they assume you spend the credit on the cheapest batch model each vendor offers. Spend it on streaming or with add-ons switched on and the hours fall accordingly. Gladia quotes its credit in euros and its rates in dollars, so that row is an approximation at whatever the exchange rate is on the day.
Transcribing an hour costs about what synthesising one does
Worth noting for anyone budgeting a full voice pipeline. At the 1,000-characters-per-minute equivalence ElevenLabs publishes, an hour of speech is about 60,000 characters. That makes an hour of synthesised audio $0.24 on Amazon Polly Standard, $0.90 on Deepgram Aura-1 or OpenAI tts-1, $3.00 on the ElevenLabs API at its Flash rate and $12.00 on an ElevenLabs Starter subscription. Against transcription at $0.10 to $0.61, the two directions cost roughly the same at the cheap end of each market and diverge by about 50x at the expensive end — which is another way of saying the money in voice AI is in generating it, not reading it.
What this page does not publish
- Google Cloud per-model rates. The main Speech-to-Text pricing page did not render its tables on fetch today. The rounding rule quoted above comes from the Google-hosted pricing document CartSignal could load, which covers the on-premises edition; it lists "$0.006 / 15 seconds" for standard models with the first 60 minutes free, but because CartSignal could not confirm that figure against the cloud API rate card, no Google rate appears in the tables above.
- Microsoft Azure rates. The Azure Speech pricing page rendered its structure but not its numbers — every dollar amount appeared as a placeholder. What did render: the F0 free tier at 5 audio hours a month, and standard commitment tiers at 2,000, 10,000 and 50,000 hours a month. This is the same client-side-rendering problem this site has recorded at VEED, Murf and Kling, and it means the third hyperscaler cannot be costed from public data.
- Accuracy. No audio was transcribed, so nothing here ranks output quality. Vendors publish competing word-error-rate claims on their own benchmarks and against their own selections of competitors; those are marketing documents, and a rate card cannot tell you which model will handle your accents, your jargon or your background noise. The only safe reading of the 6.1x spread above is that switching vendors to save money is cheap, so test on your own audio with the free credits before committing.
- Token-billed transcription. Some models are billed per token rather than per minute of audio, which cannot be converted to an hourly rate without knowing the audio-to-token ratio for your material. Those rows are excluded rather than estimated.
What to buy
- A back catalogue of long files: the cheapest batch API you can get an account on. Start on Deepgram's $200 no-expiry credit, which may well finish the job.
- A voice app with short utterances: ignore the rate card and check the billing increment first. Amazon's per-second, no-minimum rule is the friendliest published; a 15-second increment can cost you more than the vendor choice does.
- Live captions: AssemblyAI's Universal-Streaming at $0.15 an hour is the cheapest streaming row here and matches its own batch price. Deepgram is next, with the caveat that its streaming rates are marked promotional with no published end date.
- Transcripts that need speakers, chapters or summaries: price the whole configuration, not the base rate. Compare the à la carte total against a bundled vendor before assuming the cheap headline wins.
- Already paying for a studio subscription: do not transcribe on credits. An hour is $3.27 on ElevenLabs Pro credits against $0.22 through the same company's API.
For what the same vendors charge to produce speech, see AI voice tools by job and the per-million-character ranking in ElevenLabs alternatives. For translating an existing recording rather than transcribing it, see what AI dubbing costs per minute, and for making a voice of your own, what voice cloning costs.
Frequently asked questions
How much does it cost to transcribe one hour of audio?
Between $0.10 and $0.61 at published self-serve API rates. Inworld STT is $0.10 on paid plans, AssemblyAI Universal-2 $0.15, AssemblyAI Universal-3.5 Pro $0.21, ElevenLabs Scribe v2 $0.22, Deepgram Nova-3 $0.258, OpenAI gpt-transcribe $0.27, Amazon Transcribe $0.36, Cartesia Ink-2 $0.40 to $0.54 and Gladia Starter $0.61 with diarization included. Buying the same hour with ElevenLabs subscription credits instead is $3.27, about 15x that company's own API rate.
Is real-time transcription more expensive than batch?
Usually, but the premium ranges from nothing to 278%. AssemblyAI charges $0.15 for both Universal-Streaming and its Universal-2 batch model. Deepgram adds 11.6%, Gladia 23%, Amazon 66.7%, ElevenLabs 77.3%, AssemblyAI's top model 114.3%, and OpenAI 278% moving from gpt-transcribe to gpt-live-transcribe.
Why is my transcription bill higher than the advertised rate?
Features and rounding. AssemblyAI prices more than a dozen add-ons separately, so a meeting transcript with diarization, summarization, sentiment and entities is $0.36 an hour rather than $0.21. And Google documents that each request is rounded up to the nearest 15 seconds, so a three-second clip bills as fifteen. Amazon bills in one-second increments with no minimum.
Does medical transcription cost more?
By 0%, 71% or 1,150%, depending on the vendor. ElevenLabs lists Scribe v2 Medical at the same $0.22 as Scribe v2. AssemblyAI charges +$0.15 an hour for medical mode. Amazon sells Transcribe Medical at $4.50 an hour against $0.36 for standard batch. Pricing says nothing about clinical suitability, which these pages do not address.
What is the cheapest way to transcribe a large back catalogue?
A batch endpoint on a developer key, with long files rather than many short ones, starting on free credit. Deepgram's $200 credit carries no expiry and covers roughly 775 hours at its pre-recorded rate. AssemblyAI gives $50, worth 238 to 333 hours. Azure gives 5 hours a month, recurring.
Sources
- AssemblyAI pricing: model rates, the full add-on rate card and the $50 signup credit — read 20 September 2026
- Deepgram pricing: Nova-3, Flux and Whisper rates, Growth tier and the $200 credit — read 20 September 2026
- ElevenLabs API pricing: Scribe v2, Scribe v2 Medical and Scribe v2 Realtime per hour — read 20 September 2026
- ElevenLabs speech-to-text docs: included diarization, paid keyterm prompting and entity detection, 3 GB and 10-hour limits
- Amazon Transcribe pricing: batch and streaming rates, Transcribe Medical, per-second billing and the two-channel rule — read 20 September 2026
- Google Cloud Speech-to-Text pricing: the 15-second rounding rule and worked example
- Microsoft Azure Speech pricing: free tier and commitment tiers, with dollar amounts unrendered on fetch
- Gladia pricing: bundled async and real-time rates with diarization included
- Inworld pricing: STT at $0.10 on paid plans and $0.15 on demand
- Cartesia pricing: Ink-2 hours included by tier, used to derive the per-hour rate
- OpenAI API pricing: gpt-transcribe and gpt-live-transcribe per minute