MAI-Transcribe-2 Streaming: Pricing, vs Whisper, and How to Use It

Logeshwaran
—
MAI-Transcribe-2 Streaming: Pricing, vs Whisper, and How to Use It

MAI-Transcribe-2-Streaming is Microsoft's new real-time speech-to-text model, launched on October 1, 2026. It writes live transcripts in 60 languages with a 2.5% word error rate, shows the first words in just over 100 milliseconds, and ranks first for accuracy on Artificial Analysis's streaming leaderboard. It costs $0.54 per hour of audio through the end of 2026, through Microsoft Foundry and the MAI Playground. Here is what the launch posts do not line up for you: Microsoft's batch model from four weeks earlier, plain MAI-Transcribe-2, costs $0.10 an hour and scores even better on recorded audio. So for meetings, lectures and podcasts you have already recorded, the "No. 1" model is the expensive choice, at 5.4 times the price. You only want the streaming one when the words must appear while someone is still talking. Neither runs on your own PC. For that, the free answer is still Whisper, and this page shows both routes.

Ethan is in his second year of an evening accounting course, and every lecture is a two-hour Teams recording he means to rewatch and never does. On Thursday he saw "Microsoft launches No. 1 transcription AI" on his feed and came into Jake's shop with a question he thought was simple: "Can I get that thing to turn my lectures into notes? And can it run on my laptop so I'm not paying for it?" Jake has been setting up dictation and captions for customers since Windows 10, and knows that "Microsoft transcription" means at least five different things depending on which button you press. That is where this page starts: what the MAI-Transcribe models are, which one to pay for and which to skip, what it all costs in real numbers, how it compares with Whisper, and the Microsoft transcription tools already sitting on your PC that cost nothing extra.

⚡ Quick Answer

• What is MAI-Transcribe-2? → Microsoft AI's own speech-to-text models: a batch version (September 3, 2026) and a real-time Streaming version (October 1, 2026). The whole family.

• Price → Streaming $0.54 per audio hour; batch $0.10 per audio hour; both introductory rates through the end of 2026. Cost examples.

• Which one do I need? → Recorded files: batch. Live captions, voice agents, dictation: Streaming. Decide in 30 seconds.

• Open source or local? → No. Cloud API only. The free local alternative is Whisper with whisper.cpp. Run Whisper on Windows.

Already have Microsoft 365? Word's Transcribe button gives you 300 minutes a month at no extra cost, and Windows 11 Live Captions is free and on-device. The tools you already own.

A quick orientation if you are new to how these models fit together. Speech-to-text models are a close cousin of the language models behind chatbots: they turn sound into words rather than words into more words. Some, like Microsoft's, live only in the cloud and charge by the hour of audio. Others, like OpenAI's Whisper, are free downloads you can run on your own machine, the same way the free guide series on running AI locally runs chat models. Which kind suits you depends on three things: how much audio you have, whether it is live or recorded, and whether it is allowed to leave your computer.


What is MAI-Transcribe-2-Streaming?

MAI stands for Microsoft AI, the company's in-house model group, which has spent 2026 building its own models instead of relying only on partners. MAI-Transcribe is its speech-recognition line. The new Streaming model is the first in that line built for real time: it listens to audio as it arrives and keeps a running transcript, instead of waiting for a finished file.

Real-time transcription has a problem that batch transcription does not. When someone says "I'd like to book a…", the model has to show something before the sentence ends, and then fix it as more context arrives. Those early guesses are called partials. MAI-Transcribe-2-Streaming produces its first partials in just over 100 milliseconds, revises them as the speaker continues, and commits a stable final transcript about 0.13 seconds after the words are spoken. Microsoft reports a 2.5% word error rate on final transcripts and 2.8% on first partials, and says words appear on screen about twice as fast as with its closest competitor.

Two practical features come with it. It handles 60 languages and detects the language automatically and continuously, so a speaker who switches from English to Spanish mid-sentence does not need anyone to change a setting. And because partials arrive so quickly, a voice agent built on it can start acting before the speaker finishes, which is what makes a phone assistant feel responsive instead of awkward.

The launch also included two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, covered briefly further down. Together they make the two ends of a voice conversation: one model listens, the other speaks.

MAI-Transcribe 1 vs 1.5 vs 2 vs 2-Streaming: the whole family

Searches for "mai transcribe 1", "mai transcribe 1.5" and "mai transcribe 2" are all common, because Microsoft shipped four versions in six months. Here is the line-up in order.

Model Released Mode Languages Price per audio hour Headline
MAI-Transcribe-1April 2, 2026Batch25$0.36Microsoft's first in-house speech model; rolling into Copilot Voice and Teams
MAI-Transcribe-1.5In June 2026BatchMore than v1SupersededIncremental update between 1 and 2
MAI-Transcribe-2September 3, 2026 (preview)Batch60$0.10 (intro, through 2026)2.0% AA-WER, among the top three on Artificial Analysis; 72% cheaper than v1
MAI-Transcribe-2-StreamingOctober 1, 2026Real time60$0.54 (intro, through 2026)No. 1 for streaming accuracy on Artificial Analysis; 2.5% final WER

Three things stand out. First, the batch model got dramatically cheaper between versions 1 and 2, from $0.36 to $0.10 an hour. Second, both version 2 prices are introductory rates that run "through the end of the year". Microsoft has not said what they become in 2027, so treat them as launch pricing, not a promise. Third, version 1 is the one Microsoft said it was rolling into Copilot's Voice mode and Microsoft Teams. If you have used Teams transcription or talked to Copilot by voice this year, there is a fair chance an MAI model was already doing the listening.

MAI-Transcribe pricing: what it really costs

Per-hour prices are hard to feel, so here are the real jobs people ask about, at the introductory rates.

Job Audio Batch ($0.10/hr) Streaming ($0.54/hr)
One semester of lectures15 weeks x 2 lectures x 2 hours = 60 hours$6.00$32.40
A weekly one-hour podcast, for a year52 hours$5.20$28.08
A small team's meetings10 hours a week, for a month (about 43 hours)$4.30$23.22
A support phone line20 agents x 6 hours x 22 days = 2,640 hours a month$264 (after the call)$1,425.60 (live)

For a person, both prices are small. Ethan's whole semester costs less than a coffee in batch mode. The difference starts to matter at business volume, and that is where choosing the right model pays. A support line that only needs searchable call records afterward saves over $1,100 a month by sending recordings to the batch model instead of streaming them. A support line whose voice agent must answer callers live has no choice: it needs streaming.

Two pricing details are easy to miss. These are prices for audio minutes processed, so a two-hour recording costs two hours whether it holds two hours of speech or forty minutes of speech and a long silence. Trimming silence before upload is free money at volume. And Foundry is an Azure service, so the usual cloud habits apply: set a budget alert, and remember that storage and other services you use alongside it bill separately. If you are new to cloud billing generally, our comparison of the three big clouds explains how Azure's billing fits next to the others.

Streaming or batch: which MAI-Transcribe model do you need?

This is the decision that matters most, and it takes thirty seconds.

  1. Is the audio already recorded? Lectures, meetings saved to a file, podcast episodes, voice memos, interviews, court or compliance recordings. Use batch MAI-Transcribe-2. It is cheaper, and it is more accurate on the leaderboard, because it can hear the whole sentence before deciding what was said.
  2. Must the words appear while the person is still speaking? Live captions for an event, real-time subtitles, dictation into a document, a phone agent that answers callers. Use MAI-Transcribe-2-Streaming. That is the job its speed exists for.
  3. Do you need speaker labels, timestamps or custom vocabulary? The batch model offers speaker separation, word-level timestamps, keyword biasing for names and jargon, and a choice of "verbatim" or "clean" output. For a meeting transcript that should say who said what, batch is the better fit anyway.
  4. Must the audio never leave your computer? Neither. Both are cloud services. Skip to the local Whisper section.
  5. Do you just want your own Teams meetings or Word recordings transcribed? You may not need to touch either model directly. See the Microsoft tools you already have.

Ethan's lectures are recordings. Batch was the obvious answer, and more importantly, it turned out he did not need an API account at all, because his college's Microsoft 365 license already included the Word tool covered below.

How accurate is MAI-Transcribe-2? What 2.5% WER means

Word error rate, or WER, counts the words a model gets wrong, adds or drops, as a share of the words actually spoken. A 2.5% WER means roughly one mistake in every forty words. In a one-hour lecture of about 9,000 words, that is around 225 errors. Most are small, a "the" for an "a" or a dropped "um", but a few will be names, numbers or technical terms, which are exactly the words you care about.

Leaderboard figures come from test sets of clean-ish speech. Real recordings are messier: accents, crosstalk, a laptop microphone at the back of a lecture hall, background noise. Expect your own error rate to be higher than any headline number. That is true for every model, not just Microsoft's. The leaderboard is still useful for comparing models with each other, because they all face the same test.

On that comparison, Microsoft's models sit at the top. Artificial Analysis, an independent benchmarking group, ranks MAI-Transcribe-2-Streaming first for streaming accuracy on both final and partial transcripts. In its non-streaming ranking, batch MAI-Transcribe-2 scores 2.0% on the group's AA-WER test, among the top three, alongside models from ElevenLabs, StepFun and others that trade places as new versions arrive. Microsoft's earlier MAI-Transcribe-1 was already measured against Whisper large-v3, ElevenLabs Scribe v2 and Google's Gemini Flash-Lite on the multilingual FLEURS test, and came out ahead of all of them on Microsoft's numbers.

Three habits improve accuracy more than switching models. Record closer to the speaker: a cheap clip-on microphone beats an expensive model listening from across the room. Use keyword biasing where it is offered, giving the model your course names, product names and colleagues' names. And keep the "verbatim" setting off unless you need every "um", because the "clean" output reads far better as notes. If your microphone itself is acting up after a Windows update, the USB audio Code 10 fix is worth checking before you blame any model.

How to use MAI-Transcribe-2: Playground, Foundry and API

There are three ways in, from easiest to most involved.

  1. MAI Playground. Microsoft AI's own web playground lets you try the models in a browser, including a live voice-agent demo called "Chatter" that pairs the Streaming model with the new voices. This is the fastest way to hear how it handles your accent or your topic before spending anything.
  2. Microsoft Foundry. For real use, the models live in Microsoft Foundry, Azure's AI platform. You need an Azure account with billing set up. You create a Foundry project, find MAI-Transcribe-2 or MAI-Transcribe-2-Streaming in the model catalog, deploy it, and use the endpoint and key it gives you from your own code or from tools that support it. The batch model was in public preview at launch, so check its current status on the catalog page.
  3. Partner platforms. Microsoft also lists OpenRouter, Vercel and Azure Voice Live as places where its new audio models are available, with LiveKit, a popular toolkit for real-time voice apps, coming soon. If you already build on one of those, that may be the shortest path.

Whichever route you take, start with a short test file that represents your real audio: the same microphone, the same room, the same speakers. A three-minute sample tells you more than any benchmark. And set a spending alert in Azure before you send a backlog of files, because "it's only ten cents an hour" stops being true when a script loops on the same folder overnight.

MAI-Transcribe vs Whisper: which should you use?

"mai transcribe vs whisper" is the comparison most people are really asking about, because Whisper, OpenAI's speech model, is free and open, and has been the default choice since 2022. Here is how they compare in practice.

MAI-Transcribe-2 (batch) MAI-Transcribe-2-Streaming Whisper large-v3-turbo
Open sourceNoNoYes, MIT license
Runs on your PCNo, cloud onlyNo, cloud onlyYes, CPU or GPU
Cost$0.10 per audio hour (intro)$0.54 per audio hour (intro)Free; your electricity
Real timeNoYes, about 100 ms partialsNot natively; community tools approximate it
Languages60, auto-detected60, continuous detectionAbout 99, quality varies widely
AccuracyTop tier on independent testsNo. 1 for streamingGood, below the newest cloud models
PrivacyAudio goes to Microsoft's cloudAudio goes to Microsoft's cloudNever leaves your machine

The plain verdict: MAI-Transcribe-2 is more accurate, especially on noisy and multilingual audio, and cheap enough that cost rarely decides it for a person. Whisper wins on two things money cannot buy from a cloud API: privacy and independence. If your recordings are confidential, such as therapy sessions, legal interviews, medical dictation or HR meetings, or if your employer forbids uploading them, Whisper on your own machine is the answer whatever the leaderboard says. Plenty of people end up using both: the cloud model for public or low-sensitivity audio, and Whisper for anything private.

Is MAI-Transcribe open source? Can you run it locally?

No to both. MAI-Transcribe-1, 1.5, 2 and 2-Streaming are closed models, available only as cloud services. Microsoft has published no weights, no download and no on-device version, so any "MAI-Transcribe GitHub" project is someone's code for calling the API, not the model itself. The searches for "is mai transcribe open source" all end at the same place.

If you want speech-to-text that runs on your own PC, the practical choice is Whisper large-v3-turbo through whisper.cpp, a free program from the same ggml team behind llama.cpp. The turbo model has about 809 million parameters, its file is around 1.6 GB, and it runs on an ordinary laptop CPU, faster still with a graphics card. Here is the Windows route.

  1. Download whisper.cpp from its GitHub releases page. Pick the plain Windows x64 build for CPU, or a CUDA build if you have an NVIDIA card. Unzip it to a folder such as C:\whisper.
  2. Download the model. From the same project, get ggml-large-v3-turbo.bin from its Hugging Face model list, or run the bundled download-ggml-model.cmd large-v3-turbo script. Put the file in a models folder next to the program.
  3. Install ffmpeg with winget install ffmpeg. whisper.cpp prefers 16 kHz WAV audio, and ffmpeg converts anything else.
  4. Convert your recording: ffmpeg -i lecture.mp4 -ar 16000 -ac 1 -c:a pcm_s16le lecture.wav
  5. Transcribe it: whisper-cli -m models\ggml-large-v3-turbo.bin -f lecture.wav -otxt -osrt. The -otxt flag writes a plain text transcript and -osrt writes subtitles with timestamps. Older releases call the program main.exe instead of whisper-cli.

On Linux and Kali the steps are the same with the Linux build or a source build, and on an Apple Silicon Mac whisper.cpp uses the built-in GPU and is quick. A two-hour lecture takes a few minutes on a recent laptop with a graphics card and longer on CPU alone. Nothing leaves the machine at any point. If you already run chat models locally, the guide to installing AI models on a PC covers the same groundwork, and a transcript from Whisper is the perfect input for a local chat model to summarize into notes.

Microsoft transcription tools you already have (free or included)

Here is the part most launch coverage skips, and the reason "microsoft transcription tool" is searched far more than any model name. Most people do not need an API at all. Microsoft already puts transcription in places you may be paying for, or that come free with Windows.

Tool What it does Cost and limits Best for
Word TranscribeUpload or record audio; get a transcript with speakers and timestampsIncluded with Microsoft 365: 300 minutes of uploads a month (30,000 with a Copilot license); MP3, WAV, M4A, MP4, MPEG up to 300 MBLectures, interviews, voice memos
Teams meeting transcriptionLive transcript during a meeting, saved with the recordingIncluded in business and education Teams plans; the organizer's policy must allow itWork and class meetings
Windows voice typing (Win+H)Dictate into any text boxFree with WindowsWriting emails and notes by voice
Live Captions (Win+Ctrl+L)Captions for any audio playing on the PCFree with Windows 11, processed on-deviceFollowing videos and calls; accessibility
MAI-Transcribe APIProgrammatic transcription at scale$0.10 or $0.54 per audio hourDevelopers and businesses

How to use Word's Transcribe

  1. Open a new document in Word for the web, or in a current Microsoft 365 version of Word, signed in with your Microsoft 365 account.
  2. On the Home tab, open the arrow next to Dictate and choose Transcribe.
  3. Choose Upload audio and pick your recording, or Start recording to capture a live conversation.
  4. Wait while it processes. A long file can take a while, and you can keep working in another tab.
  5. Review the transcript in the side pane, rename speakers, then choose Add to document to drop it in with or without speakers and timestamps.

Three hundred minutes a month covers about two and a half two-hour lectures a week, which was just enough for Ethan's course. The limit cannot be raised on a normal Microsoft 365 plan, so a heavy user either spreads the work across months, adds Whisper for the overflow, or moves to the API.

How to turn on transcription in Teams

In a Teams meeting, open More actions (the three dots), choose Record and transcribe, then Start transcription. Everyone in the meeting sees a notice that transcription has started. Afterward the transcript appears in the meeting chat and the Recap tab. If the option is grayed out, your organization's administrator has switched it off, and only they can change that. Teams Premium and Microsoft 365 Copilot licenses add AI meeting notes on top of the transcript, which is often what people actually want from "transcription".

Microsoft transcription not working? Fixes for Word, Teams and voice typing

"Microsoft word transcribe not working" and "teams transcription not working" are among the most common searches in this whole topic, and almost every case comes down to one of a handful of causes. Work down this list in order.

  1. The Transcribe button is missing in Word. Transcribe needs a Microsoft 365 subscription and a signed-in account. It appears in Word for the web and current Microsoft 365 versions of Word, not in one-time-purchase versions such as Office 2021 or 2024. Sign in with the account that holds the subscription, or open the document in Word for the web.
  2. Word says you have reached your limit. The 300 uploaded minutes reset monthly and cannot be topped up on a normal plan. Live recording inside Word is not counted the same way as uploads, so recording a meeting directly can be a workaround for future sessions.
  3. The upload fails or never finishes. Check the file is MP3, WAV, M4A, MP4 or MPEG and under 300 MB. A long video can be shrunk to audio first with ffmpeg -i lecture.mp4 -vn -c:a aac -b:a 64k lecture.m4a, which usually lands well under the limit and uploads faster.
  4. Teams shows Start transcription grayed out. Your organization's meeting policy has transcription off, or you joined as a guest or from an account without permission. Only an administrator can change the policy. In personal and free Teams, the feature set is smaller.
  5. The Teams transcript is in the wrong language or full of nonsense. The meeting's spoken-language setting does not match what people are speaking. Change it in the transcript settings during the meeting. It applies from that point forward.
  6. Win+H voice typing does nothing. Check that a text box has focus, that Windows has microphone permission under Settings, Privacy and security, Microphone, and that the right input device is selected under Settings, System, Sound.
  7. Everything is blank or silent. The microphone itself is the problem. Test it in Settings, System, Sound, Input. If a recent Windows update broke it, the USB audio Code 10 fix covers the most common cause this autumn.

For the API, failures are usually more mundane: an expired key, a deployment name typed differently from the one in your code, an audio format the endpoint does not accept, or a spending limit reached. The Foundry portal shows the exact error for each request, and it is worth reading in full before changing anything.

Better recordings, better transcripts: what actually helps

Every model on this page, Microsoft's or Whisper, is limited by the audio you feed it. The difference between a 2% and a 10% error rate is far more often the recording than the model, and fixing the recording costs little or nothing.

Distance beats everything. A microphone at arm's length picks up the speaker. A microphone across a room picks up the room: echo from bare walls, the air conditioner, chairs scraping. In a lecture hall, a phone placed on the front desk records better than a laptop at the back. For interviews, a cheap clip-on microphone for each person beats any single expensive microphone between them.

One voice at a time, where you can. Overlapping speech is the hardest thing for any model. Meeting transcripts improve noticeably when people do not talk over each other, and when each remote participant uses a headset instead of laptop speakers that feed the other side's voice back into the microphone.

Be careful with heavy noise suppression. Some apps and headsets apply aggressive noise removal that sounds clean to a human but strips out parts of consonants the model relies on. If a transcript is oddly bad on audio that sounds fine, try recording once with suppression set lower and compare.

Do not compress twice. Each round of lossy compression costs a little detail. Record in a decent format the first time, such as WAV, or M4A at 64 kbps or more for speech, and convert once if you must. Converting an already-compressed MP3 into another compressed format does not add anything back.

Mono and 16 kHz are enough for speech. Speech models work at 16 kHz internally, so recording at higher rates or in stereo makes files bigger without improving the transcript. For uploads with a size limit, such as Word's 300 MB, a mono 16 kHz file holds many hours of speech.

Say the hard words clearly once. In a meeting, spelling out an unusual name or product the first time it comes up, or adding it to the keyword list where the tool offers one, fixes it for the rest of the transcript far more reliably than correcting it fifty times afterward.

Transcription as accessibility

For people who are deaf or hard of hearing, live transcription is not a convenience. It is how they join the conversation. Windows 11 Live Captions works on any audio playing on the PC, including video calls in apps without their own captions, and it runs on-device, so nothing is uploaded. Teams live captions can be switched on per person without anyone else needing to change anything. If you organize meetings, turning on captions or transcription by default, and telling people it is on, is one of the kindest small habits in any workplace. The Streaming model's speed is exactly what makes live captions feel natural rather than a sentence behind, which is why this launch matters beyond developers.

From transcript to useful notes: the step most people skip

A raw transcript is not notes. Ethan's first two-hour lecture came back as 14,000 words, which is more than most people will ever reread. The value comes from what you do next, and this is where transcription and the chat models covered elsewhere on this site meet.

  1. Clean it up. Use the "clean" style where offered, or ask a chat model to remove filler words and false starts without changing the meaning. This alone cuts a transcript by a fifth and makes it readable.
  2. Split it by topic. Ask for headings wherever the subject changes. A lecture usually has five to eight sections, and seeing them is half the revision.
  3. Pull out the essentials. Ask for the definitions, formulas, deadlines and any sentence that starts with "this will be on the exam". For meetings, ask for decisions made, actions assigned and who owns each one.
  4. Keep the timestamps. When a summary says something surprising, the timestamp lets you jump to that minute of the recording and hear it in context. Summaries can be wrong; the recording is the source.
  5. Do it locally when it is private. A local chat model running in Ollama can summarize a Whisper transcript without either ever leaving your machine. For a confidential meeting, that pairing gives you the whole pipeline, from audio to action list, offline.

One caution: a summary is only as good as the transcript under it. If a name or a number matters, such as a due date, a dosage or a sum of money, check it against the recording rather than trusting either the transcript or the summary. A 2.5% error rate is excellent, and it still means the one word you needed might be the wrong one.

Does Teams or Copilot use MAI-Transcribe?

Partly, and gradually. When Microsoft launched MAI-Transcribe-1 in April 2026, it said the model was rolling out in phases to Copilot's Voice mode and Microsoft Teams. It has not published which Teams features, regions or tenants use which model at any given time, and the October launch post for the Streaming model does not mention Teams or Copilot at all.

For users, this matters less than it sounds. You do not choose the model inside Teams or Word; Microsoft does, and swaps it as better ones arrive. What you can control is the audio going in, the language setting on the meeting, and whether transcription is on. For businesses, it is a useful reminder that the transcription quality inside Microsoft 365 tends to improve quietly over time. A transcript that was poor a year ago is worth testing again.

If what you actually want is less Copilot rather than more, that is a separate setting. Our guide to stopping Microsoft 365 Copilot from starting with Windows covers it, and transcription in Word and Teams keeps working either way.

MAI-Voice-2.1 and MAI-Voice-2.1-Flash: the other half of the launch

The October 1 launch paired the Streaming transcriber with two text-to-speech models.

MAI-Voice-2.1 speaks 23 languages across 26 locales and can keep one voice consistent across all of them, carrying a native accent in each, rather than sounding like a different person in every language. It supports voice cloning from a few seconds of reference audio. In Microsoft's listening test with 4,000 people, 50.3% rated it as human-like as, or more human-like than, real recordings. It costs $22 per million characters.

MAI-Voice-2.1-Flash is the fast, cheap version for high-volume work. Microsoft says it can generate 45 seconds of audio with 150 milliseconds of end-to-end delay, at $15 per million characters. Paired with the Streaming transcriber, it is the building block for phone agents that listen and answer without the awkward pause.

A word of care about voice cloning, since a few seconds of audio is all it takes. Clone only voices you have clear permission to use, and treat any unexpected voice call asking for money or passwords with suspicion, even if it sounds exactly like someone you know. A family code word for emergencies costs nothing and defeats the most common voice-clone scam.

Privacy, consent and where your audio goes

Any cloud transcription, Microsoft's or anyone else's, means your audio travels to a data center and is processed there. For a business on Azure, Microsoft's enterprise terms govern how that data is handled, and organizations with strict requirements should confirm the data region and retention settings for their Foundry deployment before sending anything sensitive. For individuals, the simple rule is: if you would not email the recording to a stranger, think twice before uploading it anywhere, and use Whisper locally instead.

Consent matters too. Recording and transcribing a conversation can require everyone's permission, depending on where you live, and the rules differ between countries and even between US states. Teams announces transcription to every participant for exactly this reason. If you record a lecture, check your college's policy; many allow recordings for personal study but not for sharing. If you record calls for a business, tell people at the start of the call.

Finally, transcripts are data, and they can outlive the meeting. A transcript sitting in a shared folder is searchable by anyone with access. Store them where the recordings would be stored, delete them on the same schedule, and keep anything about health, money or HR out of shared channels.

Who should use what: the short version

  1. Students and individuals with recordings: Word Transcribe if you have Microsoft 365, Whisper with whisper.cpp if you do not, or if the audio is private.
  2. Office workers in meetings: Teams transcription, already included with business plans. Ask IT if the option is grayed out.
  3. People who want to write by voice: Windows voice typing with Win+H. Free, and good enough for emails.
  4. Anyone following audio they cannot hear well: Live Captions with Win+Ctrl+L. On-device and private.
  5. Developers transcribing recorded files at scale: batch MAI-Transcribe-2 at $0.10 an hour, while the introductory rate lasts.
  6. Developers building live captions or voice agents: MAI-Transcribe-2-Streaming, the reason it exists.
  7. Anyone with confidential audio: Whisper, locally, whatever the leaderboard says.

If you are deciding whether paying for AI tools makes sense at all, our honest look at AI subscriptions runs the same kind of math for chat assistants.

MAI-Transcribe and Microsoft transcription: frequently asked questions

What is MAI-Transcribe-2-Streaming?

MAI-Transcribe-2-Streaming is Microsoft AI's real-time speech-to-text model, launched on October 1, 2026. It transcribes 60 languages live with automatic language detection, shows first words in about 100 milliseconds, has a 2.5% final word error rate, and costs $0.54 per audio hour through the end of 2026.

How much does MAI-Transcribe-2 cost?

Batch MAI-Transcribe-2 costs $0.10 per hour of audio and MAI-Transcribe-2-Streaming costs $0.54 per hour, both introductory prices through the end of 2026. The earlier MAI-Transcribe-1 was $0.36 per hour.

Is MAI-Transcribe open source?

No. All MAI-Transcribe models are closed and available only as cloud services through Microsoft Foundry, the MAI Playground and partner platforms. There are no downloadable weights. Whisper is the main open-source alternative.

Can I run MAI-Transcribe locally?

No. It runs only in Microsoft's cloud. To transcribe on your own PC, use Whisper large-v3-turbo with whisper.cpp, which is free, runs on a CPU or GPU, and never sends your audio anywhere.

What is the difference between MAI-Transcribe-1, 1.5 and 2?

MAI-Transcribe-1 (April 2026) covered 25 languages at $0.36 per hour. In June 2026, version 1.5 followed as an incremental update. MAI-Transcribe-2 (September 2026) covers 60 languages at $0.10 per hour with top-tier accuracy, and the Streaming version (October 2026) adds real-time transcription.

MAI-Transcribe vs Whisper: which is better?

MAI-Transcribe-2 is more accurate on independent tests and cheap at $0.10 per hour, but cloud-only. Whisper is free, open source under the MIT license, and runs on your own machine, so it wins for private or confidential audio.

What languages does MAI-Transcribe support?

MAI-Transcribe-2 and MAI-Transcribe-2-Streaming support 60 languages with automatic language identification, and the Streaming model detects language changes continuously mid-conversation. MAI-Transcribe-1 supported 25 languages.

How do I use MAI-Transcribe-2?

Try it in the MAI Playground in a browser, or deploy it from the model catalog in Microsoft Foundry with an Azure account and call the endpoint from your code. Microsoft also lists OpenRouter, Vercel and Azure Voice Live as platforms for its new audio models.

Does Microsoft have a free transcription tool?

Yes. Word's Transcribe feature is included with Microsoft 365 for 300 minutes of uploaded audio a month, Windows voice typing (Win+H) and Live Captions (Win+Ctrl+L) are free with Windows, and Teams meeting transcription is included in business and education plans.

What is the Word Transcribe limit?

Microsoft 365 subscribers can transcribe 300 minutes of uploaded audio per month in Word, and users with a Microsoft Copilot license get 30,000 minutes. Each file can be up to 300 MB in MP3, WAV, M4A, MP4 or MPEG format. The 300-minute limit cannot be raised on standard plans.

Does Microsoft Teams have transcription?

Yes. In a meeting, choose More actions, then Record and transcribe, then Start transcription. The transcript is saved to the meeting chat and Recap. If the option is grayed out, your organization's administrator has disabled it.

Why is Microsoft Teams transcription not working?

The most common causes are an organization policy that disables transcription, a guest account without permission, or an unsupported spoken-language setting. Check the meeting's spoken language under the transcript settings and ask your administrator to confirm the transcription policy.

Does Teams use MAI-Transcribe?

Microsoft said in April 2026 that MAI-Transcribe-1 was rolling out in phases to Teams and Copilot's Voice mode. It has not published which Teams features use which model, and users cannot choose the model inside Teams.

What is a good word error rate for transcription?

Under 5% is very good, and the best cloud models now score around 2 to 2.5% on standard tests. Real recordings with noise, accents or distant microphones score worse with every model, so test on your own audio.

What is MAI-Voice-2.1?

MAI-Voice-2.1 is Microsoft's text-to-speech model launched with the Streaming transcriber on October 1, 2026. It speaks 23 languages across 26 locales, supports voice cloning from seconds of audio, and costs $22 per million characters; the faster Flash version costs $15.

Is MAI-Transcribe better than ElevenLabs Scribe?

On Artificial Analysis, MAI-Transcribe-2-Streaming ranks first for streaming accuracy, and batch MAI-Transcribe-2 sits among the top three alongside ElevenLabs Scribe v2, at a lower introductory price. Rankings shift as new versions ship, so compare on your own audio before committing.

Ethan left the shop without signing up for anything. His college account already had Microsoft 365, so Jake showed him Word's Transcribe button, and his first lecture came back as a 14,000-word transcript with timestamps before he had finished his tea. Two lectures a week fit inside the 300 free minutes with room to spare. For the one recording he was not allowed to upload, a guest lecture with a confidential case study, Jake set up whisper.cpp on the laptop, and it ground through the file in the background while Ethan read the Word transcript of the week before. The No. 1 model in the world never touched his audio, and that was fine. The best transcription tool is the one that matches the job, and most of the time it is a button you already have.

 If you keep one line from this page

Streaming is for words that must appear now; for anything already recorded, the cheaper batch model, or the Word button you already pay for, does the job better.

$0.54 an hour live, $0.10 an hour recorded, and free with Whisper when the audio must stay on your machine.

Revision note. Written October 2, 2026, the day after Microsoft launched MAI-Transcribe-2-Streaming; prices get a dated update when the introductory rates end. If a recording has been sitting unwatched, I hope it becomes notes this week.

Related