Skip to main content

Free AI Voice and Audio Tools: Where the Free Tier Actually Stops

Six jobs, the tools that do them at no cost, and the exact limit the short videos never mention.

A studio microphone on a desk, representing free AI voice and audio tools.

Every one of these tools gets described the same way in a twenty-second video: free. That word is doing an enormous amount of work. Some of these products are free the way a public library is free. Others are free the way a first taste is free, and the limit arrives about four minutes into real use.

This is the version with the numbers in it. For each job below you get the tool, what the free tier actually includes, and the point at which you either pay or change approach. Where a limit is published by the vendor we cite it; where it is not published we say so instead of guessing.

Turning written text into a spoken voice

Two very different free options exist here and people conflate them constantly.

The hosted option is ElevenLabs. Its free plan costs nothing per month and includes ten thousand credits per month, which covers text to speech, speech to text, sound effects and voice design. Ten thousand credits is roughly ten minutes of generated speech, so it is enough to voice a short video or test a script, and nowhere near enough to narrate a course. There is no commercial licence on the free tier; that starts on the paid Starter plan.

Our short on generating a voiceover from text. The article above adds the credit limits the video does not have room for.

The other option is the speech engine already sitting on your machine. Every modern operating system and browser ships a text to speech engine that costs nothing and has no monthly cap at all. The voices are noticeably flatter than a hosted neural voice, but for internal use, accessibility, or listening to your own drafts read back, the quality difference stops mattering and the quota difference starts to.

The decision rule is simple: if the audio is going in front of an audience, the hosted voice is worth the credits. If the audio is for you, use the engine you already have and keep the credits.

Transcribing a meeting or a recording

This is the category where the free tiers differ most, and where most people pick wrong.

Otter.ai is the tool named most often. Its Basic plan is free and includes three hundred transcription minutes per user per month, with a separate three hundred minute monthly allowance for imported audio and video files. Neither allowance rolls over. Three hundred minutes is five hours, which sounds generous until you attend three hour-long calls a week.

Our short on live meeting transcription. The monthly minute caps are in the section above.

Google Docs has voice typing built in, free, with no monthly minute cap. It transcribes live from your microphone into an open document. It has no speaker labels and no recording, so it is a poor meeting recorder and a very good dictation tool.

The option almost nobody in these videos mentions is running the model yourself. OpenAI released Whisper as an open source speech recognition model, so you can transcribe on your own machine with no account, no upload, and no monthly limit. It needs a one-off setup and a reasonably recent computer, and it is the only option on this list where the audio never leaves your hardware. For anything confidential, that is not a nice-to-have.

So: light and occasional, use Otter's free tier. Dictation, use Google Docs. High volume or sensitive, run Whisper locally and stop counting minutes.

Separating vocals from a backing track

Vocal isolation went from a studio process to a background task in about three years. Most of the web tools that do it are free for a preview and charge for the full-length export, which is the business model the videos rarely mention.

Our short on vocal isolation. The export limits and the open source alternative are covered above.

The way around it is the open source route. Ultimate Vocal Remover is a free desktop application that runs the separation models locally, with no export cap and no watermark. It is slower than a hosted service on an old machine and it asks you to pick a model rather than pressing one button, which is the actual trade you are making.

One thing worth stating plainly: separating a track does not give you the right to publish it. The tool is free; the underlying recording is still somebody's copyright.

Turning a document into something you can listen to

Google's notebook product generates a spoken audio overview from documents you upload, in the form of two synthetic hosts discussing your material. It is free to use, and the vendor's own help documentation is the place to check current limits, because they have moved more than once and the product has since been folded into the Gemini branding.

Our short on document to audio overviews, plus the caveat about how loosely they summarise.

It is genuinely useful for turning a long report into something you can absorb while driving. It is not a research tool. The hosts sound confident about material they have summarised loosely, and that confidence is exactly the failure mode to watch for. Treat the output as a trailer for the document, not a replacement for reading it.

Live speech translation

Conversational voice mode in a general assistant app will now translate a spoken conversation in near real time, and the free tier of most assistant apps includes some voice usage. This is the category where the free allowance is least clearly documented, so plan on it changing.

Our short on live voice translation, and where turn-taking breaks it.

What actually limits this in practice is not the quota, it is the turn-taking. It works well for a two-person exchange where both people pause. It falls apart in a fast group conversation, and it will not save you in a negotiation where precision matters.

Voice cloning and synthetic singing

This is the one to think about before using, not after.

Our short on synthetic singing. Read the consent and disclosure section before you use it.

Cloning a voice from a short sample is now trivial and often free at low volume. The technical barrier is gone; the consent question is not. Cloning your own voice for your own content is uncontroversial. Cloning somebody else's, including a public figure's, exposes you to publicity-rights and platform-policy problems that no free tier covers. Most major platforms now require synthetic voice disclosure on uploads. If you use it, label it.

How to decide, in one rule

Free tiers are priced to be comfortable exactly until a tool becomes part of your routine. The useful test is not what the tool costs today, it is what happens in week six.

The short videos are accurate about what these tools can do. They are just quiet about where the free part stops, and that is usually the only detail that decides whether a tool survives contact with your actual workload.

Frequently Asked Questions

What does the ElevenLabs free plan actually include?

The published free plan costs nothing per month and includes ten thousand credits per month across text to speech, speech to text, sound effects and voice design. Commercial use is not included on the free plan; it begins on the paid Starter tier.

How many minutes does Otter.ai transcribe for free?

Otter.ai's Basic plan is free and includes three hundred transcription minutes per user each month, plus a separate three hundred minute allowance for imported files. Neither allowance rolls over to the next month.

Is there a genuinely unlimited free option for transcription?

Yes, if you run the model yourself. OpenAI released Whisper as an open source speech recognition model, so it can be run locally with no account, no upload and no monthly cap. The cost is a one-off setup and your own hardware.

Can I publish audio made with a free AI voice tool?

Check the licence rather than the price. Several free tiers exclude commercial use outright, and separating or cloning existing audio does not transfer any rights in the underlying recording or voice.

Sources

Related Articles