Skip to main content

Free AI Voice and Audio Tools: Where the Free Tier Actually Stops

Six jobs, the tools that do them at no cost, and the exact limit the short videos never mention.

A studio microphone on a desk, representing free AI voice and audio tools.

Free voice tools can be useful, but a free plan may limit minutes, file imports, exports or commercial use. Match the limits to the recording you need to process before investing time in a workflow.

This guide compares hosted services and tools you can run locally. The September 12 update checks Otter’s current limits and corrects broad quality and privacy claims. Other product sections retain their earlier research; this is not a fresh test of every service or embedded video.

Turning written text into a spoken voice

Two very different free options exist here and people conflate them constantly.

The hosted option is ElevenLabs. Its free plan costs nothing per month and includes ten thousand credits per month, which covers text to speech, speech to text, sound effects and voice design. Ten thousand credits is roughly ten minutes of generated speech, so it is enough to voice a short video or test a script, and nowhere near enough to narrate a course. There is no commercial licence on the free tier; that starts on the paid Starter plan.

Our short on generating a voiceover from text. The article above adds the credit limits the video does not have room for.

Check the text-to-speech options already available on your device. Available voices, offline support and export controls vary by operating system and application. Read the same short script through each option, including a name, a number and a pause, before deciding whether a hosted voice adds enough value for your use.

For published audio, compare intelligibility, pronunciation, editing control and the permitted use of the output. A paid plan does not by itself establish better quality; an installed voice may be sufficient for your recording.

Transcribing a meeting or a recording

For transcription, check the length of each recording, the number of files you will upload and the monthly workload separately.

Otter Basic lists 300 transcription minutes per user each month and a 30-minute limit per conversation. File uploads also have a stricter cap: 3 audio or video imports per user across the lifetime of the account. The pricing table separately lists 300 monthly imported-file minutes, but unused minutes do not create more imports. Check both limits before choosing it for recordings. Otter pricing, checked September 12, 2026.

Our short on live meeting transcription. The monthly minute caps are in the section above.

Google Docs has voice typing built in, free, with no monthly minute cap. It transcribes live from your microphone into an open document. It has no speaker labels and no recording, so it is a poor meeting recorder and a very good dictation tool.

Whisper can run locally on your own hardware without a hosted transcription account or a vendor minute allowance. Processing time, memory and setup are still real costs. A local installation can keep the transcription on the device; check the application you install and any cloud sync or integrations rather than assuming every Whisper-based service is local.

For occasional live meetings, first check whether Otter’s per-conversation and monthly limits fit. For uploaded recordings, count lifetime imports as well. For repeated transcription, compare the work of running Whisper locally with the cost and controls of a hosted plan.

Separating vocals from a backing track

Vocal isolation went from a studio process to a background task in about three years. Most of the web tools that do it are free for a preview and charge for the full-length export, which is the business model the videos rarely mention.

Our short on vocal isolation. The export limits and the open source alternative are covered above.

The way around it is the open source route. Ultimate Vocal Remover is a free desktop application that runs the separation models locally, with no export cap and no watermark. It is slower than a hosted service on an old machine and it asks you to pick a model rather than pressing one button, which is the actual trade you are making.

One thing worth stating plainly: separating a track does not give you the right to publish it. The tool is free; the underlying recording is still somebody's copyright.

Turning a document into something you can listen to

Google's notebook product generates a spoken audio overview from documents you upload, in the form of two synthetic hosts discussing your material. It is free to use, and the vendor's own help documentation is the place to check current limits, because they have moved more than once and the product has since been folded into the Gemini branding.

Our short on document to audio overviews, plus the caveat about how loosely they summarise.

It is genuinely useful for turning a long report into something you can absorb while driving. It is not a research tool. The hosts sound confident about material they have summarised loosely, and that confidence is exactly the failure mode to watch for. Treat the output as a trailer for the document, not a replacement for reading it.

Live speech translation

Conversational voice mode in a general assistant app will now translate a spoken conversation in near real time, and the free tier of most assistant apps includes some voice usage. This is the category where the free allowance is least clearly documented, so plan on it changing.

Our short on live voice translation, and where turn-taking breaks it.

Before using voice translation for a real conversation, try the language pair and turn-taking pattern you need. Check a few phrases with a fluent speaker, especially numbers and names. An app’s free or paid status is not a measure of translation accuracy.

Voice cloning and synthetic singing

This is the one to think about before using, not after.

Our short on synthetic singing. Read the consent and disclosure section before you use it.

Before cloning a voice, establish permission to use the recording and voice, and check the service’s allowed uses and the publishing platform’s disclosure requirements. Do not treat a working demo or a free allowance as permission to publish.

How to decide, in one rule

Estimate a typical month before choosing: how many recordings, how long each is, how many uploads, and whether you need an export or commercial licence. A recurring job may fit a free plan; a single long recording may not.

Treat the embedded shorts as introductions. Check the written limits and source pages before starting; this update does not independently validate every demonstration in the videos.

Frequently Asked Questions

What does the ElevenLabs free plan actually include?

The published free plan costs nothing per month and includes ten thousand credits per month across text to speech, speech to text, sound effects and voice design. Commercial use is not included on the free plan; it begins on the paid Starter tier.

How many minutes does Otter.ai transcribe for free?

Otter Basic includes 300 monthly transcription minutes per user, a 30-minute conversation limit, and only 3 lifetime audio/video file imports per user. Monthly minutes do not roll over; the lifetime import count does not reset each month.

Can I transcribe locally without a vendor minute quota?

A local Whisper installation has no hosted-service minute quota. You supply the hardware and setup, and processing remains bounded by available memory, storage and time. Check that the application runs locally rather than sending audio to a hosted service.

Can I publish audio made with a free AI voice tool?

Check the licence rather than the price. Several free tiers exclude commercial use outright, and separating or cloning existing audio does not transfer any rights in the underlying recording or voice.

Sources

Related Articles