Free voice tools can be useful, but a free plan may limit minutes, file imports, exports or commercial use. Match the limits to the recording you need to process before investing time in a workflow.
This guide compares hosted services and tools you can run locally. The September 12 update checks Otter’s current limits and corrects broad quality and privacy claims. Other product sections retain their earlier research; this is not a fresh test of every service or embedded video.
Turning written text into a spoken voice
Two very different free options exist here and people conflate them constantly.
The hosted option is ElevenLabs. Its free plan costs nothing per month and includes ten thousand credits per month, which covers text to speech, speech to text, sound effects and voice design. Ten thousand credits is roughly ten minutes of generated speech, so it is enough to voice a short video or test a script, and nowhere near enough to narrate a course. There is no commercial licence on the free tier; that starts on the paid Starter plan.
Check the text-to-speech options already available on your device. Available voices, offline support and export controls vary by operating system and application. Read the same short script through each option, including a name, a number and a pause, before deciding whether a hosted voice adds enough value for your use.
For published audio, compare intelligibility, pronunciation, editing control and the permitted use of the output. A paid plan does not by itself establish better quality; an installed voice may be sufficient for your recording.
Transcribing a meeting or a recording
For transcription, check the length of each recording, the number of files you will upload and the monthly workload separately.
Otter Basic lists 300 transcription minutes per user each month and a 30-minute limit per conversation. File uploads also have a stricter cap: 3 audio or video imports per user across the lifetime of the account. The pricing table separately lists 300 monthly imported-file minutes, but unused minutes do not create more imports. Check both limits before choosing it for recordings. Otter pricing, checked September 12, 2026.
Google Docs has voice typing built in, free, with no monthly minute cap. It transcribes live from your microphone into an open document. It has no speaker labels and no recording, so it is a poor meeting recorder and a very good dictation tool.
Whisper can run locally on your own hardware without a hosted transcription account or a vendor minute allowance. Processing time, memory and setup are still real costs. A local installation can keep the transcription on the device; check the application you install and any cloud sync or integrations rather than assuming every Whisper-based service is local.
For occasional live meetings, first check whether Otter’s per-conversation and monthly limits fit. For uploaded recordings, count lifetime imports as well. For repeated transcription, compare the work of running Whisper locally with the cost and controls of a hosted plan.
Separating vocals from a backing track
Vocal isolation went from a studio process to a background task in about three years. Most of the web tools that do it are free for a preview and charge for the full-length export, which is the business model the videos rarely mention.
The way around it is the open source route. Ultimate Vocal Remover is a free desktop application that runs the separation models locally, with no export cap and no watermark. It is slower than a hosted service on an old machine and it asks you to pick a model rather than pressing one button, which is the actual trade you are making.
One thing worth stating plainly: separating a track does not give you the right to publish it. The tool is free; the underlying recording is still somebody's copyright.
Turning a document into something you can listen to
Google's notebook product generates a spoken audio overview from documents you upload, in the form of two synthetic hosts discussing your material. It is free to use, and the vendor's own help documentation is the place to check current limits, because they have moved more than once and the product has since been folded into the Gemini branding.
It is genuinely useful for turning a long report into something you can absorb while driving. It is not a research tool. The hosts sound confident about material they have summarised loosely, and that confidence is exactly the failure mode to watch for. Treat the output as a trailer for the document, not a replacement for reading it.
Live speech translation
Conversational voice mode in a general assistant app will now translate a spoken conversation in near real time, and the free tier of most assistant apps includes some voice usage. This is the category where the free allowance is least clearly documented, so plan on it changing.
Before using voice translation for a real conversation, try the language pair and turn-taking pattern you need. Check a few phrases with a fluent speaker, especially numbers and names. An app’s free or paid status is not a measure of translation accuracy.
Voice cloning and synthetic singing
This is the one to think about before using, not after.
Before cloning a voice, establish permission to use the recording and voice, and check the service’s allowed uses and the publishing platform’s disclosure requirements. Do not treat a working demo or a free allowance as permission to publish.
How to decide, in one rule
Estimate a typical month before choosing: how many recordings, how long each is, how many uploads, and whether you need an export or commercial licence. A recurring job may fit a free plan; a single long recording may not.
- One-off or occasional job — check whether the entire file fits the free allowance before uploading it.
- Something you will do weekly — compare monthly allowances, file limits and local processing time using your actual workload.
- Anything confidential — check permission, storage, retention and data-use terms. Use a properly configured local workflow when the material must remain on your device; a free price alone tells you nothing about data practices.
- Anything published — check the licence, not just the price. Several free tiers explicitly exclude commercial use.
Treat the embedded shorts as introductions. Check the written limits and source pages before starting; this update does not independently validate every demonstration in the videos.
Frequently Asked Questions
What does the ElevenLabs free plan actually include?
The published free plan costs nothing per month and includes ten thousand credits per month across text to speech, speech to text, sound effects and voice design. Commercial use is not included on the free plan; it begins on the paid Starter tier.
How many minutes does Otter.ai transcribe for free?
Otter Basic includes 300 monthly transcription minutes per user, a 30-minute conversation limit, and only 3 lifetime audio/video file imports per user. Monthly minutes do not roll over; the lifetime import count does not reset each month.
Can I transcribe locally without a vendor minute quota?
A local Whisper installation has no hosted-service minute quota. You supply the hardware and setup, and processing remains bounded by available memory, storage and time. Check that the application runs locally rather than sending audio to a hosted service.
Can I publish audio made with a free AI voice tool?
Check the licence rather than the price. Several free tiers exclude commercial use outright, and separating or cloning existing audio does not transfer any rights in the underlying recording or voice.
Sources
- ElevenLabs Pricing Page – ElevenLabs Pricing for Creators & Businesses of All Sizes
- Otter.ai Pricing Page – Pricing | Otter.ai
- OpenAI Whisper Repository – GitHub - openai/whisper: Robust Speech Recognition via Large-Scale Weak Supervision · GitHub
- Google Docs Editors Help – Type & edit with your voice - Google Docs Editors Help
- Gemini Notebook Help – Frequently asked questions - Gemini Notebook Help
- Ultimate Vocal Remover Repository – GitHub - Anjok07/ultimatevocalremovergui: GUI for a Vocal Remover that uses Deep Neural Networks. · GitHub