About TextToSpeech.tools

What this is

texttospeech.tools turns written English into natural-sounding speech. Paste or import text, pick one of sixteen American or British voices, and listen or download the result as MP3, WAV or SRT subtitles. There is no account, no character limit, no watermark and no charge.

Everything runs inside your browser. The text is turned into speech on your own device and the audio file is written straight to your disk; nothing is sent to a server. The first time you use it, the voice engine is downloaded once — about 330 MB on most computers, about 90 MB on phones and devices without fast graphics — and your browser keeps it, so later visits start instantly and work offline.

Why it is built this way

Online text-to-speech services generate audio on their own servers, and every character costs them something. That is why free plans are capped, downloads carry watermarks and accounts are required. Moving the work to your device removes that cost, which is what makes unlimited and free honest rather than a trial. It also makes the privacy claim checkable: after the first use, disconnect from the internet and the tool keeps working.

The goal is for every part of the tool to be as good as the best alternative, not merely to work: playback that starts while the rest is still being generated, editing a single sentence without regenerating the whole script, pauses you can place exactly, and subtitles timed to the real audio.

What you can expect

The voices are natural for narration, explainers and long reading, and they read numbers, dates, money and common abbreviations in their spoken form. They are English only. Unusual names and brand names are sometimes mispronounced — spell them the way they sound and regenerate just that part. There is no voice cloning and no custom voices.

The site is free and supported by the ads on the page. We don't sell anything, don't collect email addresses and don't run affiliate links. The privacy policy says exactly what is measured and by whom. Every message sent through the contact page is read, and what breaks gets fixed.

Known issues

Phones and older computers generate speech more slowly than it plays; playback then starts after a short wait, and long texts take a while to download. The first generation on a computer can take up to a minute of one-time preparation. Very long texts (several hours of audio) use a lot of memory; if a browser tab struggles, generate a chapter at a time. Scanned PDFs without a text layer can't be imported.

The software components the tool is built on, and their licences, are listed on the licences page.