← Back to the island

The audio tower · the atrium

Text to speech, voice cloning and transcription

This is the entrance to the audio tower. Text to speech that does not sound like a railway announcement, voices cloned from a few minutes of recording, music generated from a sentence, and the reverse of all of it — turning speech back into text.

Audio is where the change of the last few years is most audible, and where the ethical questions are the least abstract. A synthetic voice reading an article is useful and harms nobody. The same technology applied to somebody else’s voice without their agreement is a different act entirely, and it is the reason one wing of this tower has consent in its title rather than in a footnote.

The text-to-speech wing itself is largely good news, and unusually so. The built-in voices on every phone and computer have improved enormously, they cost nothing, they work offline, and for the most common uses — listening to an article, proofreading your own writing by ear, accessibility — they are now genuinely sufficient. The floors say that before naming a single paid service.

The transcription wing has changed just as much and is less noticed. Turning a recording into accurate text used to be a paid service with a turnaround time; it is now something a laptop does by itself, free, offline, in most languages. Those floors are mostly about accuracy, speaker separation and what to do with the result.

The music wing is the newest and the most unsettled. What may be generated, what may be published, and what the terms permit are all genuinely open questions, and those floors say so rather than guessing.

Three levels, and you are on the first. This atrium lists the floors; a floor covers one subject completely and holds the shorter guides underneath it. Nothing here is more than three clicks from the front door.

One practical warning belongs at the door, because it applies across three of the four wings and almost nobody checks it. When you upload audio to a service — a recording of a meeting, an interview with a source, a message from a relative — you are handing somebody else a recording of people who did not agree to that. Several of these services state plainly that uploaded audio may be retained or used to improve their systems, and a few say the opposite just as plainly. The difference is written down in each case, every floor here quotes it with the date, and where a free offline tool does the same job the floor names that first. For anything sensitive, offline is not a preference on these floors, it is the recommendation.

Where to start

Seven ways in. The first two have better free answers than most people realise.

“I want an article read aloud to me.”
Start at reading aloud
“I need a voice for a video.”
Go to reading aloud
“I need a recording turned into text.”
That is speech into text
“Can I clone my own voice?”
That is cloned voices
“Someone used my voice without asking.”
Read the consent floor in cloned voices
“I need background music I can actually use.”
That is generated music
“I need subtitles for a video.”
That is a floor in speech into text

Reading aloud

Text to speech, from the voices already on your devices to the paid services. Each floor says what it costs, what it can be used for commercially, and whether it works without a connection.

The voices you already haveWhat every phone and computer does for free, which is more than most people know.Being built
Text to speech servicesWhat each publishes about voices, pricing and permitted use of the output.Being built
Listening to articlesTurning reading into listening, on a commute, without a subscription.Being built
Voices for videoWhat an AI voice generator needs to do for narration that a screen reader does not, and disclosure.Being built
AccessibilityThe built-in tools designed for this, and how to set them up properly.Being built
Languages and accentsWhere the quality is good, and where it is still obviously wrong.Being built

Cloned voices

The wing with the consent problem. Every floor here states the rule before the method: your own voice is yours to clone, and somebody else’s is not yours to use.

Cloning your own voiceWhat it takes, what it is genuinely useful for, and what the service keeps.Being built
Consent and the lawWhere using somebody’s voice becomes unlawful, and how that differs by country.Being built
Voice scamsThe call that sounds like a relative, and the family password that defeats it.Being built
Detecting synthetic audioWhat detection can and cannot do, stated with the same caution as elsewhere.Being built

Speech into text

Transcription, subtitles and notes from recordings. The wing where free and offline tools now match paid services for most work.

Transcribing a recordingSpeech to text on your own machine, free and offline, and how accurate it actually is.Being built
SubtitlesGenerating and correcting them, and the formats different platforms want.Being built
Meeting notesWhat the automatic note-takers produce, and what everybody in the room should be told.Being built
Interviews and researchSpeaker separation, timestamps, and keeping the recording private.Being built
Cleaning up audioRemoving noise and levelling voices, mostly with free tools.Being built

Generated music

The newest wing and the most legally unsettled. Floors here describe what the tools do and what their terms say, and are explicit about what has not been decided.

What music generators doHow they are used in practice, and where the output is obviously limited.Being built
Using it in a videoWhat the terms say you may publish, and what platforms do about it.Being built
Free alternativesLicensed libraries that avoid the question entirely, and how their licences work.Being built
Separating a trackIsolating vocals or instruments, what it is used for, and where it is not allowed.Being built

What this tower will not do

It will not help anybody clone a voice that is not theirs. Every floor in that wing leads with consent, and the floor about voice scams exists because this technology is already being used against families.

It will not sell a subscription for something your laptop does free. Both the reading-aloud and the transcription wings begin with the built-in and offline options, because for most readers that is where the honest answer ends.

It will not pretend the music rights question is settled. It is being argued in courts in several countries, the terms of the services differ from one another, and floors that touch it say what is unresolved.

And it will not carry affiliate links, including for the voice platforms, which pay recurring commissions and account for most of what currently ranks in this field.

A last note on disclosure, which runs through the whole tower. Where synthetic audio is used in something an audience will hear, these floors take the position that saying so costs nothing and settles the question before anybody asks it. What holds instead is simple: the free and offline options come first in every wing, and no floor carries an affiliate link.

Written by Alberto Gulotta

Founder and editor of AI Tools Primer, writing from Palermo, Italy. Thirty-five years of taking computers apart, starting with a Commodore 64 — the long version is on the about page.

Something wrong on this page? Write to aitoolsprimer@gmail.com and it gets fixed.

Independence and limits

No affiliate links and no paid placements anywhere on this site. Nobody pays to appear here, and no company has seen this page before you did.

This is general information, not professional advice. Where a page touches money, health, safety or the law, it names its source and the date it was read — and your situation may still differ. See the privacy page and the cookie policy.