The audio tower · the atrium
Text to speech, voice cloning and transcription
This is the entrance to the audio tower. Text to speech that does not sound like a railway announcement, voices cloned from a few minutes of recording, music generated from a sentence, and the reverse of all of it — turning speech back into text.
Audio is where the change of the last few years is most audible, and where the ethical questions are the least abstract. A synthetic voice reading an article is useful and harms nobody. The same technology applied to somebody else’s voice without their agreement is a different act entirely, and it is the reason one wing of this tower has consent in its title rather than in a footnote.
The text-to-speech wing itself is largely good news, and unusually so. The built-in voices on every phone and computer have improved enormously, they cost nothing, they work offline, and for the most common uses — listening to an article, proofreading your own writing by ear, accessibility — they are now genuinely sufficient. The floors say that before naming a single paid service.
The transcription wing has changed just as much and is less noticed. Turning a recording into accurate text used to be a paid service with a turnaround time; it is now something a laptop does by itself, free, offline, in most languages. Those floors are mostly about accuracy, speaker separation and what to do with the result.
The music wing is the newest and the most unsettled. What may be generated, what may be published, and what the terms permit are all genuinely open questions, and those floors say so rather than guessing.
Three levels, and you are on the first. This atrium lists the floors; a floor covers one subject completely and holds the shorter guides underneath it. Nothing here is more than three clicks from the front door.
One practical warning belongs at the door, because it applies across three of the four wings and almost nobody checks it. When you upload audio to a service — a recording of a meeting, an interview with a source, a message from a relative — you are handing somebody else a recording of people who did not agree to that. Several of these services state plainly that uploaded audio may be retained or used to improve their systems, and a few say the opposite just as plainly. The difference is written down in each case, every floor here quotes it with the date, and where a free offline tool does the same job the floor names that first. For anything sensitive, offline is not a preference on these floors, it is the recommendation.
Where to start
Seven ways in. The first two have better free answers than most people realise.
- “I want an article read aloud to me.”
- Start at reading aloud
- “I need a voice for a video.”
- Go to reading aloud
- “I need a recording turned into text.”
- That is speech into text
- “Can I clone my own voice?”
- That is cloned voices
- “Someone used my voice without asking.”
- Read the consent floor in cloned voices
- “I need background music I can actually use.”
- That is generated music
- “I need subtitles for a video.”
- That is a floor in speech into text
Reading aloud
Text to speech, from the voices already on your devices to the paid services. Each floor says what it costs, what it can be used for commercially, and whether it works without a connection.
Cloned voices
The wing with the consent problem. Every floor here states the rule before the method: your own voice is yours to clone, and somebody else’s is not yours to use.
Speech into text
Transcription, subtitles and notes from recordings. The wing where free and offline tools now match paid services for most work.
Generated music
The newest wing and the most legally unsettled. Floors here describe what the tools do and what their terms say, and are explicit about what has not been decided.
What this tower will not do
It will not help anybody clone a voice that is not theirs. Every floor in that wing leads with consent, and the floor about voice scams exists because this technology is already being used against families.
It will not sell a subscription for something your laptop does free. Both the reading-aloud and the transcription wings begin with the built-in and offline options, because for most readers that is where the honest answer ends.
It will not pretend the music rights question is settled. It is being argued in courts in several countries, the terms of the services differ from one another, and floors that touch it say what is unresolved.
And it will not carry affiliate links, including for the voice platforms, which pay recurring commissions and account for most of what currently ranks in this field.
A last note on disclosure, which runs through the whole tower. Where synthetic audio is used in something an audience will hear, these floors take the position that saying so costs nothing and settles the question before anybody asks it. What holds instead is simple: the free and offline options come first in every wing, and no floor carries an affiliate link.
Written by Alberto Gulotta
Founder and editor of AI Tools Primer, writing from Palermo, Italy. Thirty-five years of taking computers apart, starting with a Commodore 64 — the long version is on the about page.
Something wrong on this page? Write to aitoolsprimer@gmail.com and it gets fixed.
Independence and limits
No affiliate links and no paid placements anywhere on this site. Nobody pays to appear here, and no company has seen this page before you did.
This is general information, not professional advice. Where a page touches money, health, safety or the law, it names its source and the date it was read — and your situation may still differ. See the privacy page and the cookie policy.