Audio · guides
Text to speech, voice cloning and transcription
By Alberto Gulotta · Updated · 6 min read
This is the entrance to the audio tower. Text to speech that does not sound like a railway announcement, voices cloned from a few minutes of recording, music generated from a sentence, and the reverse of all of it — turning speech back into text.
Audio is where the change of the last few years is most audible, and where the ethical questions are the least abstract. A synthetic voice reading an article is useful and harms nobody. The same technology applied to somebody else’s voice without their agreement is a different act entirely, and it is the reason one wing of this tower has consent in its title rather than in a footnote.
Reading text aloud is largely good news, and it is the reason no wing here is devoted to selling it. The built-in voices on every phone and computer have improved enormously, they cost nothing, they work offline, and for the most common uses — listening to an article, proofreading your own writing by ear, accessibility — they are now genuinely sufficient. Nothing on this page will name a paid service before saying that.
The transcription wing has changed just as much and is less noticed. Turning a recording into accurate text used to be a paid service with a turnaround time; it is now something a laptop does by itself, free, offline, in most languages. That wing is mostly about accuracy, speaker separation and what to do with the result.
Generated music is the newest corner of this subject and the most unsettled. What may be generated, what may be published, and what the terms permit are all genuinely open questions, and no page here will answer them until the answers can be quoted from somebody who decides them.
Three levels are planned, and you are on the first. Five guides exist so far — voice scams, transcribing a recording, voice typing in Google Docs, recording a call, and Audacity plugins — and they are linked below. The rest of the map is written out here as a plan: as each floor arrives it will cover one subject completely and hold the shorter guides underneath it, and nothing will be more than three clicks from the front door.
One practical warning belongs at the door, because it applies to everything in this building and is worth checking. When you upload audio to a service — a recording of a meeting, an interview with a source, a message from a relative — you are handing somebody else a recording of people who did not agree to that. Several of these services state plainly that uploaded audio may be retained or used to improve their systems, and a few say the opposite just as plainly. The difference is written down in each case, every floor here will quote it with the date, and where a free offline tool does the same job the guide names that first. For anything sensitive, offline is not a preference in this tower, it is the recommendation.
Where to start
Four ways in. The first has a free answer that is better than the paid ones assume.
- “I need a recording turned into text.”
- That is speech into text
- “Can I clone my own voice?”
- That is cloned voices
- “Someone used my voice without asking.”
- Read the consent floor in cloned voices
- “I need subtitles for a video.”
- That is a floor in speech into text
Cloned voices
The wing with the consent problem. Every floor here will state the rule before the method: your own voice is yours to clone, and somebody else’s is not yours to use.
Speech into text
Transcription, subtitles and notes from recordings. The wing where free and offline tools now match paid services for most work.
Not covered here. It will not help anybody clone a voice that is not theirs. Every floor in that wing will lead with consent, and the guide about voice scams exists already because this technology is being used against families now.
It will not sell a subscription for something your laptop does free. Reading aloud and transcription both begin with the built-in and offline options, because that is often where the honest answer ends.
It will not pretend the music rights question is settled. It is being argued in courts in several countries, the terms of the services differ from one another, and any floor that touches it will say what is unresolved.
And it will not carry affiliate links, including for the voice platforms, which pay recurring commissions and account for most of what currently ranks in this field.
A last note on disclosure, which runs through the whole tower. Where synthetic audio is used in something an audience will hear, these floors take the position that saying so costs nothing and settles the question before anybody asks it. What holds instead is simple: the free and offline options come first in every wing, and no floor carries an affiliate link.
Written by Alberto Gulotta
Founder and editor of AI Tools Primer, writing from Palermo, Italy. Thirty-five years of taking computers apart, starting with a Commodore 64 — the long version is on the about page.
Something wrong on this page? Write to aitoolsprimer@gmail.com and it gets fixed.
Written on 20 August 2026 · last checked 9 September 2026.
Independence and limits
No affiliate links and no paid placements anywhere on this site. Nobody pays to appear here, and no company has seen this page before you did.
This is general information, not professional advice. Where a page touches money, health, safety or the law, it names its source and the date it was read — and your situation may still differ. See the privacy page and the cookie policy.