AI voice and radio imaging: station IDs that don’t sound synthetic
A practical guide to AI radio imaging: writing AI voice scripts for the ear, comping takes, hiding synthetic texture under a bed, and the consent rules for cloning.
Imaging is the best use of AI in small radio, because it’s the one job that used to mean booking another human being. Station IDs, sweepers and sponsor reads are all voice work, and a synthetic voice can now do most of them convincingly. Most of them. The gap between imaging that sounds like a station and imaging that sounds like a text-to-speech demo is almost never the model you picked — it’s the script, the take selection, and about ninety seconds of processing.
1. What radio imaging actually is
Imaging is everything on your station that isn’t a song and isn’t a show — the two seconds of voice over an intro, the station name between records, the promo for Saturday night. It’s called imaging because it’s literally what builds the image of the station in a listener’s head. It’s why a station with a great library and no imaging sounds like a playlist, and one with an average library and tight imaging sounds like radio.
The vocabulary is old, and every AI voice tool and production forum assumes you already have it:
| Element | What it is | Typical length |
|---|---|---|
| Station ID (legal ID) | Names the station, plainly. “You’re listening to Harbour FM.” | 3–6 seconds |
| Sweeper | Voice over music that “sweeps” one song into the next. Carries a slogan, not information. | 5–10 seconds |
| Liner | A dry spoken line over a song’s intro or outro, usually from the host rather than the imaging voice. | 4–8 seconds |
| Bed | An instrumental loop that plays *under* speech. Never heard consciously; entirely responsible for whether speech sounds produced. | 30–120 s, loopable |
| Stinger (shotgun) | A one-hit musical punctuation mark. No voice. Ends a segment or covers an edit. | 1–3 seconds |
| Promo | A trailer for something on your own station — a show, a competition, a guest. | 20–30 seconds |
| Sponsor spot | A paid read for someone else’s business. The element that pays for the others. | 20–30 seconds |
For the music half — jingle packages, how they’re commissioned, what they cost — see radio imaging and jingles. This guide is the voice half, produced with AI, without being caught.
The amateur tell is silence, not music
A new station usually has good songs and nothing between them. Song ends, half a second of nothing, next song. That gap is what listeners hear as “not a real station”, long before they judge your music taste. Four sweepers and one ID fixes it.
2. Why AI voice sounds fake — and the six fixes
People assume synthetic voice is caught by its timbre. It usually isn’t — modern models get the grain of a voice right. What they get wrong is rhythm: every sentence arrives at the same speed, with the same energy arc, landing neatly on a full stop. Humans don’t do that. Humans rush a clause, lean on a word for no reason, and run out of breath in the wrong place. Every fix below is a way of putting that disorder back in.
Fix 1 — Write for the ear, not the eye
Scripts that read well on a page perform worst out loud. Aim for seven to twelve words a line, one idea a line, contractions everywhere. A sentence with two subordinate clauses gives the model a long unbroken run to flatten; cut it into three lines and it has to breathe three times. Read every script aloud once at pace first. If you stumble, the model will too — it will just stumble smoothly, which is worse.
Fix 2 — Treat punctuation as pacing, not grammar
In a voice script, punctuation is a timing instruction. A full stop is a beat. A comma is half a beat. An ellipsis is a breath. A line break is air. So put full stops where grammar wants commas, and a line break wherever you want the voice to land.
Two things reliably break the illusion: semicolons and parentheses, both of which render as an even, unstressed run. Delete them from imaging scripts. Numbers and web addresses are the other trap — write “one-oh-two point five” and “harbour f m dot com”, and spell station names phonetically when the model mangles them (“KAY-ex-eff-em”).
Fix 3 — Deliberately avoid perfectly formed sentences
Real announcers speak in fragments. “All the hits. All day. Harbour FM.” is three grammatical crimes and one good sweeper. The complete version — “You’re listening to Harbour FM, where we play all the hits, all day long” — is longer, flatter and unmistakably generated. Fragments also break up the second-biggest tell, uniform sentence length. Vary it hard. One word. Then eleven. Then three.
Fix 4 — Give the model a character brief, not an adjective
Asking for a “warm, professional, energetic” voice gets you the average of every voice-over ever recorded, which is exactly what it sounds like. Adjectives aren’t direction. A person is.
Write two lines and keep them in a file: who this person is, how old, where from, what they’ve been doing for twenty years, and what they think of what they’re reading. “A 45-year-old night-shift announcer from Manchester who’s done this since analogue, likes the record coming up, and isn’t impressed by much.” Then add the situation — what just happened on air, and who they’re talking to. Whether your tool takes that as a prompt or only as guidance for you, the brief is what makes you pick the *right* voice instead of the first pleasant one.
Fix 5 — Generate many takes and comp the best
This is what separates people who get good results from people who conclude AI voice “isn’t there yet”. Generation is nearly free, so make six to ten takes of every line and throw away nine. Take-to-take variance is real and large: one will land a genuinely odd, human emphasis somewhere in the middle, and that’s the keeper — not the take that is most correct.
For anything longer than a sweeper, comp it: build the finished read from the best take of each line, not the best single pass. That’s ordinary voice-over practice — nobody airs an unedited studio take either. Your bottleneck stops being the model and becomes your own ear, which is the right place for it.
Fix 6 — A bed hides more than a better model does
The least intuitive and most important thing here. Synthetic texture lives in the quiet parts — the unnaturally clean noise floor between words, the too-even sibilance, the room that isn’t there. Put a music bed under the voice and duck it and most of that evidence is simply covered. Add a de-esser and a light compressor and you’ve removed the rest.
People spend weeks shopping for a more expensive voice to fix what a free instrumental loop solves in one pass. Dry AI voice, alone, in silence, is the hardest possible test, and no real station airs dry voice anyway. Don’t sit an exam nobody set.
The listening test that matters
Never judge imaging on headphones in a quiet room. Play it through a phone speaker, in a kitchen, between two of your actual songs. That’s where it will be heard, and it’s far more forgiving — which is why headphone perfectionism wastes so much time.
3. The production recipe, start to finish
One sweeper, blank page to on air, is about ten minutes once you’ve done it twice:
- 1Write the script to the rules above — short lines, fragments allowed, punctuation as pacing, numbers spelled phonetically. Say it out loud before you generate anything.
- 2Generate six to ten takes of the same script. Change nothing between them; you want the model’s variance, not your settings’.
- 3Pick and comp. Choose the take with the odd emphasis. For longer reads, assemble the best line from each pass.
- 4Level and de-ess. Trim the head and tail silence to nothing, de-ess, then a gentle compressor. Keep peaks around −1 dBTP so nothing clips downstream.
- 5Duck it under a bed. Start the loop about half a second before the voice and pull it down 8–12 dB while the voice speaks. Let the bed run a beat past the last word — that tail is what makes it sound like it came from a station.
- 6Export as WAV where you can, MP3 at 256 kbps or better otherwise, matched to the loudness of your music so the ID doesn’t jump out of the mix. Nobex accepts MP3, WAV, FLAC, M4A and OGG.
- 7Schedule it into rotation. Upload it, file it in its own imaging category, and let rule-based rotation place it with separation so the same sweeper can’t come round twice in an hour. Cloud AutoDJ handles the placement and the crossfade; scheduling and shows covers the pieces that have to hit a specific clock time.
Batch, don’t drip
Produce a whole set in one sitting — the voice, the bed and your ear are all calibrated then, and the results sound like they belong together. A good starter pack is eight sweepers, three legal IDs, two time-check lines and one promo. That’s an afternoon, and it covers a month.
4. Which imaging jobs AI voice is actually good at
AI voice isn’t uniformly good across imaging. It’s excellent at short, factual, repeated elements and weak at anything that has to sound like a person who cares. Knowing the line saves you an obviously synthetic promo:
| Element | Length | How often to air it | AI voice fit |
|---|---|---|---|
| Legal / station ID | 3–6 s | Top of the hour, plus every third or fourth break | Excellent — short, factual, identical every time. |
| Sweeper | 5–10 s | Every 2–3 songs | Excellent. Under a bed, at that length, nobody can tell. |
| Time or weather check | 3–5 s | Once or twice an hour | Excellent, and tedious enough that you’d never voice it by hand. |
| Liner | 4–8 s | A few an hour, over song intros | Good — though a real host voice is warmer here if you have one. |
| Podcast top and tail | 5–15 s | Every episode | Good. Consistency beats personality on a repeated intro — see recording shows. |
| Sponsor spot | 20–30 s | Per the contract | Mixed. Fine for factual reads — address, hours, offer. Weak once it needs warmth. Ask the sponsor first. |
| Show promo | 20–30 s | 2–4 an hour in the week before | Weak. Promos sell excitement, and excitement is what generated reads fake worst. Voice this yourself. |
| Memorial, apology, appeal | Any | Rare | Don’t. A synthetic voice here reads as contempt, however good it sounds. |
The pattern: AI takes the volume work, you keep the moments. That’s the same trade as the rest of the stack — AI for radio stations maps where the line falls across programming, metadata and promotion, and AI tools for radio stations covers what to actually run.
5. Pick one voice and stay with it
A station voice works by repetition. It takes a listener something like a dozen exposures before your imaging voice stops being noise and starts being *your station* — the audio equivalent of a logo. A station that changes voice every month never gets there; it just sounds permanently under construction.
The temptation is severe, because a better model ships every few weeks and always sounds nicer in isolation. It won’t sound nicer next to the fourteen pieces already on air. Treat your voice as a brand decision with a switching cost, not a settings dropdown.
- Write down the exact configuration — voice name or ID, model version, every parameter you changed — in a text file next to your scripts. You will need it in six months, and voice libraries do get reshuffled.
- Keep a style sheet of your standard lines and phonetic spellings, so a script written next year sounds like one written today.
- Use a different voice for imaging than for hosting. The imaging voice should sound like the *station*; the host should sound like a person. Same voice for both and listeners can’t tell the furniture from the people.
- If you must switch, re-produce everything in one afternoon. Half-old, half-new imaging is worse than either.
- Two voices maximum — one main, one contrast voice for promos or a specialist show. Three is a mess.
6. Cloning, consent and disclosure
Using a stock synthetic voice from a commercial provider is legally uneventful: check the licence covers broadcast and commercial use — some free tiers pointedly don’t — keep a copy of the terms, move on. Cloning a real person’s voice is the one thing here that can genuinely put your station in trouble, and the law has tightened fast.
Tennessee’s ELVIS Act, effective July 2024, added voice explicitly to that state’s right-of-publicity statute and made unauthorised voice simulation directly actionable. California’s AB 2602, effective January 2025, voids contract terms that let a producer generate a digital replica of a performer without specific, informed consent. Federal digital-replica bills keep being introduced and their status keeps moving — check, don’t assume. In the EU, the AI Act carries transparency duties for published AI-generated audio. This moves quickly, none of it is legal advice, and the sensible posture for a small station is to stay well inside the line rather than near it.
The rule that keeps you out of court
Never clone a voice you don’t have written permission to clone. Not a celebrity’s. Not a politician’s. Not the well-known presenter across town. Not “a voice that sounds a bit like” one either — sound-alike claims are how these cases have historically been won, and they long pre-date AI.
Cloning your own voice, or a co-host’s, is the good use of this: record thirty minutes once, and your station has a voice that produces imaging at three in the morning. Even then, do the paperwork.
- Get consent in writing from anyone whose voice you clone — co-hosts and volunteers included, and including when they’re enthusiastic about it. A one-page note is fine.
- Say what it may be used for. Imaging and promos, yes. Sponsor reads that imply a personal endorsement, probably not — that’s a separate conversation and often a separate fee.
- Put an end date and a deletion clause in it. When someone leaves the station, their clone goes with them. This is the clause people forget and the one that causes the argument.
- Get the sponsor’s agreement before a synthetic voice reads their ad. Some advertisers care a great deal, and finding out afterwards is expensive.
- Follow the disclosure norm. You needn’t caption every sweeper. But if a synthetic voice could be mistaken for a real named person, or is delivering something that sounds like news or a personal endorsement, say so on air or on your station page.
None of this touches your music rights — synthetic voice over a commercial track is still a use of that track. Internet radio licensing covers what you need, and why a licence in your own name is an asset of your station rather than something borrowed from a host.
7. Getting it on air without wearing it out
Produced imaging that airs badly sounds worse than none, and two failure modes account for nearly all of it: too much, and too repetitive. The workable ratio for a music station is one imaging element every two to three songs, with the legal ID at the top of the hour. Push past that and you sound like a jingle demo reel; drop below it and the gaps come back. Carry enough that nothing repeats inside an hour — eight sweepers does it — and file them as their own category so rotation rules enforce the spacing instead of you hoping.
Then rotate the set itself. Replace roughly a third of your sweepers each quarter, and kill anything seasonal the week it expires. A promo for a show that ended in March, still airing in June, undoes more credibility than the promo ever built. That’s the real running cost of imaging — about an hour a quarter, affordable only because you’re no longer booking studio time to change one line.
Steal the structure, not the audio
Listen to a commercial station in your format and write down, minute by minute, what sits between the songs. That pattern of ID, sweeper, liner and promo is decades of audience testing you can copy for free. Copying their actual audio is a different matter, and it is not free.
8. AI imaging questions, asked and answered
Can I use AI voice for my radio station IDs and sweepers?
Yes, and it’s the strongest use of AI in small radio. Short factual elements — legal IDs, sweepers, time checks — are the easiest thing for a synthetic voice to carry convincingly, especially placed under a music bed. Longer emotional pieces like show promos still sound better voiced by a person.
Why does my AI voice sound fake even on a good model?
Almost always the script and the mix, not the model. Synthetic voice is caught by uniform rhythm, so write in short fragments with one idea per line, use punctuation as pacing, generate six to ten takes and keep the one with an odd emphasis. Then duck it under a music bed — synthetic texture hides in the quiet parts, and a bed covers them.
Is it legal to clone someone’s voice for radio imaging?
Only with that person’s written permission. Voice-likeness protection has tightened sharply — Tennessee’s ELVIS Act added voice to right-of-publicity law in 2024, and California’s AB 2602 restricted digital-replica contract terms from 2025 — and sound-alike imitation has long been actionable in its own right. Cloning your own or a consenting co-host’s voice is fine; cloning a celebrity or a rival presenter is the thing that will actually cause you trouble.
How much imaging does a small radio station need?
A workable starter set is eight sweepers, three legal IDs and a couple of time-check lines — enough that nothing repeats within an hour. Air roughly one element every two or three songs, with the legal ID at the top of the hour, and replace about a third of the set each quarter so it doesn’t go stale.
Do I have to tell listeners a voice is AI-generated?
You don’t need to label every sweeper, but you should disclose whenever a listener could be misled about who is speaking: a synthetic voice presented as a named real person, anything resembling news, or a personal endorsement in an advert. Disclosure expectations for AI-generated content are expanding, so a short line on your station page is cheap insurance.
How long should a radio sweeper be?
Five to ten seconds — long enough to carry a slogan, short enough to sit over a song intro without stepping on the vocal. Station IDs are shorter at three to six seconds, and promos and sponsor spots run twenty to thirty.
Keep reading
Ready to put this into practice?
Start free and be on air in under five minutes — no software to install.
Start streaming free