Suno prompts for audiobook and storytelling background music

Writing “instrumental only, no vocals” in Suno’s Styles field is a suggestion, not a control. The Styles field describes what you want the music to sound like — Suno reads the phrase as a stylistic hint and will still hand you a track with a choir pad in it. The vocals are switched off by the Instrumental toggle and the Exclude field, both of which sit in Custom Mode.

That distinction is why most Suno prompt lists for narration don’t work. Below is the order of operations that kills the vocals at the control, gets the arrangement out of your voice’s way, solves the length problem, and only then picks a palette.

Step 1: Switch off vocals at the control

In Custom Mode, toggle Instrumental. This is the reliable method — it changes what the model generates rather than describing a preference to it. Suno’s own documentation lists it as the alternative to entering lyrics, not as a style descriptor.

The Instrumental toggle removes lead vocals. It does not reliably remove wordless voice — choir pads, breathy “ahh” textures, and vocal chops are treated by the model as instrumentation, and they will still land on top of your narration. For those, use Exclude: click Advanced Options in Custom Mode, and the menu opens with the Exclude field. Enter what you don’t want.

Exclude field — paste this as your narration baseline:

vocals, choir, vocal pad, vocal chops, humming, whistling, spoken word, lead melody

Keep the Exclude list short. It is a steer, not a guaranteed filter — a list of thirty terms dilutes each one. Six to ten is the working range.

Step 2: Get the arrangement out of the voice’s way

An instrumental track is not automatically a narration bed. Most instrumentals are written to be listened to, which means they have a melody, a dynamic arc, and energy in the same frequency range as a speaking voice. All three fight narration.

Register

An adult speaking voice has a fundamental frequency in the 85–255 Hz range, with harmonics that contribute to intelligibility reaching several kilohertz above that. Instruments that live in the same space — solo piano in its middle octaves, acoustic guitar strummed in open position, tenor sax, muted trumpet — will mask consonants. Instruments that sit clearly above or below it, such as a low string drone, sub-bass pulse, celeste, glockenspiel, or high sustained pad, leave the voice a clear channel.

This is why “solo piano, melancholic” is the single most over-recommended and worst-performing narration prompt in circulation.

Melody

A listener cannot follow two melodic lines at once, and narration is a melodic line. Ask for texture instead: drones, sustained pads, ostinatos, arpeggio figures, and slow harmonic movement all lend a scene emotional color without demanding attention. Put lead melody and melodic hook in Exclude and use words like underscore, bed, drone, sustained, and textural in Styles.

Dynamics

Cinematic music is built on contrast — a swell that goes from a whisper to a wall. Under narration, that means either the quiet passages vanish, or the loud ones bury you, and no amount of ducking automation fixes a track whose loudest moment is 20 dB above its quietest. Ask for a flat dynamic profile explicitly: steady dynamics, no swells, no crescendo, consistent level throughout.

The two-field template

Narration work is split across two inputs, not one. Styles describes the music; Exclude protects the voice.

Styles:
[setting or genre], [2-4 named instruments], [mood], [BPM] BPM,
sparse arrangement, steady dynamics, textural underscore

Exclude:
vocals, choir, vocal chops, lead melody, crescendo, cymbal crashes

Name specific instruments rather than families — “solo cello and low string drone” produces a far more predictable result than “orchestral.” Give tempo as a number. And use one mood, not three; contradictory emotional directions get averaged into something characterless.

Step 3: Solve length before you solve style

Suno does not generate loops. It generates songs — with an intro, an arc, and an ending. Current models produce up to around eight minutes in a single generation, which sounds like plenty until you try to score a forty-minute chapter.

Three approaches, in order of how much work they cost you:

  1. Generate long, then cut a loop. Take a full-length generation into your editor, find a section where the arrangement has settled into steady state, and cut a 60–90 second region at a musically sensible point. Crossfade the seam. This produces the cleanest loops because you are choosing the calmest passage rather than accepting whatever the intro happens to be.
  2. Use Extend. Open the track’s menu, choose Remix/Edit, then Extend. Drag the marker to set how much of the original to keep and where the extension begins. Quick Extend continues without a new prompt; entering style details lets you steer the continuation. Once you have the desired length, use Get Whole Song to render it as a single file.
  3. Extract stems. If a generation is right except for one element, choose Get Stems from the Edit menu. Suno splits the track into separate parts you can mute or rebalance — the fix for a bed that’s perfect apart from a snare that keeps poking through your dialogue.

Step 4: Pick a palette

These are ordered by how hard each palette is to keep out of the way of a voice, not by genre. The first group is nearly foolproof. The last group will fight you, and the notes explain why.

Sleep, meditation and ASMR — easiest

Almost no melodic content, very slow harmonic movement, and a natural home in the frequency extremes. This palette is the most forgiving under narration and the best place to start if you’re testing your chain.

Sleep story ambience, slow evolving synth pads, distant soft rain,
deeply calming, 40 BPM, sparse arrangement, steady dynamics, no percussion

Guided meditation underscore, low drone, single sustained singing bowl,
spacious and still, 45 BPM, textural, steady dynamics

Forest night ambience, low string drone, distant wind texture,
warm and hypnotic, 42 BPM, sparse, no percussion, steady dynamics

Tension, horror and thriller — easy

Dread is built from sustained dissonance rather than melody, which is exactly what you want. The trap is the stinger: ask for tension, and Suno will often deliver sudden string stabs that spike the level mid-sentence. Exclude them.

Horror narration underscore, dissonant low strings, sub bass drone,
unsettling and sparse, 40 BPM, steady dynamics, no percussion
Exclude: stingers, string stabs, crescendo, sudden hits, vocals, choir

Slow-burn thriller bed, sparse cello pulses, quiet ticking texture,
restrained and anxious, 70 BPM, steady dynamics, minimal percussion

Haunted interior ambience, detuned music box, tape hiss, long reverb tail,
eerie, 48 BPM, sparse, no lead melody

Sci-fi and futuristic — easy

Synth pads occupy the high and low registers naturally and rarely carry a tune. Watch for arpeggiated sequences that are fast enough to become rhythmically distracting when read with a measured voice.

Deep space ambient, slow analog synth pads, sparse low pulses,
cold and vast, 50 BPM, textural drone, steady dynamics

Cyberpunk street underscore, dark synth bass, granular glitch texture,
tense, 90 BPM, sparse arrangement, no lead melody

Ship interior ambience, low synth hum, faint mechanical texture,
calm and technical, 65 BPM, steady dynamics, no percussion

Fantasy, historical and children’s stories — moderate

These palettes are defined by melodic instruments — harp, flute, lute, glockenspiel, music box — so the melody problem is structural rather than incidental. The fix is to ask for the instrument’s texture rather than its tune: arpeggios, ostinatos and rolled chords give you the color of a harp without a hook competing with the narrator.

Enchanted forest underscore, harp arpeggios, soft flute texture,
ethereal pads, dreamy, 70 BPM, no lead melody, steady dynamics

Dark epic fantasy bed, solo cello, low string drone,
ancient and mysterious, 60 BPM, sparse, minimal percussion

Medieval village ambience, lute ostinato, hand drum, recorder texture,
rustic and warm, 92 BPM, steady dynamics, no lead melody

Victorian drawing room underscore, string quartet, harpsichord figure,
formal and restrained, 70 BPM, sparse, steady dynamics

Bedtime story ambience, soft music box texture, gentle harp,
warm and soothing, 55 BPM, no lead melody, steady dynamics

Talking animals adventure bed, pizzicato strings, light woodwind texture,
bouncy, 100 BPM, steady dynamics, minimal percussion

Noir, jazz and romance — hardest

Jazz is an improvisational idiom, so the model’s instinct is to give you a soloist — and a muted trumpet or tenor sax sits directly in the speaking register, playing an expressive line, at unpredictable volume. Romance has the same problem via the swelling string arrangement. Both are workable if you strip them to the rhythm section and explicitly refuse the solo.

1940s noir underscore, upright bass walk, brushed snare, quiet piano comping,
smoky and cool, 78 BPM, steady dynamics
Exclude: saxophone solo, trumpet solo, lead melody, improvisation, vocals

Rain-soaked city bed, slow jazz piano voicings, soft double bass,
reflective, 70 BPM, sparse, no lead melody, steady dynamics

Slow-burn romance underscore, warm sustained strings, low piano texture,
tender, 62 BPM, no crescendo, steady dynamics, no lead melody

Intro and outro themes — the exception

Themes are the one place where every rule above inverts. Nothing is being spoken over them, so you want a melody, a dynamic arc, and a motif memorable enough that listeners recognise episode two. Drop the sparse-and-steady language entirely.

Audiobook series opening theme, orchestral swell, memorable string motif,
cinematic and bold, 90 BPM, full dynamic range, strong melodic hook

Podcast intro theme, warm acoustic guitar, light percussion,
inviting and confident, 100 BPM, clear melody, builds to a peak

Chapter outro, soft piano and strings, gentle resolve,
calm sign-off, 66 BPM, fades to silence

Troubleshooting

There are still vocals in the track

Check the Instrumental toggle is actually on — describing the track as instrumental in Styles does not set it. If the toggle is on and you’re hearing wordless voice, that’s choir or vocal pad being generated as instrumentation; add those terms to Exclude specifically rather than relying on the word “vocals.” If a generation is otherwise perfect, extract stems and drop the vocal stem.

The music keeps swelling over the narration

Genre words carry implied dynamics — “epic,” “cinematic,” “trailer” and “building” all instruct the model to write an arc. Remove them from Styles, add crescendo and swell to Exclude, and ask for steady dynamics in their place.

The loop point is audible

You’re most likely looping across a section boundary. Cut your loop from the middle of a settled passage, not from the opening — and cut on a bar line, with a short crossfade over the seam. Sparse, drone-based beds loop far more cleanly than rhythmic ones, which is another argument for the first two palettes above.

The bed is fine alone but muddy under the voice

This is a register collision, not a prompt failure. Regenerate with instruments that sit outside the speaking range, or extract stems and remove the offending part. A gentle high-pass or a narrow cut around the low mids on the music track will also open up room for the voice without dropping the bed’s overall level.

Questions

It depends on your plan. Suno states that music made on the free Basic plan is owned by Suno and licensed for non-commercial use only; songs made while subscribed to Pro or Premier are owned by you with a commercial licence. Separately, Suno notes that fully AI-generated music may not be eligible for copyright protection in jurisdictions such as the US, which is a different question from whether you may monetise it. Check your current plan terms before publishing.

Long enough that the listener doesn’t register the repeat. A 60–90 second loop is generally the shortest that survives a long chapter; anything under 30 seconds becomes obvious within a few minutes. Sparse, non-rhythmic material tolerates shorter loops than anything with a clear beat.

A number gives the model a far tighter target than “slow” or “medium,” but treat it as a strong steer rather than a setting — generations will land in the neighbourhood rather than exactly on it. For narration the useful range is roughly 40–70 BPM, which sits below conversational speech rhythm and so doesn’t compete with the reading pace.

Both, for different jobs. The Instrumental toggle handles lead vocals. Exclude handles everything else you want kept out — wordless choir, melodic hooks, crescendos, specific percussion. Neither is a guaranteed audio filter, so expect to regenerate occasionally rather than assuming a failed generation means you wrote the prompt wrong.

Where to start

Take one prompt from the sleep and meditation group, run it with the Instrumental toggle on and the baseline Exclude list, and mix it under thirty seconds of your own narration before generating anything else. That one test tells you more about your voice, your levels, and your chain than any prompt library will.