Writing “instrumental only, no vocals” in Suno’s Styles field is a suggestion, not a control. The Styles field describes what you want the music to sound like — Suno reads the phrase as a stylistic hint and will still hand you a track with a choir pad in it. The vocals are switched off by the Instrumental toggle and the Exclude field, both of which sit in Custom Mode.
That distinction is why most Suno prompt lists for narration don’t work. Below is the order of operations that kills the vocals at the control, gets the arrangement out of your voice’s way, solves the length problem, and only then picks a palette.
Step 1: Switch off vocals at the control
In Custom Mode, toggle Instrumental. This is the reliable method — it changes what the model generates rather than describing a preference to it. Suno’s own documentation lists it as the alternative to entering lyrics, not as a style descriptor.
The Instrumental toggle removes lead vocals. It does not reliably remove wordless voice — choir pads, breathy “ahh” textures, and vocal chops are treated by the model as instrumentation, and they will still land on top of your narration. For those, use Exclude: click Advanced Options in Custom Mode, and the menu opens with the Exclude field. Enter what you don’t want.
Exclude field — paste this as your narration baseline:
vocals, choir, vocal pad, vocal chops, humming, whistling, spoken word, lead melody
Keep the Exclude list short. It is a steer, not a guaranteed filter — a list of thirty terms dilutes each one. Six to ten is the working range.
Step 2: Get the arrangement out of the voice’s way
An instrumental track is not automatically a narration bed. Most instrumentals are written to be listened to, which means they have a melody, a dynamic arc, and energy in the same frequency range as a speaking voice. All three fight narration.
Register
An adult speaking voice has a fundamental frequency in the 85–255 Hz range, with harmonics that contribute to intelligibility reaching several kilohertz above that. Instruments that live in the same space — solo piano in its middle octaves, acoustic guitar strummed in open position, tenor sax, muted trumpet — will mask consonants. Instruments that sit clearly above or below it, such as a low string drone, sub-bass pulse, celeste, glockenspiel, or high sustained pad, leave the voice a clear channel.
This is why “solo piano, melancholic” is the single most over-recommended and worst-performing narration prompt in circulation.
Melody
A listener cannot follow two melodic lines at once, and narration is a melodic line. Ask for texture instead: drones, sustained pads, ostinatos, arpeggio figures, and slow harmonic movement all lend a scene emotional color without demanding attention. Put lead melody and melodic hook in Exclude and use words like underscore, bed, drone, sustained, and textural in Styles.
Dynamics
Cinematic music is built on contrast — a swell that goes from a whisper to a wall. Under narration, that means either the quiet passages vanish, or the loud ones bury you, and no amount of ducking automation fixes a track whose loudest moment is 20 dB above its quietest. Ask for a flat dynamic profile explicitly: steady dynamics, no swells, no crescendo, consistent level throughout.
The two-field template
Narration work is split across two inputs, not one. Styles describes the music; Exclude protects the voice.
Styles:
[setting or genre], [2-4 named instruments], [mood], [BPM] BPM,
sparse arrangement, steady dynamics, textural underscore
Exclude:
vocals, choir, vocal chops, lead melody, crescendo, cymbal crashes
Name specific instruments rather than families — “solo cello and low string drone” produces a far more predictable result than “orchestral.” Give tempo as a number. And use one mood, not three; contradictory emotional directions get averaged into something characterless.
Step 3: Solve length before you solve style
Suno does not generate loops. It generates songs — with an intro, an arc, and an ending. Current models produce up to around eight minutes in a single generation, which sounds like plenty until you try to score a forty-minute chapter.
Three approaches, in order of how much work they cost you:
- Generate long, then cut a loop. Take a full-length generation into your editor, find a section where the arrangement has settled into steady state, and cut a 60–90 second region at a musically sensible point. Crossfade the seam. This produces the cleanest loops because you are choosing the calmest passage rather than accepting whatever the intro happens to be.
- Use Extend. Open the track’s menu, choose Remix/Edit, then Extend. Drag the marker to set how much of the original to keep and where the extension begins. Quick Extend continues without a new prompt; entering style details lets you steer the continuation. Once you have the desired length, use Get Whole Song to render it as a single file.
- Extract stems. If a generation is right except for one element, choose Get Stems from the Edit menu. Suno splits the track into separate parts you can mute or rebalance — the fix for a bed that’s perfect apart from a snare that keeps poking through your dialogue.
Step 4: Pick a palette
These are ordered by how hard each palette is to keep out of the way of a voice, not by genre. The first group is nearly foolproof. The last group will fight you, and the notes explain why.
Sleep, meditation and ASMR — easiest
Almost no melodic content, very slow harmonic movement, and a natural home in the frequency extremes. This palette is the most forgiving under narration and the best place to start if you’re testing your chain.
Sleep story ambience, slow evolving synth pads, distant soft rain,
deeply calming, 40 BPM, sparse arrangement, steady dynamics, no percussion
Guided meditation underscore, low drone, single sustained singing bowl,
spacious and still, 45 BPM, textural, steady dynamics
Forest night ambience, low string drone, distant wind texture,
warm and hypnotic, 42 BPM, sparse, no percussion, steady dynamics
Tension, horror and thriller — easy
Dread is built from sustained dissonance rather than melody, which is exactly what you want. The trap is the stinger: ask for tension, and Suno will often deliver sudden string stabs that spike the level mid-sentence. Exclude them.
Horror narration underscore, dissonant low strings, sub bass drone,
unsettling and sparse, 40 BPM, steady dynamics, no percussion
Exclude: stingers, string stabs, crescendo, sudden hits, vocals, choir
Slow-burn thriller bed, sparse cello pulses, quiet ticking texture,
restrained and anxious, 70 BPM, steady dynamics, minimal percussion
Haunted interior ambience, detuned music box, tape hiss, long reverb tail,
eerie, 48 BPM, sparse, no lead melody
Sci-fi and futuristic — easy
Synth pads occupy the high and low registers naturally and rarely carry a tune. Watch for arpeggiated sequences that are fast enough to become rhythmically distracting when read with a measured voice.
Deep space ambient, slow analog synth pads, sparse low pulses,
cold and vast, 50 BPM, textural drone, steady dynamics
Cyberpunk street underscore, dark synth bass, granular glitch texture,
tense, 90 BPM, sparse arrangement, no lead melody
Ship interior ambience, low synth hum, faint mechanical texture,
calm and technical, 65 BPM, steady dynamics, no percussion
Fantasy, historical and children’s stories — moderate
These palettes are defined by melodic instruments — harp, flute, lute, glockenspiel, music box — so the melody problem is structural rather than incidental. The fix is to ask for the instrument’s texture rather than its tune: arpeggios, ostinatos and rolled chords give you the color of a harp without a hook competing with the narrator.
Enchanted forest underscore, harp arpeggios, soft flute texture,
ethereal pads, dreamy, 70 BPM, no lead melody, steady dynamics
Dark epic fantasy bed, solo cello, low string drone,
ancient and mysterious, 60 BPM, sparse, minimal percussion
Medieval village ambience, lute ostinato, hand drum, recorder texture,
rustic and warm, 92 BPM, steady dynamics, no lead melody
Victorian drawing room underscore, string quartet, harpsichord figure,
formal and restrained, 70 BPM, sparse, steady dynamics
Bedtime story ambience, soft music box texture, gentle harp,
warm and soothing, 55 BPM, no lead melody, steady dynamics
Talking animals adventure bed, pizzicato strings, light woodwind texture,
bouncy, 100 BPM, steady dynamics, minimal percussion
Noir, jazz and romance — hardest
Jazz is an improvisational idiom, so the model’s instinct is to give you a soloist — and a muted trumpet or tenor sax sits directly in the speaking register, playing an expressive line, at unpredictable volume. Romance has the same problem via the swelling string arrangement. Both are workable if you strip them to the rhythm section and explicitly refuse the solo.
1940s noir underscore, upright bass walk, brushed snare, quiet piano comping,
smoky and cool, 78 BPM, steady dynamics
Exclude: saxophone solo, trumpet solo, lead melody, improvisation, vocals
Rain-soaked city bed, slow jazz piano voicings, soft double bass,
reflective, 70 BPM, sparse, no lead melody, steady dynamics
Slow-burn romance underscore, warm sustained strings, low piano texture,
tender, 62 BPM, no crescendo, steady dynamics, no lead melody
Intro and outro themes — the exception
Themes are the one place where every rule above inverts. Nothing is being spoken over them, so you want a melody, a dynamic arc, and a motif memorable enough that listeners recognise episode two. Drop the sparse-and-steady language entirely.
Audiobook series opening theme, orchestral swell, memorable string motif,
cinematic and bold, 90 BPM, full dynamic range, strong melodic hook
Podcast intro theme, warm acoustic guitar, light percussion,
inviting and confident, 100 BPM, clear melody, builds to a peak
Chapter outro, soft piano and strings, gentle resolve,
calm sign-off, 66 BPM, fades to silence
Troubleshooting
There are still vocals in the track
Check the Instrumental toggle is actually on — describing the track as instrumental in Styles does not set it. If the toggle is on and you’re hearing wordless voice, that’s choir or vocal pad being generated as instrumentation; add those terms to Exclude specifically rather than relying on the word “vocals.” If a generation is otherwise perfect, extract stems and drop the vocal stem.
The music keeps swelling over the narration
Genre words carry implied dynamics — “epic,” “cinematic,” “trailer” and “building” all instruct the model to write an arc. Remove them from Styles, add crescendo and swell to Exclude, and ask for steady dynamics in their place.
The loop point is audible
You’re most likely looping across a section boundary. Cut your loop from the middle of a settled passage, not from the opening — and cut on a bar line, with a short crossfade over the seam. Sparse, drone-based beds loop far more cleanly than rhythmic ones, which is another argument for the first two palettes above.
The bed is fine alone but muddy under the voice
This is a register collision, not a prompt failure. Regenerate with instruments that sit outside the speaking range, or extract stems and remove the offending part. A gentle high-pass or a narrow cut around the low mids on the music track will also open up room for the voice without dropping the bed’s overall level.
Questions
It depends on your plan. Suno states that music made on the free Basic plan is owned by Suno and licensed for non-commercial use only; songs made while subscribed to Pro or Premier are owned by you with a commercial licence. Separately, Suno notes that fully AI-generated music may not be eligible for copyright protection in jurisdictions such as the US, which is a different question from whether you may monetise it. Check your current plan terms before publishing.
Long enough that the listener doesn’t register the repeat. A 60–90 second loop is generally the shortest that survives a long chapter; anything under 30 seconds becomes obvious within a few minutes. Sparse, non-rhythmic material tolerates shorter loops than anything with a clear beat.
A number gives the model a far tighter target than “slow” or “medium,” but treat it as a strong steer rather than a setting — generations will land in the neighbourhood rather than exactly on it. For narration the useful range is roughly 40–70 BPM, which sits below conversational speech rhythm and so doesn’t compete with the reading pace.
Both, for different jobs. The Instrumental toggle handles lead vocals. Exclude handles everything else you want kept out — wordless choir, melodic hooks, crescendos, specific percussion. Neither is a guaranteed audio filter, so expect to regenerate occasionally rather than assuming a failed generation means you wrote the prompt wrong.
Where to start
Take one prompt from the sleep and meditation group, run it with the Instrumental toggle on and the baseline Exclude list, and mix it under thirty seconds of your own narration before generating anything else. That one test tells you more about your voice, your levels, and your chain than any prompt library will.