A single looping music track stops being enough the moment the player's situation changes. The music that fits a quiet exploration scene does not fit a combat encounter, and a hard cut between the two sounds like a mistake. Vertical layering solves this by stacking multiple synchronized music stems and adjusting their volumes in real time based on game state, so the score shifts intensity without ever stopping.
Wwise handles this better than most tools, because the Interactive Music Hierarchy was built specifically for this kind of structure. The hierarchy separates musical content from game logic, which means a composer can write the stems and a programmer can wire them to parameters without either needing to understand the other's domain. The result is a system where the music responds to gameplay continuously rather than switching between fixed states.
The Wwise first-week guide covers the basics of getting a project started with the middleware. This tutorial assumes that foundation and walks through the specific steps for building a working vertical layering system from scratch: preparing the audio stems, setting up the Interactive Music Hierarchy, mapping RTPCs to track volumes, adding stingers for critical moments, and implementing the whole thing in Unity. The game audio production pipeline guide covers where music production fits in a larger project; this article is about the implementation layer.
The Audio Preparation That Determines Whether It Works
Vertical layering only works if every stem shares the same tempo, key center, and bar-aligned structure. This is not a stylistic preference; it is a hard technical requirement. If one layer loops at 32 bars and another at 16 bars, they will drift out of phase within a few minutes and the player will hear an unintentional polyrhythm.
The practical requirement is that all stems are exported from a single DAW session with a locked tempo map, starting at bar one. Each layer is exported as its own file at identical length, even if a given layer is mostly silence for long stretches. Silence that is correctly aligned is far more useful to Wwise than a shorter file that is musically cleaner but does not line up.
A practical stem structure for a vertical layering system with three intensity levels looks like this:
- Layer 1 — Ambient Pad. Always playing. Establishes the emotional baseline. Low volume, wide stereo, no rhythm.
- Layer 2 — Melody and Soft Percussion. Fades in when the player begins engaging with the game world. Adds motion without demanding attention.
- Layer 3 — Full Rhythm and Brass. Fades in during combat or high-intensity moments. The loudest, most rhythmically active layer.
The total file size of a three-stem system that loops for 64 bars at 120 BPM, stereo, 44.1 kHz, is roughly 13 MB per layer as uncompressed PCM, or about 40 MB for the full set. Compressed with Vorbis at 192 kbps, the same set drops to about 4 MB, which is manageable even for a mobile project. The file size reduction guide covers the compression choices that apply here; music compression is less aggressive than sound effect compression because the ear is more sensitive to artifacts in sustained tonal content.
Setting Up the Interactive Music Hierarchy
The Interactive Music Hierarchy is a separate tree from the Actor-Mixer Hierarchy. You switch to it by clicking the Interactive Music Layout in the Wwise toolbar, or by using the Layout menu. The hierarchy uses the same tree structure as the rest of Wwise, but the objects inside it are music-specific.
The three objects you will use for vertical layering:
Music Segment is the basic building block. It contains one or more Music Tracks, each with audio clips. For vertical layering, you create one Music Segment that contains all the stems as separate tracks. The segment defines the musical timeline: tempo, meter, and loop points.
Music Track holds the audio for one layer. In a three-layer system, the segment has three tracks. Each track has its own volume, pitch, and high/low-pass properties that can be adjusted independently.
Music Playlist Container organizes multiple segments and handles playback order. For a simple vertical layering system, the Playlist Container holds a single segment on loop. For a more complex system that combines vertical layering with horizontal re-sequencing, the container holds multiple segments that transition between each other.
The setup sequence in Wwise:
- In the Interactive Music Hierarchy, right-click and select New Child then Music Playlist Container. Name it something like
Music_Gameplay. - Inside the container, create a Music Segment. Name it
Segment_ExploreLoopor whatever describes the musical content. - Import the stems as audio files by dragging them from the file browser into the segment. Each stem becomes a separate Music Track inside the segment.
- Set the segment's tempo and meter to match the DAW session. This is critical for ensuring the loop points align correctly.
- Set the loop points by dragging the yellow loop markers at the top of the segment timeline. For a seamless loop, the loop start and end points should be at bar boundaries.
The Music Segment in Wwise is a container for the musical timeline; the actual audio data lives in the Music Tracks inside it. The segment's properties, including tempo, meter, and loop points, apply to everything inside it, which is why setting them before importing the stems is the right order.
Making the Layers Respond to Gameplay
A Music Segment with three tracks is still static. The layers are all playing at full volume, and nothing changes. The next step is to make the layer volumes respond to a game parameter, which is where RTPCs (Real-Time Parameter Controls) come in.
An RTPC is a named numeric value that the game sends to Wwise. The game does not send commands like "increase the drums"; it sends a number, and Wwise maps that number to a property through a curve. The mapping lives in Wwise, which means the audio designer can tune the response without changing a line of code.
To set up an RTPC for vertical layering:
- In the Game Syncs tab of the Project Explorer, right-click the Game Parameters section and create a new Game Parameter. Name it
Intensity. - In the Project Explorer, select the Music Track you want to control — for example, the percussion layer.
- In the Property Editor, find the Volume property and click the RTPC icon next to it (the small blue circle).
- Select the Intensity RTPC from the dropdown. This creates an RTPC curve for the track volume.
- In the RTPC curve editor that appears below the property, set the curve. For a percussion layer that should be silent at low intensity and loud at high intensity, set the curve to start at -96 dB (inaudible) when Intensity is 0, and ramp up to 0 dB (full volume) when Intensity is 100.
The -96 dB value is Wwise's threshold for inaudibility. Setting the track volume to -96 dB effectively mutes the layer without removing it from the mix, which means it is still playing at a sample level and the transition is instant when the RTPC value rises. This is different from muting the track, which would cause a discontinuity when the track is unmuted.
Repeat the RTPC setup for each layer, with different curves:
- Ambient Pad: Volume stays at 0 dB regardless of Intensity. This layer is always audible.
- Melody Layer: Ramps from -96 dB to 0 dB as Intensity goes from 20 to 60.
- Percussion/Brass Layer: Ramps from -96 dB to 0 dB as Intensity goes from 60 to 100.
The overlapping ranges are deliberate. At Intensity 30, the melody is partially audible and the percussion is silent. At Intensity 70, the melody is at full volume and the percussion is partially audible. The layers crossfade against each other, which is what makes the transition feel smooth rather than stepped. The step between layers should overlap by at least 15 to 20 RTPC units to avoid an audible gap where neither layer is at a useful volume.
Wwise also supports Switch changes as an alternative to RTPCs for discrete state transitions. The difference is that an RTPC is continuous and a Switch is discrete. For a vertical layering system where the intensity changes gradually with gameplay, RTPCs are the right tool. For a system where the music changes between two clearly defined states — exploration and combat, for example — a Switch might be simpler to manage. The Interactive Music documentation describes both approaches; the choice depends on whether the game's intensity value is continuous or discrete.
Adding Stingers for Critical Moments
Stingers are short musical phrases that play over the currently playing music without interrupting it. A stinger fires on a specific trigger — finding a treasure, landing a critical hit, discovering a new area — and it plays on top of whatever layer configuration is currently active.
In Wwise, a stinger is a Music Segment associated with a Trigger. The Trigger is what the game calls; Wwise handles the rest.
- Create a Music Segment for the stinger. Keep it short — typically two to four bars.
- In the Interactive Music Hierarchy, select the Music Playlist Container that holds your main music.
- In the Property Editor, go to the Stingers tab.
- Drag the stinger Music Segment from the Project Explorer into the Stingers list. Wwise automatically creates a Trigger and a stinger association.
- Set the Play At property to determine when the stinger fires. The options are Immediate, Next Grid, Next Bar, or Next Marker. For most game stingers, Next Bar or Next Grid is the right choice, because it ensures the stinger lands on a musically sensible beat.
The Play At setting matters more than it seems. A stinger that fires immediately can land in the middle of a measure, which sounds jarring. A stinger that fires on the next bar lands on a downbeat, which sounds intentional. The tradeoff is that the player waits up to a few seconds for the stinger to play, which can weaken the feedback if the event is time-sensitive.
The practical resolution is to use Immediate for stingers that need to feel responsive and Next Bar for stingers that are part of the musical texture. A hit-landing stinger that fires on the next bar might feel disconnected from the action; a discovery stinger that fires on the next bar feels like part of the score.
Generating the SoundBank and Integrating with Unity
The Wwise project is an authoring environment. The game engine does not read it directly. The audio has to be packaged into a SoundBank, which is a binary file containing the audio data and the logic that the game needs at runtime.
To generate the SoundBank, select the Events tab, find the Event that plays the music, and drag it into a SoundBank in the SoundBanks layout. Then click Generate SoundBanks in the SoundBanks toolbar. The generated SoundBank appears in the project's Generated SoundBanks folder.
The Event for the music is what the game calls. It is typically created in the Events tab and named something like Play_Music_Gameplay. The Event should be set to play the Music Playlist Container, not the individual Music Segment. The Playlist Container handles the looping and the segment selection.
For the Unity integration:
- In the Unity project, create an empty GameObject to hold the music system. Name it something like
MusicManager. - Add an AkBank component to the GameObject. Set the SoundBank to the one generated from Wwise.
- Add an AkEvent component to the same GameObject. Set the Event to the music event created in Wwise.
- Create a C# script that gets a reference to the AkEvent component and calls
AkSoundEngine.PostEvent()to start the music. The Event name is the string identifier used in Wwise. - To change the intensity, use
AkSoundEngine.SetRTPCValue("Intensity", value)where value is a float from 0 to 100.
The AkBank and AkEvent components are what the Unity integration provides; they are the bridge between the Unity scene and the Wwise runtime. The integration takes about three to five days for an experienced programmer to set up for a typical project, though most of that time is spent on the sound effects and dialogue system rather than the music.
One implementation detail that trips people up: the RTPC value should be set in the Update() method or on a timer, not in a one-shot call. The game's intensity value changes over time, and the RTPC needs to be updated continuously to track it. A common pattern is to lerp the RTPC value toward the target value over a fraction of a second, which smooths out abrupt changes and prevents the music from jumping between intensity levels. A lerp time of 0.5 to 1.5 seconds works well for most games; the right value depends on how quickly the gameplay intensity actually changes. The Adaptive Sound System guide covers the broader implementation patterns for runtime parameter control.
Where Vertical Layering Falls Short
The technique is not the right answer for every game, and knowing the failure cases prevents a lot of wasted production time.
The most common failure is a game without a continuous intensity value. A turn-based RPG, a puzzle game, a visual novel — these do not have a metric that ramps smoothly between states. Combat begins and ends abruptly, and there is no natural continuum for the RTPC to ride. A vertical layering system in that context produces either a static track (because the intensity value never changes) or a jarring one (because the value jumps between discrete points). The right approach for those games is a Switch-based system or a set of separate music tracks that crossfade.
The second failure case is sparse or ambient music. Vertical layering works because the layers reinforce each other — the pad provides continuity, the melody adds motion, the percussion adds energy. If the music is minimal by design, the layers do not have enough content to reinforce each other, and the system produces what sounds like a single track with slight volume fluctuations. A sparse ambient score is better served by horizontal re-sequencing or by a single looped piece that does not try to respond to gameplay.
The third case is a project without the production budget to write multiple layers. Vertical layering multiplies the composer's work: a three-layer system requires three times the musical content. For a small indie project, that multiplication may not be justified. A single looping track with a well-placed stinger often serves the same purpose at a fraction of the cost.
The fourth case is a composer who writes linearly rather than in loops. Vertical layering requires the music to be conceived as a set of synchronized layers from the start. A composer who writes a single piece and then tries to split it into layers afterward almost always produces a system that sounds wrong — the layers do not sit correctly because they were not written to be heard independently. The composer has to be involved in the layering plan from the beginning, not brought in to cut up an existing track.
If the game does not have a continuous gameplay metric, or the music is intentionally sparse, or the budget does not support multiple stem recordings, the vertical layering tutorial above is the wrong place to start. The Wwise first-week guide covers the simpler patterns that are appropriate for those projects, and the broader comparison of audio middleware covers the tools that handle each case best.
When the technique is the right fit, though, it produces a score that responds to the game in a way that no single looping track can match. The player does not hear the layers — they hear music that fits the moment, which is the entire point.
Generate Sound Effects to Complement Your Adaptive Music
Open the SfxMaker generator and create the sound effects that sit alongside the music system — combat impacts, UI feedback, and environmental sounds. These are generated from parameters rather than recordings, so they import into Wwise the same way any WAV file does.
Open SfxMaker Generator →