The previous two articles in this series treated sound and haptics as separate signals that need to be aligned. The synchronization guide covered the timing relationship: how many milliseconds of offset the perceptual system tolerates, and how to compensate for the different latencies of audio drivers and haptic actuators. The controller haptic design guide covered the signal itself: the DualSense actuator's frequency response, DAW-based authoring, and the filter settings that produce a strong tactile sensation.
Both of those articles assume the sound and the haptic are designed independently and then brought together. That assumption is useful for production, but it misses the question that determines whether the combination actually works: does the brain treat the two signals as a single event, or as two events that happen to coincide?
Cross-modal design is the practice of answering that question deliberately. It is not about making the haptic signal match the audio waveform. It is about making the two signals congruent in the dimensions the brain actually uses to fuse them — timing, semantic content, and intensity scaling — while accepting that the two channels will always be physically different.
The Brain Does Not Receive Two Signals. It Receives One Event.
The starting point for cross-modal design is understanding what happens when a sound and a vibration arrive at the same time from the same source. The brain does not process them as separate streams and then decide they are related. It fuses them into a single perceptual event, using a process that has been studied extensively under the label of multisensory integration.
The dominant theoretical framework for this fusion is Bayesian integration, sometimes called the maximum likelihood estimate model. The brain treats each sensory channel as a noisy estimate of some underlying property — the location of a source, the hardness of an object, the intensity of an impact — and combines the estimates by weighting each one according to its reliability. A channel that is more precise contributes more to the final percept. A channel that is less precise contributes less.
This has a specific and testable consequence for game design: the perceived property of the combined event is not the same as either channel alone. If the audio suggests a soft impact and the haptic suggests a hard impact, the brain does not experience a contradiction. It integrates them, and the result depends on the relative reliability of the two channels for that particular judgment.
For hardness judgments, research has shown that auditory cues can influence the perceived hardness of a contact that also has a haptic component. The audio and haptic channels are not redundant; they inform each other.
The practical implication is that a haptic signal does not need to be a faithful reproduction of the audio waveform. It needs to be a reliable estimate of the same underlying event, in a form the actuator can deliver.
Semantic Congruence Beats Signal Fidelity
The most common mistake in cross-modal design is trying to make the haptic signal as similar as possible to the audio signal. A sound designer who has designed a sharp, bright weapon impact will try to reproduce that brightness in the haptic actuator. The result is usually a weak, buzzy vibration that feels disconnected from the sound.
The reason is that the two channels have different perceptual vocabularies. The audio channel carries information primarily through frequency content: a high frequency reads as bright, a low frequency reads as heavy, a fast attack reads as sharp, a slow attack reads as soft. The haptic channel carries information primarily through amplitude envelope and temporal pattern. The actuator can reproduce timing with high precision, but its frequency response is narrow and non-linear.
Research on audio-tactile integration has found that the critical factor for a convincing combined percept is semantic harmony between the channels, not physical similarity. A study on phantom sensations found that penetration-related sounds significantly enhanced perceived realism and satisfaction when combined with vibrotactile feedback, and the authors highlighted the role of semantic harmony between auditory and tactile information[reference:0]. The sounds and vibrations “agreed” on what the event was, even though they were physically different signals.
The design principle that follows is straightforward: the haptic signal should be designed to communicate the same event as the audio signal, using the parameters the haptic channel controls well. The event is “a metal impact.” The audio communicates the metal through bright mid-range content and a ringing decay. The haptic communicates the metal through a sharp attack and a short, crisp decay. Neither channel reproduces the other. Both describe the same thing.
The Three Dimensions That Determine Fusion
Apple's audio-haptic design guidance from WWDC 2019 identifies three principles that determine whether a sound and a haptic will be perceived as a unified event: causality, harmony, and utility[reference:1]. These are not abstract design values. They correspond to specific parameters that can be adjusted.
Causality is about attribution. The player should be able to identify what caused the feedback. A weapon recoil haptic that fires at the exact moment the weapon fires, with the sound of the shot, has clear causality. A haptic that fires during a menu transition, with no visible cause, has none.
The design check is simple: can the player point to the moment in the game that produced the sensation? If the answer is no, the haptic is noise, regardless of how well it is designed in isolation.
Harmony is about congruence between the channels. Apple's phrasing is that “things should feel the way they look, the way they sound”[reference:2]. For audio-haptic design, this means the intensity, sharpness, and duration of the haptic should match the intensity, brightness, and decay of the sound.
The mapping is not one-to-one, because the channels have different ranges. A loud sound does not translate to a strong haptic by copying the amplitude. It translates by scaling the haptic intensity to match the perceived loudness. A bright sound does not translate by copying the high-frequency content. It translates by increasing the haptic sharpness — the rate of onset and the crispness of the decay.
Utility is about information. A haptic that repeats information the player already has from audio and visuals is redundant. A haptic that adds information the other channels do not carry is useful. The events that benefit most from haptics are the ones where the tactile channel contributes something the player would otherwise have to infer: the weight of a landing, the recoil of a weapon, the direction of an impact.
Mapping Audio Parameters to Haptic Parameters
The practical work of cross-modal design is deciding which audio parameter maps to which haptic parameter. The mappings that have emerged from practice and from the Wwise Motion workflow are not arbitrary. They follow from what each channel measures well.
Audio amplitude → haptic intensity. The loudness of the sound maps to the strength of the vibration. This is the most direct mapping and the one that needs the least adjustment. A loud explosion produces a strong haptic; a quiet footstep produces a subtle one. The absolute levels are different, but the relative scaling should be consistent across the sound set.
Audio attack time → haptic attack time. A sharp transient in the audio — a click, a crack, a snap — maps to a fast onset in the haptic. A soft, rounded sound maps to a gradual onset. This mapping is what makes the combined event feel like a single impact rather than a sound followed by a buzz.
Audio decay time → haptic decay time. A short sound maps to a short haptic pulse. A long, reverberant sound maps to a sustained or repeating vibration. This is where the actuator's limitations matter most: a sound with a two-second decay cannot be faithfully reproduced by a haptic actuator that has a meaningful response time of tens of milliseconds. The haptic version of a long decay is usually a shortened version that communicates the character without reproducing the full length.
Audio frequency content → haptic sharpness. This is the least direct mapping, and the one where copying the audio waveform fails most obviously. A bright sound does not map to a high-frequency haptic, because the actuator cannot reproduce high frequencies. It maps to a haptic with a fast attack and a crisp cutoff — what the haptic vocabulary calls sharpness. A dull sound maps to a soft, rounded haptic.
The Wwise Motion workflow supports these mappings directly. A Motion Source plugin can generate a dedicated haptic signal from a synthesis source, independent of the audio signal, while still being triggered by the same event[reference:3]. The designer can route a weapon sound to a Motion Bus with a low-pass filter that keeps only the frequencies the actuator can reproduce, then add a synthesized low-frequency layer to fill the actuator's sweet spot[reference:4]. The result is a haptic that shares the rhythm and envelope of the sound while having its own spectral content designed for the actuator.
What Cross-Modal Design Changes About Testing
The testing process for a cross-modal effect is different from testing a sound or a haptic in isolation. The relevant question is not “does the sound sound good” or “does the haptic feel good.” It is “does the combination feel like one event.”
The subjective test is the fusion test. Play the sound and the haptic together. Does the player experience one sensation or two? If two, the most likely causes are timing misalignment (covered in the synchronization guide), intensity mismatch, or semantic incongruence — the sound and the haptic describing different events.
The diagnostic test is the channel isolation test. Play the sound alone. Play the haptic alone. Then play them together. If the combination feels like the sound plus a vibration that happens to occur at the same time, the channels are not fusing. If the combination feels like a single more substantial event that neither channel produces alone, the fusion is working.
The repetition test applies here as it does everywhere in game audio. A cross-modal effect that is convincing once can become fatiguing after repeated exposure. The haptic component is often the first thing to become irritating, because the tactile channel has no equivalent of “turning the volume down” — the player either feels the vibration or they do not.
The variation test is the most important one for cross-modal design. If the audio has variation — small pitch shifts, different samples — the haptic needs corresponding variation, or the two channels will drift apart over repeated playbacks. A footstep that sounds slightly different each time but vibrates identically every time breaks the fusion. The variation in the haptic does not need to be large, but it needs to be present.
When to Fuse and When to Separate
Cross-modal fusion is not always the goal. Some sounds and haptics should be designed to remain perceptually distinct, and the designer should know the difference.
Fusion is the right target for events that are physically unified: a weapon firing, a footstep landing, a collision, a hit. These are single physical events with both an acoustic and a tactile component, and the brain expects them to be integrated.
Separation is the right target for events that are different in kind. A music track and a controller rumble are not the same event, and trying to fuse them produces a confused experience. A UI confirmation sound and a haptic pulse that fires at the same time are two different communications — “the action succeeded” and “something happened” — and forcing them into a single percept removes the redundancy that makes the feedback reliable.
The most useful heuristic is to ask whether the sound and the haptic share a physical cause. A gunshot sound and a recoil vibration share the cause of the gun firing. A coin sound and a vibration share the cause of the coin being collected. A background ambience and a controller hum do not share a cause; they are two unrelated signals that happen to be playing at the same time.
The psychology of game audio covers the broader principle: the player's perceptual system is constantly looking for causal relationships between events. When the sound and the haptic share a cause, the system integrates them automatically. When they do not, the system keeps them separate, and the result is a busier, less coherent experience.
A Working Order for Cross-Modal Design
The design sequence that produces the most coherent result starts with the event, not the signals.
- Identify the physical event. What happened in the game world? A weapon fired, a door opened, a character landed. The event is the common referent that both channels are describing.
- Decide which channels carry the event. Not every event needs audio and haptics. A UI click might need audio only. A subtle environmental change might need neither. The events that benefit from both are the ones where the physicality of the event is part of what the player is supposed to feel.
- Design the audio first. The audio channel carries more information and is easier to iterate on. The haptic will be designed to agree with the audio, not the other way around.
- Extract the envelope, not the waveform. The haptic signal should share the attack time, the decay time, and the relative intensity of the audio signal. It should not share the frequency content, because the actuator cannot reproduce it.
- Synthesize the haptic carrier. Use a frequency in the actuator's sweet spot — roughly 80 to 250 Hz for a DualSense actuator — and shape it with the extracted envelope. The result is a haptic that has the rhythm of the sound but the spectral content the actuator can actually deliver.
- Test the fusion, not the channels. Evaluate the combination as a single percept. If it feels like two events, adjust the timing first, then the intensity scaling, then the semantic content.
- Add variation to both channels. The audio and the haptic should vary together. A footstep that varies in pitch should have a haptic that varies in intensity or sharpness by a corresponding amount.
The final step is the one that most projects skip. Variation in the haptic is easy to forget because the haptic channel is less consciously noticed. But it is precisely because the haptic is less conscious that the mismatch becomes noticeable over time. The player may not be able to say why the footsteps feel wrong after ten minutes, but the identical vibration accompanying slightly different sounds is the cause.
Create a Sound with a Clear Envelope for Haptic Pairing
Open the SfxMaker generator and create a short impact sound with a distinct attack and a clean decay. Sounds with a simple, readable envelope are the easiest to translate to a haptic signal, because the envelope is what the haptic channel reproduces most accurately.
Open SfxMaker Generator →Common Mistakes
- Trying to make the haptic signal match the audio waveform. The actuator cannot reproduce high frequencies or complex spectra. The haptic should share the envelope, not the waveform.
- Designing the sound and haptic in isolation. The two channels describe the same event. Designing them separately and hoping they combine is less effective than designing them with the shared event in mind.
- Fusing events that do not share a cause. A UI sound and a haptic pulse are two different communications. Forcing them into a single percept removes information rather than adding it.
- Forgetting variation in the haptic. If the audio varies and the haptic does not, the two channels drift apart over repeated playbacks. The mismatch is subtle but accumulates.
- Judging the effect by the channels in isolation. A sound that is good alone and a haptic that is good alone can produce a combined experience that fails. The fusion is the unit of evaluation.
- Ignoring the accessibility case. A cross-modal effect that only works when both channels are present fails for players who have disabled one of them. Each channel should carry the event on its own, even if the combination is stronger.
What to Check Before Shipping
Play the game with both audio and haptics enabled. Does the combination feel like a single experience, or does the haptic feel like an addition that was applied on top of the audio? The distinction is subtle but consistent: a fused effect feels like the event has weight and presence that the audio alone does not provide. An unfused effect feels like the game is vibrating in sync with the sound.
Play with haptics disabled. Does the game still communicate everything it needs to? The audio and visual channels should be sufficient on their own. The haptic supplements; it does not carry information that exists nowhere else.
Play a long session with the same effects repeating. Do the cross-modal effects remain comfortable, or does the haptic component become the first thing the player wants to turn off? The haptic channel has a lower tolerance for repetition than the audio channel, because there is no equivalent of turning the volume down.
The most successful cross-modal effects are the ones the player never notices as cross-modal. They experience the event as more substantial and more physical than it would be with audio alone, without ever thinking about the vibration as a separate thing. The goal is not to make the haptic noticeable. The goal is to make the event believable.