Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To make a VRM avatar’s mouth move with audio using only RMS, measure short-window audio amplitude, map it to an opening weight, smooth that weight, and apply it to the VRM aa expression. This produces a simple speech-driven open-and-close motion—not phoneme-accurate lip sync. Natural results depend on calibrating the range to your audio, tuning the response and smoothing, and managing competing expressions and playback state.

What RMS-driven aa lip sync can—and cannot—do

RMS (root mean square) describes the strength of waveform samples over a window. Driving one mouth shape with that value is a compact way to make a mouth move during speech and close during silence. It does not identify vowels or infer the timing of individual phonemes. An “i” sound, for example, still drives the same aa shape as other sounds. The implementation approach and its limits are described in the RMS lip-sync implementation article.

VRM 1.0 defines five procedural lip-sync expression keys: aa, ih, ou, ee, and oh. Those keys are standardized, but the avatar’s configured shapes determine how each expression looks. See the VRM 1.0 specification and UniVRM blend-shape documentation.

With aa alone, think of the result as amplitude-driven mouth opening, not accurate lip sync. RMS cannot reliably recover phoneme timing or closures such as the lips coming together before “m.” The implementation article also identifies “n,” geminate “tsu,” and devoiced vowels as cases amplitude alone cannot reliably express. If those articulations matter, use a separate articulation estimator or viseme timing derived from text or audio.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Meta Quest 3S 128GB | Virtual Reality — VR Headset — Gorilla Tag Bundle
  • CARDBOARD MONKENAUT — Get our best Gorilla Tag bundle yet with this Amazon exclusive deal. Purchase Meta Quest 3S to get exclusive items, including the Gorilla Space Program Suit and Helmet, plus 2,000 SHINY ROCKS.
  • NO WIRES, MORE FUN — Break free from cords. Game, play and explore immersive worlds — untethered and without limits.
  • 2X GRAPHICAL PROCESSING POWER — Enjoy lightning-fast load times and next-gen graphics for smooth gaming powered by the Snapdragon XR2 Gen 2 processor.
  • EXPERIENCE VIRTUAL REALITY — Take gaming to a new level and blend virtual objects with your physical space to experience two worlds at once in your VR headset.
  • 2+ HOURS OF BATTERY LIFE — Charge less, play longer and stay in the action with an improved battery that keeps up. *Based on the graphic performance of the Qualcomm Snapdragon XR2 Gen 2 platform vs the Meta Quest 2 platform.

Minimal implementation pattern

The core loop is: analyze the audio being played, compute RMS, map it to a bounded opening weight, smooth the result, then assign it to aa. The pseudocode below shows the calculation and state update; connect rms to the waveform data from your audio-analysis path and adapt expression access to the VRM runtime and version in your project.

function getRms(samples) {
  let sumSquares = 0;
  for (const sample of samples) {
    sumSquares += sample * sample;
  }
  return Math.sqrt(sumSquares / samples.length);
}

function getTarget(rms, floor, reference) {
  const level = Math.max(0, Math.min(1,
    (rms - floor) / (reference - floor)
  ));
  return Math.sqrt(level); // Optional response curve
}

// Each update:
const target = isPlaying ? getTarget(rms, floor, reference) : 0;
opening += (target - opening) * follow;
vrm.expressionManager.setValue('aa', opening);

RMS squares each sample before averaging, so positive and negative waveform values do not cancel as they would in a plain arithmetic mean. The mapping clamps the normalized signal to 0–1: values at or below floor close the mouth, and values at or above reference reach the maximum target. The square-root curve shown is optional; use a linear level instead if you do not want to raise weaker values.

Rank #2
Meta Quest 3S 128GB | Virtual Reality — VR Headset (Renewed Premium)
  • NO WIRES, MORE FUN — Break free from cords. Game, play, exercise and explore immersive worlds — untethered and without limits.
  • 2X GRAPHICAL PROCESSING POWER — Enjoy lightning-fast load times and next-gen graphics for smooth gaming powered by the SnapdragonTM XR2 Gen 2 processor.
  • EXPERIENCE VIRTUAL REALITY — Take gaming to a new level and blend virtual objects with your physical space to experience two worlds at once.
  • 2+ HOURS OF BATTERY LIFE — Charge less, play longer and stay in the action with an improved battery that keeps up.
  • 33% MORE MEMORY — Elevate your play with 8GB of RAM. Upgraded memory delivers a next-level experience fueled by sharper graphics and more responsive performance.

Connect analysis to the audio that is playing

Read waveform samples from the audio analysis path during updates. The implementation article describes branching an analysis path from playback without adding a second connection to the audio destination, because its setup relies on the existing <audio> playback path as the acoustic echo cancellation reference. That is a constraint of the described playback/AEC setup, not a general Web Audio rule. Follow the routing requirements of your own playback and capture design.

Update expressions and close cleanly

Choose an intentional order for expression updates and the regular VRM update in your runtime. If another subsystem can write ih, ou, ee, or oh, clear or coordinate those weights so they do not leave stale mouth poses alongside aa. Set the target to zero as soon as playback becomes inactive; otherwise a stopped render loop or missed update can leave the last nonzero expression visible. Disconnect and dispose of analysis resources when the audio or avatar is cleaned up.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Meta Quest 3 512GB, VR Without Wires, Gorilla Tag Cardboard Monkenaut Bundle, Amazon Exclusive, 3-Month Trial of Meta Horizon+ Included
  • CARDBOARD MONKENAUT — Get our best Gorilla Tag bundle yet with this Amazon exclusive deal. Purchase Meta Quest 3 to get exclusive items, including the Gorilla Space Program Suit and Helmet, plus 2,000 SHINY ROCKS.
  • NEARLY 30% LEAP IN RESOLUTION — Experience every thrill in breathtaking detail with sharp graphics and stunning 4K+ Infinite Display.
  • NO WIRES, MORE FUN — Break free from cords. Game, play and explore in immersive worlds — untethered and without limits.
  • 2X GRAPHICAL PROCESSING POWER — Enjoy lightning-fast load times and next-gen graphics for smooth gaming powered by the Snapdragon XR2 Gen 2 processor.
  • EXPERIENCE VIRTUAL REALITY — Blend virtual objects with your physical space and experience two worlds at once in your VR headset.

Calibrate the mapping against your audio

There is no universal RMS threshold. A floor and reference that work for one TTS voice, microphone, or playback level may make another source barely move or stay open too often. Choose them using the actual audio and inspect both quiet and loud passages.

  • Set the floor: find a level below which the mouth should be closed. This can suppress low-level noise or silence, but a floor set too high can erase quiet speech.
  • Set the reference: choose the signal level that should produce the intended maximum opening. A reference set too low causes frequent saturation; one set too high can make speech look weak.
  • Check the distribution: inspect representative quiet and loud sections, or RMS percentiles, rather than relying only on an average. Watch the avatar at both ends of the range.
  • Choose a curve deliberately: linear mapping preserves normalized differences. A square-root curve makes low levels more visible while compressing differences at the high end, and may increase the proportion of frames at or near maximum opening.

One implementation author reported a median frame RMS of 0.214, a 25th percentile of 0.024, and a 90th percentile of 0.403 for a particular TTS and on-device tuning setup; the excerpt does not expose the year. In that setup, a local baseline of 0.15 reportedly saturated 58.5% of frames. In a separate curve comparison, linear mapping had an average maximum weight of 0.537, a bottom-quartile weight of 0.375, and 3.5% saturation; square-root mapping had corresponding figures of 0.647, 0.531, and 6.1%. These are the author’s reported observations, not general thresholds or expected results, and the curve comparison did not test the exact aa-only code shown above. See the implementation article.

Rank #4
Meta Quest Pro Headset with Virtual Reality Field Trips 1-Month Subscription
  • Your purchase of this item includes a new Meta Quest Pro 256 GB VR headset and a 12-month subscription to Optima Academy Online (OAO) field trips.
  • Optima Academy Online (OAO) harnesses the power of virtual reality to make previously impossible learning opportunities just a few clicks away. Our VR Field Trips provide powerful ways of engaging users on a whole new level while providing learning experiences. With our VR Field Trips, we deliver users directly into an immersive educational experience that engages them like never before. We offer a one-month subscription to our VR Field Trips. During your subscription, you can spend as much time in our uniquely created Metaverse environments as you like. Each environment has its own theme, learning experiences, and adventures.
  • High resolution mixed reality passthrough uses full-color sensors to let you see and engage with the physical world around you, even as you connect, work and play in virtual spaces.
  • Share your true emotions and reactions with real time natural avatar expressions. Meta Avatars translate your natural facial expressions into VR so you can bring your true personality to meetings and gatherings with friends.
  • Meta Quest Touch Pro Controllers translate instinctive hand gestures and detailed finger actions directly into VR with self-tracking cameras and precision controls. Multi-point, advanced haptics make virtual interactions feel entirely real

Tune smoothing without losing synchronization

The update opening += (target - opening) * follow eases the current opening toward the target. A stronger follow coefficient tracks changes more quickly but can look jittery; a weaker coefficient looks steadier but can lag speech. A fixed per-frame coefficient also behaves differently at different frame rates. For a production animation, consider smoothing based on elapsed time so the response is less dependent on how often the render loop runs. Judge the trade-off while watching the avatar with the actual audio.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Prevent expression conflicts and inspect the avatar

The VRM specification warns that applying aa at the same time as happy can open the mouth too far and look strange. VRM 1.0 includes overrideMouth behavior to block or attenuate procedural lip-sync presets while an emotion is active; use it or otherwise coordinate expression weights when emotions and lip sync overlap. The specification’s guidance is, “Do not lip sync during happy.” See the VRM 1.0 specification.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Meta Quest 3 512GB | Virtual Reality — VR Headset — Renewed Premium
  • NEARLY 30% LEAP IN RESOLUTION — Experience every thrill in breathtaking detail with sharp graphics and stunning 4K Infinite Display.
  • NO WIRES, MORE FUN — Break free from cords. Play, explore and exercise in immersive worlds — untethered and without limits.
  • 2X GRAPHICAL PROCESSING POWER — Enjoy lightning-fast load times and next-gen graphics for smooth gaming powered by the Snapdragon XR2 Gen 2 processor.
  • EXPERIENCE VIRTUAL REALITY — Blend virtual objects with your physical space and experience two worlds at once.
  • 2+ HOURS OF BATTERY LIFE — Charge less, play longer and stay in the action with an improved battery that keeps up.

Also check the avatar’s authored expressions rather than assuming the preset name guarantees a particular deformation. UniVRM documents that blend shapes can be combined into expressions, so the final mouth movement depends on the shapes configured for that model. The UniVRM documentation covers that mapping.

When to use more than RMS and one expression

Approach What it estimates Benefits Limits
RMS driving aa Signal strength mapped to one mouth-opening shape Small implementation; language-independent amplitude response; no separate text-timing track required No vowel identification; weak consonant closure and phoneme timing; requires audio-specific calibration and visual tuning. Source: implementation article.
Multi-viseme software path Multiple vowel visemes estimated from audio The three-vrm-lip-sync README documents MFCC-based vowel classification, writes aa/ih/ou/ee/oh, and describes releasing mouth control while silent. More package and runtime integration; a vowel set does not guarantee accurate consonant articulation; check compatibility with installed versions and avatar shapes. Source: library README.

Choose based on the articulation detail you need, integration effort, available avatar shapes, timing requirements, and runtime behavior. The README documents inputs including audio-file URLs, AudioBuffer, <audio>, microphone, and MediaStream; its example updates the animation mixer, then lip-sync weights, then vrm.update, and demonstrates stop and dispose calls. This is documented usage, not an independent compatibility or performance guarantee. Check the current API and version against your installed stack.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.