Free tools Windows power users keep installed
One-click scans. No signup required.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
If your VRM avatar’s mouth opens too widely during speech, check both the audio-to-mouth mapping and any other active mouth expression. A mapping that hits its maximum on too many frames loses room to show louder and quieter speech differently; meanwhile, expressions such as happy can compound the opening from aa. Measure your audio distribution, tune the mapping against it, and judge the rendered result rather than relying on a universal threshold.
Why a VRM mouth can open too far
There are two common mechanisms to check. First, an amplitude-based lip-sync mapping may send many input frames to its maximum output. Once that ceiling is reached, additional loudness cannot produce additional variation, so the mouth can appear stuck wide. Second, another active expression can add to the mouth opening. The VRM 1.0 specification explicitly warns that combining happy and aa can make the mouth open too much; it describes the result as “The mouth opens too much and becomes strange.” VRM 1.0 expression specification
VRM standardizes expression names and weights, not how an application analyzes audio. Its optional procedural lip-sync presets are aa, ih, ou, ee, and oh; expression weights range from 0 to 1, and implementations should clamp values outside that range. The specification also defines overrideMouth, which lets an application control procedural lip-sync weights while another expression is active. The actual audio analysis and calibration remain application choices. VRM 1.0 expression specification The official feature overview describes standardized a-i-u-e-o facial operations and audio-generated lip sync, but does not prescribe an RMS threshold or tuning algorithm. VRM features overview
Calibrate the mapping from representative audio
Quantiles describe the distribution of input levels over time. The median gives the middle frame level; lower and upper quantiles show quieter and louder portions. For a simple normalized amplitude-to-mouth mapping, a baseline can represent the quiet or noise floor, while a reference level determines where the mouth reaches its chosen maximum. These are calibration parameters, not VRM-wide defaults.
#1 Best Overall
- CARDBOARD MONKENAUT — Get our best Gorilla Tag bundle yet with this Amazon exclusive deal. Purchase Meta Quest 3S to get exclusive items, including the Gorilla Space Program Suit and Helmet, plus 2,000 SHINY ROCKS.
- NO WIRES, MORE FUN — Break free from cords. Game, play and explore immersive worlds — untethered and without limits.
- 2X GRAPHICAL PROCESSING POWER — Enjoy lightning-fast load times and next-gen graphics for smooth gaming powered by the Snapdragon XR2 Gen 2 processor.
- EXPERIENCE VIRTUAL REALITY — Take gaming to a new level and blend virtual objects with your physical space to experience two worlds at once in your VR headset.
- 2+ HOURS OF BATTERY LIFE — Charge less, play longer and stay in the action with an improved battery that keeps up. *Based on the graphic performance of the Qualcomm Snapdragon XR2 Gen 2 platform vs the Meta Quest 2 platform.
- Use representative speech. Include quiet and loud sections from the same microphone, processing chain, and speaking conditions you expect during use. A single average or one loud moment will not show how often the mapping reaches its limit.
- Inspect the level distribution. If your lip-sync system uses frame RMS, record a median and useful lower and upper quantiles from that representative audio. Treat the values as descriptions of your signal, not thresholds that should be copied to another setup.
- Set a baseline and reference. Use the baseline to avoid opening the mouth on the quiet floor, and the reference to set the input level that reaches your intended mouth-opening ceiling. Map the useful part of the input range into the expression range rather than letting ordinary speech sit at maximum.
- Measure ceiling hits. Track the fraction of frames whose mapped expression reaches its cap. If that fraction is high, lower the mapping gain or revise the baseline or reference, then inspect the result again.
- Check quiet speech and the rendered avatar. Reducing gain can restore headroom but also make quiet speech hard to see. Confirm that the mouth still moves at low levels and that the rendered expression looks natural with other active expressions.
One published TTS example by orca_forge reported median frame RMS of 0.214, 25th-percentile frame RMS of 0.024, and 90th-percentile frame RMS of 0.403. In that article’s on-device tuning example, a 0.15 volume baseline produced 58.5% saturated frames. The search result for the article did not establish its publication year. Those are results from that particular example—not standards, targets, or expected values for other voices, microphones, avatars, or pipelines. orca_forge’s RMS tuning example
No universally correct saturation rate is established by the cited material. Use ceiling-hit frequency as a diagnostic: frequent hits suggest lost output headroom, but the acceptable balance depends on the audio and the look you want. Pair the measurement with visual checks, especially at quiet levels.
Rank #2
- NO WIRES, MORE FUN — Break free from cords. Game, play, exercise and explore immersive worlds — untethered and without limits.
- 2X GRAPHICAL PROCESSING POWER — Enjoy lightning-fast load times and next-gen graphics for smooth gaming powered by the SnapdragonTM XR2 Gen 2 processor.
- EXPERIENCE VIRTUAL REALITY — Take gaming to a new level and blend virtual objects with your physical space to experience two worlds at once.
- 2+ HOURS OF BATTERY LIFE — Charge less, play longer and stay in the action with an improved battery that keeps up.
- 33% MORE MEMORY — Elevate your play with 8GB of RAM. Upgraded memory delivers a next-level experience fueled by sharper graphics and more responsive performance.
Check expression overlap after tuning
If the audio mapping is no longer saturating excessively but the mouth still looks too wide, inspect which expressions are active together. In particular, test the procedural lip-sync expression without the emotion expression, then with it, to see whether their combined effect is responsible. Where the application supports it, overrideMouth can control procedural lip-sync weights while another expression is active; how to configure that behavior depends on the application.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteDo not assume that every control labeled “mouth” responds to speech. VRMViewMeister 2.18.0 documents opening and closing speeds for its auxiliary mouth motion and explicitly says that feature does not make the model move its mouth according to user speech. Those controls are not a substitute for calibrating an audio-reactive mapping. VRMViewMeister manual
Rank #3
- CARDBOARD MONKENAUT — Get our best Gorilla Tag bundle yet with this Amazon exclusive deal. Purchase Meta Quest 3 to get exclusive items, including the Gorilla Space Program Suit and Helmet, plus 2,000 SHINY ROCKS.
- NEARLY 30% LEAP IN RESOLUTION — Experience every thrill in breathtaking detail with sharp graphics and stunning 4K+ Infinite Display.
- NO WIRES, MORE FUN — Break free from cords. Game, play and explore in immersive worlds — untethered and without limits.
- 2X GRAPHICAL PROCESSING POWER — Enjoy lightning-fast load times and next-gen graphics for smooth gaming powered by the Snapdragon XR2 Gen 2 processor.
- EXPERIENCE VIRTUAL REALITY — Blend virtual objects with your physical space and experience two worlds at once in your VR headset.
Know when RMS is the wrong tool
RMS measures signal energy; it does not identify the phoneme being spoken. An amplitude-driven mapping can make the mouth open and close with loudness, but it cannot reliably reproduce lip closures and other articulation that depend on speech sounds. If phoneme-accurate movement matters, use a viseme- or phoneme-aware approach rather than expecting RMS tuning to supply those gestures.
The three-vrm-lip-sync project documents runtime controls for gain, smoothness, minimum and maximum volume, and viseme gain. These controls are specific to that project and are not a comparative benchmark or a universal VRM feature. Check the current documentation for the library and application you use before relying on a particular control or API. three-vrm-lip-sync README
Quick Recap
Best Value
- NEARLY 30% LEAP IN RESOLUTION — Experience every thrill in breathtaking detail with sharp graphics and stunning 4K Infinite Display.
- NO WIRES, MORE FUN — Break free from cords. Play, explore and exercise in immersive worlds — untethered and without limits.
- 2X GRAPHICAL PROCESSING POWER — Enjoy lightning-fast load times and next-gen graphics for smooth gaming powered by the Snapdragon XR2 Gen 2 processor.
- EXPERIENCE VIRTUAL REALITY — Blend virtual objects with your physical space and experience two worlds at once.
- 2+ HOURS OF BATTERY LIFE — Charge less, play longer and stay in the action with an improved battery that keeps up.
Rank #4
- Your purchase of this item includes a new Meta Quest Pro 256 GB VR headset and a 12-month subscription to Optima Academy Online (OAO) field trips.
- Optima Academy Online (OAO) harnesses the power of virtual reality to make previously impossible learning opportunities just a few clicks away. Our VR Field Trips provide powerful ways of engaging users on a whole new level while providing learning experiences. With our VR Field Trips, we deliver users directly into an immersive educational experience that engages them like never before. We offer a one-month subscription to our VR Field Trips. During your subscription, you can spend as much time in our uniquely created Metaverse environments as you like. Each environment has its own theme, learning experiences, and adventures.
- High resolution mixed reality passthrough uses full-color sensors to let you see and engage with the physical world around you, even as you connect, work and play in virtual spaces.
- Share your true emotions and reactions with real time natural avatar expressions. Meta Avatars translate your natural facial expressions into VR so you can bring your true personality to meetings and gatherings with friends.
- Meta Quest Touch Pro Controllers translate instinctive hand gestures and detailed finger actions directly into VR with self-tracking cameras and precision controls. Multi-point, advanced haptics make virtual interactions feel entirely real
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

