iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Murmur prepares speech and music while current audio plays, then schedules voice and music locally. When a listener interrupts, it clears queued talk and invalidates outdated preparation so a late result cannot take over the conversation. Its design separates content decisions from playback: a model proposes structured segments and selections, while a local Director manages queues and an AudioEngine controls sound.
How Murmur divides content decisions from playback
Murmur is a TypeScript application running on Node.js, with a separate terminal UI process built with Bun and OpenTUI. Its Brain uses the Claude Agent SDK; task-specific tools submit model outputs for schema validation. The model writes talk segments and selects content, but does not directly control speakers. The local Director coordinates preparation and scheduling, and the AudioEngine owns playback state.
Inference and production speech synthesis use external services. Selected context is sent to the inference service, while text to be spoken is sent to the speech service, so those inputs leave the machine. The implementation account describes the design as checked against revision d6c3619. Read the Murmur implementation account.
Free tools Windows power users keep installed
One-click scans. No signup required.
How prefetching reduces waits—and what it cannot hide
Talk buffer
The Director aims to keep two talk segments prepared. Each queue entry contains text and a speech-synthesis Promise whose work has already started. After one segment is consumed, the Director refills the buffer in the background, with no more than one refill task in flight.
#1 Best Overall
- Critically acclaimed sonic performance praised by top audio engineers and pro audio reviewers
- Proprietary 45 millimeter large aperture drivers with rare earth magnets and copper clad aluminum wire voice coils
- Exceptional clarity throughout an extended frequency range, with deep, accurate bass response
- Circumaural design contours around the ears for excellent sound isolation in loud environments
- 90 degree swiveling earcups for easy, one ear monitoring, and professional grade earpad and headband material delivers more durability and comfort
That distinction matters: a segment being in the queue does not mean its speech clip is ready. If playback reaches the entry before synthesis finishes, it must wait for the Promise. Prefetch overlaps preparation with current playback; it does not guarantee uninterrupted audio.
Music prefetch
Music uses a one-slot prefetch. Search, selection, and source resolution proceed in the background. If the next track is not ready at a planned boundary, Murmur plays another talk segment and checks again at the next boundary.
A deeper talk buffer could cover more variation in preparation time, but it also means more generated content and a greater chance that queued material becomes stale after the listener changes direction. The author says the current depth has not been established as a global optimum.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteHow much extra waiting remains
A useful handoff model is W = max(0, P - R), where R is the time left in the current audio when preparation begins and P is preparation time. This estimates extra waiting beyond that remaining playback time; it ignores playback startup overhead and retries. The article’s example of a 30-second segment and 12-second preparation time is illustrative, not a measured benchmark.
Rank #2
- Neodymium magnets and 40 millimeter drivers for powerful, detailed sound.Specific uses for product : Professional audio system,Home audio system
- Closed ear design provides comfort and outstanding reduction of external noises
- 9.8 foot cord ends in gold plated plug and it is not detachable; 1/4 inch adapter included
- Folds up for storage or travel in provided soft case
- Frequency Response: 10 Hertz to 20 kilohertz
What happens when a listener interrupts
Previously queued talk may no longer fit a new request. Murmur clears that queue and invalidates refill work already in progress. While the reply is generated and synthesized, current audio continues. Once reply audio is ready, remaining old voice playback is stopped and the reply begins; afterward, the queue refills using the updated conversation context.
If another line arrives while the reply is being prepared, Murmur merges it into the reply and invalidates the superseded preparation. In an ordinary interruption, the song continues underneath the process, with its volume lowered while voice plays.
Why queue clearing is not enough
Asynchronous work can finish after a queue has been cleared. Murmur guards against that with an incrementing epoch. A refill captures the current epoch, waits for generation, and enqueues its result only if the epoch is still current. An interruption increments the epoch, making older unfinished work ineligible for the queue.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteThis prevents stale results from changing the queue after they return, but it does not undo actions that have already completed. Requests already sent to a model service may continue consuming resources. Old voice can also continue while the new reply is prepared, so the time to a reply and any silence around it are distinct things to measure.
Rank #3
- Premium, around-ear, open back headphones: Audiophile sound combined with premium design and materials
- Padded headband and luxurious velour covered ear pads perfect for long listening sessions with no pressure on the ears
- Multiple connectivity options: Robust 3 meter detachable cable and 6.3 millimeter jack and additional 1.2 meter detachable cable with 3.5 millimeter Jack
- Timeless design cues: Ivory color, matte finish together with the brown headband stitching and matte metallic detail convey quality at first glance
- Premium Components: Sennheiser engineered transducers use aluminum voice coils delivering high efficiency, excellent dynamics and extremely low distortion
How voice and music share the audio timeline
Murmur builds its audio graph with node-web-audio-api, combining voice, the main song, and a background bed. It schedules gain changes ahead on the audio clock rather than relying on JavaScript timers. In the described implementation, the main song’s linear gain falls to 0.3 over about 0.3 seconds for speech, then returns over 2.5 seconds after speech ends. A linear gain of 0.3 is an amplitude ratio, not a claim that the song sounds “30% as loud.” These are listening-adjusted implementation settings, not general audio standards.
The background bed stays steady during speech and crossfades only when the main song enters or leaves. Speech synthesis returns a complete clip before playback begins. Knowing the clip’s duration helps schedule music recovery, interruptions, and joins, but the first line and each reply must wait for the full clip. By contrast, long music sources are decoded and queued in chunks, so an entire song need not be loaded before playback.
Song changes and introductions
For a transition to a new song, Murmur waits for the engine to confirm that audio was queued before it updates “now playing,” records the song, and plays its introduction. That confirmation establishes scheduling, not that sound reached the speakers.
After a song starts, the system generates a short coda. That lets the eventual return to talk use context from the current song instead of relying on a segment written before the song began.
Rank #4
- Open-back design with an extremely wide, dimensional sound stage and ultra-precise localization
- Uncolored frequency response (5 - 36,000 Hz) for honest, dynamic sound reproduction across the full spectrum. Innovative low-frequency cylinder system for full, accurate, and clearly defined low end./
- Sustainability-inspired with washable, replaceable pads and FSC-certified, forest-friendly packaging
- Sennheiser Open-frame Architecture reduces total harmonic distortion (THD) and minimizes resonance, improving audio accuracy
- Two unique sets of ear pads for producing or mixing help eliminate ear fatigue and pinpoint frequencies
How to interpret Murmur’s reported timing
The figures below are the author’s historical implementation records, not independently reproduced results. The article does not state a year for these logs.
| Measure | Reported result | What it represents |
|---|---|---|
| First two-segment batch | 24.5 seconds and 33.9 seconds in two full runs | Text generation through completed speech synthesis |
| One talk-segment refill | 9 to 14 seconds | Model call alone, before speech synthesis |
| Prefetched talk boundaries | 13 across two runs | Playback began in the same logged second as “talk.buffer warm”; logs have one-second resolution and do not establish zero latency |
| Startup to first song | Before changes: 136 seconds in a cold-start run and 195 seconds in a subsequent run with prior-session memory. After changes: 71 seconds and 78 seconds, respectively | Multiple changes were combined, so the comparison does not isolate one optimization |
| Music preparation after changes | 40.2 seconds and 54.7 seconds in the two after-change runs | Music preparation itself |
| First audible voice | Roughly 29 to 39 seconds | The first batch was not covered by prefetch |
| Music selection in another log | Roughly 82 to 192 seconds for each of five selections | Talk could continue while selection proceeded |
The before-and-after startup figures combine changes including earlier music selection, simpler search, and a limit on selection context. They are not evidence that any one change caused the full improvement. The runs also used separate data directories, preset personas, cached background beds, and a fixed “listener present” signal. The author did not rerun these figures for the article.
Which latency question are you trying to answer?
- “Why hasn’t anyone started talking?” Look at startup-to-first-audible-voice. Prefetching did not cover Murmur’s first batch in the reported measurements.
- “Why did it stop?” Measure extra wait beyond the configured pause at a playback boundary. A warm-buffer log entry in the same second as playback is not proof of zero waiting.
- “When will it answer me?” Measure from listener input to reply playback. This is not the same as startup latency or inter-segment waiting, and old voice may continue while a reply is prepared.
The author says the sample is too small for a meaningful long-run P95, and the one-second log resolution cannot establish zero latency. Natural transitions and failures in real services or audio sources require real runs followed by listening.
Recommended Free Tools
What this design suggests for other audio systems
- Separate content generation from playback control so model output does not directly manipulate audio state.
- Track whether queued speech is merely being synthesized or is actually ready to play.
- Invalidate stale asynchronous work when the listener’s intent changes; clearing a queue alone cannot stop old work from returning later.
- Measure startup, boundary waiting, and input-to-reply latency separately, because each captures a different user experience.
- Schedule gain changes on the audio clock when timing matters, and treat a scheduling confirmation as distinct from proof of audible output.
Apple’s AVFAudio documentation gives a platform-specific example of automatic interruption handling: “For example, AVPlayer monitors your app’s audio session and automatically pauses playback in response to interruption events.” Apple Developer Documentation: Handling audio interruptions. Murmur’s described behavior is different: during an ordinary listener interruption, its song continues under the reply at reduced gain.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

