Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A local LLM can make Home Assistant voice control more flexible, but it does not have to handle every request. Home Assistant’s built-in intent matching is designed for routine commands; an LLM conversation agent can help with open-ended language and conversational replies. In my setup, I tried a local LLM and then stopped routing most of what I say through it. The exact model, hardware, and reasons for that choice are specific to my experience, not a universal verdict on local AI.
Why use a local LLM for Home Assistant voice control?
A local LLM is one possible conversation agent within Home Assistant Assist, not the whole voice assistant. A typical voice request passes through several stages: a microphone captures speech, speech-to-text transcribes it, a conversation agent interprets the text, Home Assistant executes an intent or action, and text-to-speech may speak a response. Home Assistant’s voice architecture overview describes these as distinct parts of the system.
With an LLM-based agent, the recognized text goes to the model, which can use Home Assistant tools through the Assist API. That API provides access to intents and entity capabilities available to the built-in conversation agent; it does not grant administrative access. The built-in Ollama integration connects Home Assistant to a separately running local Ollama server. When control is enabled, the model can provide information about and control only the entities exposed to it. See Home Assistant’s Ollama integration documentation and Assist API documentation.
That division matters: installing an LLM does not automatically make speech recognition or spoken responses more capable. Those components are configured separately. Home Assistant describes a fully local setup as one in which audio is transcribed, interpreted, and spoken back locally; that promise applies only when each part of the pipeline is configured to run locally. Its local voice setup guide explains the pieces.
#1 Best Overall
- Designed for Home Assistant Voice & Music Workflows: Preloaded with Home Assistant Voice Assistant and Music Assistant. Functions as both a voice input terminal and an audio playback endpoint.
- Dual Microphones for Voice Capture: Built with dual digital microphones for wake word or button-activated voice capture. Audio is streamed to the Home Assistant voice pipeline.
- Integrated 3W Speaker for Direct Playback: The built-in 3W/4Ω speaker supports TTS playback, Music Assistant streaming, and system audio without external speakers.
- Linux-Based Local Operation: Runs a lightweight Linux system on a quad-core ARM A53 CPU with 256MB RAM and 512MB flash for local audio processing.
- Development & Debugging Capabilities: Supports firmware flashing, and also provides access to live logs, on-device editing—suitable for routine development or issue diagnosis.
Why I stopped sending most requests to the LLM
For ordinary home control, a predictable intent match can be a better fit than asking a language model to interpret every phrase. Turning on a light or adjusting a thermostat is usually a bounded task. An LLM becomes more attractive when wording is less predictable, when a request is open-ended, or when a conversational answer is useful. That is the distinction behind my choice to use the LLM selectively rather than as the default for everything I say.
Home Assistant’s own cautions make that selective approach reasonable, though they do not explain my personal decision or prove that every local setup behaves the same way. The Ollama integration labels Home Assistant control experimental. It requires a model that supports tool calling, and the documentation warns that smaller models can make more mistakes and may not reliably sustain a conversation when control is enabled. Home Assistant recommends exposing fewer than 25 entities while experimenting.
Rank #2
- 🎙️ Hands-Free Voice Typing for Windows & Mac – Powered by iOS & Android dictation technology, AI VoiceWriter allows fast, accurate speech-to-text directly on your desktop. Simply speak, and your words appear in real time. Compatible with Windows 10 & above, macOS 13 & above.
- ✍️ AI Writing Assistant for Effortless Editing – Boost productivity with AI proofreading, rephrasing, and formatting. Perfect for emails, reports, creative writing, and professional content.
- 💻 Works Seamlessly in Any Desktop App – Type with your voice in Microsoft Word, Google Docs, PowerPoint, Teams, emails, and more. Just place your cursor in any text field and start speaking!
- 📱 Mobile App for Enhanced Voice Input – The AI VoiceWriter mobile app enhances voice recognition by using your phone’s microphone as an input device for clearer, more accurate dictation—while typing on your desktop. Supports iOS 15 & above, Android 9.0 & above.
- 🌎 Multilingual Voice Typing & AI Assistance – Supports 33 languages for dictation, plus AI-powered features in Chinese, English, Japanese, Korean, French, German, Spanish, Italian and, Swedish.
- Entity exposure: the model can work only with the entities you make available to it, so exposing a small, relevant set limits its scope.
- Tool support: a model must support tools to control Home Assistant through this integration.
- Reliability: Home Assistant warns that smaller models may make mistakes and struggle to maintain a conversation with control enabled.
- Experimental status: the integration’s device-control feature is explicitly described as experimental, rather than a guaranteed replacement for built-in intents.
Home Assistant also documents a way to separate the roles: configure the same model twice, once for conversation without control and once with control. That can keep open-ended conversation distinct from commands that can affect devices. The documentation does not establish that this arrangement is right for every installation; it is an option to consider when deciding which requests should have access to home-control tools.
Built-in intents and an LLM are different tools
The choice is not simply “AI or no AI.” Home Assistant’s built-in conversation agent matches recognized text to intents. Integrations can add external conversation agents, including LLM-based ones. For specific commands, custom sentences and intents can provide explicit handling without relying on a general-purpose model. The developer overview explains the built-in and integration-based roles.
Rank #3
- Your favorite music and content – Play music, audiobooks, and podcasts from Amazon Music, Apple Music, Spotify and others or via Bluetooth throughout your home.
- Alexa is happy to help – Ask Alexa for weather updates and to set hands-free timers, get answers to your questions and even hear jokes. Need a few extra minutes in the morning? Just tap your Echo Dot to snooze your alarm.
- Keep your home comfortable – Control compatible smart home devices with your voice and routines triggered by built-in motion or indoor temperature sensors. Create routines to automatically turn on lights when you walk into a room, or start a fan if the inside temperature goes above your comfort zone.
- Do more with device pairing – Fill your home with music using compatible Echo devices in different rooms, or create a home theatre system with Fire TV.
- Say goodbye to drop-offs and buffering - With eero Built-in, Echo Dot doubles as a mesh wifi extender, adding up to 1,000 sq. ft. of wifi coverage to your existing eero network.
| Need | Built-in intents or custom sentences | LLM conversation agent |
|---|---|---|
| Routine, bounded device commands | Matches text to supported intents; custom sentences can cover chosen phrases. | Can control exposed entities if the model supports tools, but device control is experimental. |
| Open-ended phrasing | Works within the intents and sentences configured for the system. | Can interpret more varied language, subject to the model’s capability and reliability. |
| Conversational follow-up | Home Assistant’s documentation does not characterize built-in intent matching as an open-ended dialogue system. | May enable more conversational responses, but Home Assistant cautions that smaller models may not reliably maintain a conversation when control is enabled. |
| Sentence triggers | Custom sentences and intents can support explicit local handling for selected phrases. | Ollama does not integrate with sentence triggers; external agents use them only when “Prefer handling commands locally” is enabled. |
The sentence-trigger behavior is documented in the Ollama integration and Conversation integration pages. In practice, choose the handler by the job: keep frequent, well-defined commands explicit and local where that suits your setup; reserve an LLM for language or replies that benefit from its flexibility.
Local speech recognition affects the experience too
If a request feels slow or gets transcribed incorrectly, the conversation agent may not be the only factor. Home Assistant’s local voice guide distinguishes between two speech-to-text options:
Rank #4
- Speech-to-Phrase is a closed-ended model that transcribes supported phrases and covers a subset of Assist commands. Home Assistant reports processing in under one second on Home Assistant Green or Raspberry Pi 4. It is positioned for home control rather than unrestricted dictation.
- Whisper is open-ended. Home Assistant reports around eight seconds for transcription on Raspberry Pi 4 and under one second on an Intel NUC. The guide recommends it for households with more powerful hardware that want to extend voice beyond simple control, for example by pairing it with an LLM.
These are Home Assistant’s published figures for the named hardware, not guarantees for other machines, languages, or configurations. The guide also says performance and speech quality vary by device and language. For spoken replies, Piper is a separate local text-to-speech option; Home Assistant reports that medium-quality models generate 1.6 seconds of speech per second on a Raspberry Pi. See the local voice guide for its setup context.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →As a separate point of reference, a 2025 Aalborg University paper by Rune Birkmose, Nathan Mørkeberg Reece, Esben Hofstedt Norvin, Johannes Bjerva, and Mike Zhang evaluated fine-tuned on-device LLMs for Home Assistant. It reports approximately 80–86% accuracy on noisy human prompts and out-of-domain intents, with an average inference time of 5–6 seconds per query. The authors describe that latency as acceptable for one-shot commands but suboptimal for multi-turn dialogue. Those results apply to the paper’s models and evaluation tasks, not to every local model or an individual installation: the study.
Best Value
- Your favorite music and content – Play music, audiobooks, and podcasts from Amazon Music, Apple Music, Spotify and others or via Bluetooth throughout your home.
- Alexa is happy to help – Ask Alexa for weather updates and to set hands-free timers, get answers to your questions and even hear jokes. Need a few extra minutes in the morning? Just tap your Echo Dot to snooze your alarm.
- Keep your home comfortable – Control compatible smart home devices with your voice and routines triggered by built-in motion or indoor temperature sensors. Create routines to automatically turn on lights when you walk into a room, or start a fan if the inside temperature goes above your comfort zone.
- Do more with device pairing – Fill your home with music using compatible Echo devices in different rooms, or create a home theatre system with Fire TV.
- Say goodbye to drop-offs and buffering - With eero Built-in, Echo Dot doubles as a mesh wifi extender, adding up to 1,000 sq. ft. of wifi coverage to your existing eero network.
How to decide what should use the LLM
Start with the requests you actually make and separate them by purpose rather than routing everything through one agent by default.
- List routine commands. Identify recurring requests with clear outcomes, such as controlling a light or thermostat. Check whether Home Assistant already handles them through built-in intents or whether a custom sentence would cover them.
- Identify requests that need interpretation. Keep examples where phrasing is varied, the request is open-ended, or a conversational reply adds value. Those are stronger candidates for an LLM agent.
- Limit the LLM’s reach. If enabling control, expose only the entities the agent needs. Home Assistant recommends fewer than 25 entities for experimentation, and its integration requires a tool-capable model.
- Check the whole pipeline. Consider speech-recognition coverage and latency, the conversation agent, and text-to-speech separately. A change to the LLM does not itself change the capabilities of the other stages.
- Try the split deliberately. Compare representative routine phrases and open-ended requests with the handling you have configured. Keep the LLM where its flexibility helps, and use local intent handling for phrases that are better served explicitly.
Home Assistant’s Ollama integration guide also describes using two configurations of the same model with different prompts—one without device control and one with it—if separating conversation from control is useful.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

