Voice
The Voice tab controls how Aegis AI speaks out loud. You can use the browser's built-in speech synthesis or download a local AI voice model for much higher quality.
Open via ⚙️ Settings → Persona → Voice.
Auto-TTS
When enabled, Aegis automatically speaks its responses out loud in addition to displaying text. This is useful for:
- Hands-free monitoring — hear alerts without looking at the screen
- Ambient awareness — the agent announces events while you're in another room
- Accessibility — screen reader alternative for visual descriptions
Disable this to use text-only chat. Voice announcements are sent to the selected audio output device.
TTS Models
Browser Default
Uses the Web Speech API built into your browser/OS. Zero download, works immediately.
Select a voice from the dropdown — available voices depend on your operating system:
- macOS: Samantha, Alex, Daniel, and many more (including Siri voices on newer versions)
- Windows: David, Zira, and additional voices installed via Windows Language Settings
- Linux: Varies by distribution (commonly espeak or Festival voices)
The browser default is the fastest option with zero setup, but sounds robotic compared to AI voices.
Downloadable AI Models
Higher-quality voices that run locally. Each model must be downloaded before use.
| Model | Size | Latency | License | Best For |
|---|---|---|---|---|
| Kokoro 82M | 82 MB | ~300ms | Apache 2.0 | Fast, lightweight voice. Best for older hardware or when speed matters most. |
| Chatterbox Turbo | 350 MB | ~200ms | MIT | Most natural conversational tone. Includes laughs, sighs, and breathing. Feels like a real person. |
| XTTS-v2 | 467 MB | ~150ms | CPML (non-commercial) | Multilingual support (13 languages). Can clone voices from a short audio sample. Best for non-English setups. |
| Qwen3-TTS 1.7B | 1.7 GB | ~97ms | Apache 2.0 | Highest quality overall. Instruction-controllable — you can tell it to whisper, speak excitedly, etc. Requires more RAM. |
Actions per model:
- Download — downloads the model with a progress bar showing percentage and speed
- Delete — removes a downloaded model from disk
- Test — plays a sample sentence with the selected voice so you can hear how it sounds before committing
Choosing a TTS Model
| Priority | Recommended Model | Why |
|---|---|---|
| Speed + minimal resources | Kokoro 82M | Smallest download, lowest latency, runs well even on CPU |
| Natural conversation | Chatterbox Turbo | Most human-like with natural speech patterns |
| Non-English language | XTTS-v2 | 13-language support with accent preservation |
| Best quality, have RAM/GPU | Qwen3-TTS 1.7B | Instruction-controllable, lowest latency, highest fidelity |
AI Voice (Orpheus)
Toggle AI Voice to use the Orpheus 3B neural voice model. This provides the most natural-sounding speech but requires more compute resources.
Select a voice (e.g. tara) from the voice dropdown. The model loads when first used — a progress indicator shows loading status. Orpheus produces exceptionally natural speech with appropriate emotion and pacing, making it ideal for extended conversations.
Audio Output Device
Select which audio device Aegis speaks through. The dropdown shows all available output devices on your system. Default: System Default.
This is useful when:
- You want alerts to play through a specific speaker (e.g., a smart speaker in the hallway)
- You have multiple audio outputs (headphones, speakers, Bluetooth)
- You're using Aegis on a setup with dedicated monitoring speakers
Voice Interaction
Listening to Alerts
When Auto-TTS is enabled and an event handler fires, the agent speaks the alert description through your selected audio device. This means you hear "Package delivery at the front door" even if you're not looking at the screen.
Push-to-Talk
In the Aegis Home chat, you can use push-to-talk for voice input. Hold the microphone button, speak your question or command, and release. Your speech is transcribed and sent as a chat message. The agent responds both in text and (if Auto-TTS is enabled) out loud.
This creates a fully voice-driven interaction — ask "what's happening outside?" and hear the answer spoken back to you.
Tips
- Start with Browser Default to test the feature, then download an AI model for better quality
- Kokoro 82M is the best starting point for AI voice — small download, Apache license, and works on all hardware
- Chatterbox Turbo is recommended if you interact with the agent frequently — the conversational tone makes extended interactions more pleasant
- If voice is enabled but you don't hear anything, check that the correct Audio Output Device is selected and system volume is up