Skip to main content

Voice

The Voice tab controls how Aegis AI speaks out loud. You can use the browser's built-in speech synthesis or download a local AI voice model for much higher quality.

Open via ⚙️ Settings → Persona → Voice.


Auto-TTS

When enabled, Aegis automatically speaks its responses out loud in addition to displaying text. This is useful for:

  • Hands-free monitoring — hear alerts without looking at the screen
  • Ambient awareness — the agent announces events while you're in another room
  • Accessibility — screen reader alternative for visual descriptions

Disable this to use text-only chat. Voice announcements are sent to the selected audio output device.


TTS Models

Browser Default

Uses the Web Speech API built into your browser/OS. Zero download, works immediately.

Select a voice from the dropdown — available voices depend on your operating system:

  • macOS: Samantha, Alex, Daniel, and many more (including Siri voices on newer versions)
  • Windows: David, Zira, and additional voices installed via Windows Language Settings
  • Linux: Varies by distribution (commonly espeak or Festival voices)

The browser default is the fastest option with zero setup, but sounds robotic compared to AI voices.

Downloadable AI Models

Higher-quality voices that run locally. Each model must be downloaded before use.

ModelSizeLatencyLicenseBest For
Kokoro 82M82 MB~300msApache 2.0Fast, lightweight voice. Best for older hardware or when speed matters most.
Chatterbox Turbo350 MB~200msMITMost natural conversational tone. Includes laughs, sighs, and breathing. Feels like a real person.
XTTS-v2467 MB~150msCPML (non-commercial)Multilingual support (13 languages). Can clone voices from a short audio sample. Best for non-English setups.
Qwen3-TTS 1.7B1.7 GB~97msApache 2.0Highest quality overall. Instruction-controllable — you can tell it to whisper, speak excitedly, etc. Requires more RAM.

Actions per model:

  • Download — downloads the model with a progress bar showing percentage and speed
  • Delete — removes a downloaded model from disk
  • Test — plays a sample sentence with the selected voice so you can hear how it sounds before committing

Choosing a TTS Model

PriorityRecommended ModelWhy
Speed + minimal resourcesKokoro 82MSmallest download, lowest latency, runs well even on CPU
Natural conversationChatterbox TurboMost human-like with natural speech patterns
Non-English languageXTTS-v213-language support with accent preservation
Best quality, have RAM/GPUQwen3-TTS 1.7BInstruction-controllable, lowest latency, highest fidelity

AI Voice (Orpheus)

Toggle AI Voice to use the Orpheus 3B neural voice model. This provides the most natural-sounding speech but requires more compute resources.

Select a voice (e.g. tara) from the voice dropdown. The model loads when first used — a progress indicator shows loading status. Orpheus produces exceptionally natural speech with appropriate emotion and pacing, making it ideal for extended conversations.


Audio Output Device

Select which audio device Aegis speaks through. The dropdown shows all available output devices on your system. Default: System Default.

This is useful when:

  • You want alerts to play through a specific speaker (e.g., a smart speaker in the hallway)
  • You have multiple audio outputs (headphones, speakers, Bluetooth)
  • You're using Aegis on a setup with dedicated monitoring speakers

Voice Interaction

Listening to Alerts

When Auto-TTS is enabled and an event handler fires, the agent speaks the alert description through your selected audio device. This means you hear "Package delivery at the front door" even if you're not looking at the screen.

Push-to-Talk

In the Aegis Home chat, you can use push-to-talk for voice input. Hold the microphone button, speak your question or command, and release. Your speech is transcribed and sent as a chat message. The agent responds both in text and (if Auto-TTS is enabled) out loud.

This creates a fully voice-driven interaction — ask "what's happening outside?" and hear the answer spoken back to you.


Tips

  • Start with Browser Default to test the feature, then download an AI model for better quality
  • Kokoro 82M is the best starting point for AI voice — small download, Apache license, and works on all hardware
  • Chatterbox Turbo is recommended if you interact with the agent frequently — the conversational tone makes extended interactions more pleasant
  • If voice is enabled but you don't hear anything, check that the correct Audio Output Device is selected and system volume is up