Getting Started
When you first open Aegis, a friendly guided walkthrough greets you and shows you around. Here's what happens step by step.
The Guided Walkthrough
Step 1: Welcome
"Hi, I'm Aegis 👋 — I'm your personal AI surveillance agent. I watch, analyze, and alert — so you don't have to. Let me show you how to bring me to life in just a few steps."
You'll see a centered card with two options:
- Show me around — starts the walkthrough
- I'll explore myself — skips it (you can always set up manually)
Step 2: Give Me a Brain
The walkthrough points to the ⚙️ gear icon at the bottom of the sidebar.
"I need an LLM to think. Connect an API key (OpenAI, Anthropic, or a local model) and I'll be able to understand your camera feeds, describe what I see, and answer your questions."
What to do: Click the gear → expand Persona → click LLM. Paste an OpenAI API key, or enable the built-in local model. Full LLM setup guide →
Choosing between local and cloud:
| Scenario | Best Option | Why |
|---|---|---|
| You want full privacy, no internet | Built-in local model | Everything stays on your machine |
| You have a powerful GPU (RTX 4060+) | Built-in local model | Fast inference, zero API costs |
| You have a basic laptop, no GPU | Cloud API (OpenAI) | Fast responses without local compute |
| You want the best quality | Cloud API (GPT-5.x, Claude) | Largest, most capable models |
| You already run LM Studio | OpenAI-Compatible endpoint | Use your existing server |
Step 3: Give Me Eyes
"Cameras are how I see the world. Connect an RTSP stream, link your Blink or Ring account, or even use your webcam — I'll start watching and protecting immediately."
What to do: Click the gear → expand Cameras → choose your camera type:
- Blink cameras — link your Amazon/Blink account
- Ring cameras — link via the Ring mobile app integration
- LAN / RTSP / ONVIF cameras — auto-discover or manually add IP cameras
- Webcam or mobile device — use your built-in camera or a phone
Quickest path: If you just want to try Aegis, click Webcam (Builtin & USB) and select your laptop camera. Zero configuration — you'll have a working AI security feed in under 30 seconds.
Step 4: Your Visual Memory
The walkthrough points to the Timeline icon in the sidebar.
"Every motion event and clip I capture lands here — organized chronologically. You can browse, search, or ask me to analyze any moment. Nothing gets past me."
The Timeline is your visual history of everything Aegis has recorded. Each clip card shows a thumbnail, timestamp, camera name, duration, and — most importantly — the AI-generated description of what happened. This description is what powers intelligent search and contextual alerts.
Step 5: Let's Talk
The walkthrough points to the Aegis Home icon in the sidebar.
"This is our command center. Chat with me, ask 'what happened last night?', get summaries, or give me instructions. I'm always here."
Aegis Home is not a generic chatbot. The agent has real-time awareness of:
- What every camera currently sees (latest VLM analysis)
- Your complete clip history with AI descriptions
- Everything it has learned and stored in memory
- Active event handlers and their trigger history
- Available tools (web search, weather, camera control, etc.)
This means you can ask questions like "Was there anyone on the porch between 2pm and 4pm?" and get a real, data-backed answer.
The Setup Nudge
If you skip any steps, Aegis won't nag — but you'll see a subtle banner at the top of the screen:
Set up your LLM & Vision Model to unlock AI monitoring.
[Set Up Now →]
The setup nudge adapts based on what's missing. It checks for:
- LLM configuration — is a language model connected?
- VLM configuration — is a vision model active?
- Camera sources — is at least one camera added?
Click it to jump straight to whatever's missing. Once everything is configured, the banner disappears for good.
Quick Reference: The Sidebar
After setup, here's what you'll see in the left sidebar:
| Icon | View | What It Does |
|---|---|---|
| 🖥️ | Monitor | Live camera feeds in a configurable grid layout with drag-and-drop reordering |
| 📅 | Timeline | All captured clips and events, organized by time with camera filters and favorites |
| 🤖 | Aegis Home | Chat with your AI agent — ask questions, get summaries, control your system |
| 🧩 | Skills (bottom) | Open the Skills panel to browse, install, and manage AI capabilities |
| ⚙️ | Settings (bottom) | All configuration — agent persona, cameras, system settings |
Skill-Unlocked Views
Some views only appear after you install specific skills from the marketplace:
| View | Requires Skill | What It Does |
|---|---|---|
| Detection Studio | YOLO Detection | Real-time bounding box overlays on camera feeds with 80+ object classes |
| 3D Depth Vision | Depth Estimation | Privacy-preserving depth maps — anonymize feeds while preserving spatial awareness |
| Segmentation Studio | SAM2 Segmentation | Interactive click-to-segment with pixel-perfect masks and video tracking |
| Annotation Studio | Dataset Annotation | Label your camera data for custom model training with COCO export |
| OpenClaw | Camera Claw | Robotic arm control and monitoring |
Quick Reference: Settings Gear
The gear icon opens a flyout menu with two sections:
Agent
Expandable groups with sub-items:
Persona — configures the agent's identity and senses
- Soul — personality name, description, tone
- LLM — language model selection (built-in, OpenAI, or compatible endpoint)
- VLM — vision model selection (built-in, cloud providers, or external server)
- Channels — messaging integrations (Telegram, Discord, Slack)
- Voice — text-to-speech model and audio output device
- Memory — view and manage what the agent remembers
Abilities — additional agent capabilities
Cameras — camera source management
- Security Cameras — add and configure Blink, Ring, or LAN cameras
- Webcam — add webcams or desktop capture sources
System
Top-level system settings:
- AI Engine — manage the local inference engine
- Storage — disk usage, storage mode, retention policy
- Hardware — view CPU/GPU/RAM info and system recommendations
- Privacy — crash reporting toggle and blind mode
Next Steps
Once you've completed the walkthrough, here are the most valuable things to explore:
- Set up alerts — create event handlers like "notify me when someone approaches the front door"
- Configure messaging — receive alerts on Telegram, Discord, or Slack
- Customize the agent's personality — give it a name, tone, and behavioral instructions
- Install skills — add object detection, depth estimation, or segmentation
- Download additional models — try larger VLMs for more detailed scene descriptions