# Jarvis AI + socket.io-computer Integration Plan for WakelAI.com

## Executive Summary

This plan details the integration of two powerful technologies into the WakelAI platform:

1. **Jarvis AI** ([eadmin2/jarvis_ai](https://github.com/eadmin2/jarvis_ai)) - Iron-Man-style voice assistant with holographic HUD for Hermes Agent
2. **socket.io-computer** ([kevin-roark/socket.io-computer](https://github.com/kevin-roark/socket.io-computer)) - Collaborative virtual machine with QEMU streaming

Together, these will enable WakelAI users to have voice-controlled access to collaborative virtual machines, creating a unique "AI-powered cloud workspace" experience.

---

## Part 1: Jarvis AI Integration

### 1.1 What Jarvis AI Does

Jarvis AI is a sophisticated voice assistant overlay for Hermes Agent that provides:

- **Live Voice Interaction**: Click-to-talk interface with real-time transcription display
- **Hermes Agent Integration**: Full access to Hermes Agent's capabilities (files, terminal, web search, memory)
- **Holographic HUD**: Iron-Man-style control center with live agent activity monitoring
- **Media Panels**: Agent-summoned content displays (videos, images, dashboards)
- **Multi-device Support**: Works on any browser on the LAN with push-to-talk clients
- **Interrupt-aware**: Barge-in capability with precise memory of what was heard
- **Privacy Protection**: Secret redaction before cloud TTS, LAN-only architecture

### 1.2 Technical Architecture

```
Browser HUD (any LAN device)          Host Machine
 ── https/wss :443 ──────────┐   ┌──────────────────────────────────────┐
   mic · speaker · panels    ├──►│ Jarvis Voice Pipeline Server         │
                             │   │  - STT: faster-whisper (local)       │
 Push-to-talk client         │   │  - TTS: ElevenLabs Flash (streaming) │
 ── ws :8765 ────────────────┘   │  - HUD + auth + dashboard TLS proxy  │
                                 └────────────────┬───────────────────────┘
                                                  │
                                                 ─┼─
                                                  │
                                     ┌────────────┴────────────┐
                                     │ Hermes Agent API :8642   │
                                     │  - Memory · Tools       │
                                     │  - Skills · 80+         │
                                     │  - Terminal · Web       │
                                     └─────────────────────────┘
```

### 1.3 Key Components

**Server (jarvis_ai/server/)**
- `server.py` - FastAPI voice pipeline server
- `hud/` - Single-file HUD (vanilla JS, no build step)
- `scripts/` - Start/stop/health/cert & boot audio generators
- `config/server.yaml` - Main configuration

**Client (jarvis_ai/client/)**
- Optional Windows/Linux push-to-talk Python client
- Wake word capable

**Worker (jarvis_ai/worker/)**
- GPU sidecars for big-model STT
- Stats agent for Machines panel

**Hermes Plugin (jarvis_ai/hermes-plugin/)**
- `hud_display` tool - Lets agent summon/dismiss HUD media panels

### 1.4 Requirements for Integration

**Hermes Agent Setup:**
```bash
# Required in ~/.hermes/.env
API_SERVER_ENABLED=true
API_SERVER_KEY=<generated-key>
JARVIS_HUD_TOKEN=jarvis-<token>
ELEVENLABS_API_KEY=<user-key>
```

**Python Dependencies:**
```
fastapi, uvicorn, requests, pyyaml, numpy, anthropic
RealtimeSTT, faster-whisper, silero-vad, websockets, psutil
```

**TLS Certificates** (required for browser mic):
- Self-signed certs for HTTPS
- Must be trusted on each device

**System Requirements:**
- Python 3.11+
- ~500MB disk space (Whisper models)
- Port 8765 (WebSocket), 443 (HTTPS)
- Optional: NVIDIA GPU for faster STT

### 1.5 API Endpoints

| Endpoint | Purpose |
|---|---|
| `/hud/` | The HUD interface (static, single file) |
| `/ws` | WebSocket for voice streaming |
| `/api/hermes/{path}` | Allowlist proxy to Hermes API |
| `/api/chat` | Typed chat on shared session |
| `/api/machines` | Host stats + remote workers |
| `/api/usage` | Local token/char tally |
| `/api/summon` | Broadcast media panels to HUD |

### 1.6 WebSocket Protocol

**Client → Server:**
```json
{"type":"start","sample_rate":16000,"format":"pcm_s16le","channels":1}
{"type":"stop"}
{"type":"stop_run"}
{"type":"approval_decision","run_id":"...","approval_id":"...","decision":"allow|deny"}
```

**Server → Client:**
```
status · partial_transcript · transcript · run_started{run_id}
agent_status{state: thinking|tool_use|speaking|stopped}
approval_request{data,run_id} · error · done{timing}
```

---

## Part 2: socket.io-computer Integration

### 2.1 What socket.io-computer Does

socket.io-computer is a collaborative virtual machine where:
- Multiple users can take turns controlling a VM
- QEMU runs the actual VM
- Server streams video output to browsers via WebSocket
- Users interact through VNC-like controls in the browser
- Turn-based system for multi-user collaboration

### 2.2 Technical Architecture

```
Browser                              Server
 ───────────────┐               ┌─────────────────┐
 Canvas + VNC   │◄── ws ────────│ QEMU Instance    │
 Controls       │               │ (VM emulation)  │
 ───────────────┘               └─────────────────┘
                                          │
                                         ─┼─
                                          │
                                     ┌────┴────┐
                                     │ Redis    │
                                     │ (state)  │
                                     └───────────┘
```

### 2.3 Key Components

**Core Files:**
- `app.js` - Web server (Express)
- `io.js` - Socket.IO server
- `qemu.js` - QEMU instance management
- `emu-runner.js` - Emulator communication
- `computer.js` - Computer state management
- `vnc.js` - VNC protocol handling
- `presence.js` - User presence/turn management
- `redis.js` - Redis integration

**Dependencies:**
- `qemu` - System package for virtualization
- `redis-server` - State management
- `express` - Web framework
- `socket.io` - WebSocket framework

### 2.4 Requirements for Integration

**System Requirements:**
- QEMU installed (system package)
- Redis server running
- Node.js environment
- Sufficient disk space for VM images

**VM Image Setup:**
```bash
qemu-img create -f qcow2 vm.img 3G
```

**Runtime Environment Variables:**
```bash
COMPUTER_ISO=path/to/install.iso
COMPUTER_IMG=path/to/disk.img
```

### 2.5 Architecture Notes

- **Separate Processes**: web server, Socket.IO server, QEMU instance, emulator runner
- **Collaborative**: Turn-based control with presence indicators
- **Browser-based**: No client installation needed
- **Video Streaming**: Raw frame data over WebSocket

---

## Part 3: Integration Strategy for WakelAI

### 3.1 Proposed Architecture

```
┌─────────────────────────────────────────────────────────────────┐
│                        WakelAI Platform                          │
├─────────────────────────────────────────────────────────────────┤
│                                                                   │
│  ┌───────────────────────────────────────────────────────────┐  │
│  │              WakelAI Dashboard / Instances                  │  │
│  │  - User creates "AI Workspace" instance                    │  │
│  │  - Subdomain: workspace-xxx.wakelai.com                    │  │
│  └───────────────────────────────────────────────────────────┘  │
│                              │                                    │
│                              ▼                                    │
│  ┌───────────────────────────────────────────────────────────┐  │
│  │              Docker Compose Stack                          │  │
│  ├───────────────────────────────────────────────────────────┤  │
│  │                                                             │  │
│  │  ┌──────────────┐    ┌──────────────┐    ┌──────────────┐  │  │
│  │  │ Jarvis Voice │    │ socket.io-   │    │   VM         │  │  │
│  │  │   Server     │◄──►│   Computer    │◄──►│  Instance    │  │  │
│  │  │  (FastAPI)   │    │   (Node.js)  │    │   (QEMU)     │  │  │
│  │  │              │    │              │    │              │  │  │
│  │  │ Port: 8765   │    │ Port: 5000   │    │ VNC: 5900   │  │
│  │  └──────────────┘    └──────────────┘    └──────────────┘  │  │
│  │                                                             │  │
│  │  ┌──────────────┐    ┌──────────────┐                     │  │
│  │  │ Hermes Agent │    │    Redis     │                     │  │
│  │  │  (API:8642)  │    │  (:6379)     │                     │  │
│  │  └──────────────┘    └──────────────┘                     │  │
│  │                                                             │  │
│  └───────────────────────────────────────────────────────────┘  │
│                              │                                    │
│                              ▼                                    │
│  ┌───────────────────────────────────────────────────────────┐  │
│  │              User Experience                               │  │
│  ├───────────────────────────────────────────────────────────┤  │
│  │  - Voice HUD overlay for Jarvis                            │  │
│  │  - VM control panel embedded                               │  │
│  │  - Unified authentication                                   │  │
│  │  - Persistent workspaces                                   │  │
│  └───────────────────────────────────────────────────────────┘  │
│                                                                   │
└─────────────────────────────────────────────────────────────────┘
```

### 3.2 Integration Points

**1. Jarvis ↔ Hermes Agent**
- Jarvis connects to Hermes Agent API
- Voice commands trigger Hermes skills
- VM control becomes a Hermes skill

**2. Jarvis ↔ socket.io-computer**
- Jarvis displays VM panel as media panel
- Voice commands control VM through Hermes skill
- VM state visible in Jarvis HUD

**3. socket.io-computer ↔ VM**
- Standard QEMU integration
- VNC protocol for display/control
- Redis for collaborative state

**4. Unified Dashboard**
- Single WakelAI instance manages all components
- Subdomain routing to different services
- Shared authentication via JARVIS_HUD_TOKEN

### 3.3 Data Flow

**Voice Command to VM Control:**
```
1. User speaks: "Start the Windows VM"
2. Jarvis STT transcribes
3. Hermes receives transcript
4. Hermes skill "vm_control" processes command
5. Skill calls socket.io-computer API
6. QEMU VM starts/receives command
7. Video streams to browser
8. Jarvis displays VM panel
9. User sees VM booting
```

**Collaborative Session:**
```
1. User creates "AI Workspace" instance
2. Jarvis authenticates user
3. socket.io-computer initializes VM
4. Jarvis shows VM panel with "Your turn"
5. Other users join via subdomain
6. Presence system manages turn queue
7. Voice commands during active turn
```

---

## Part 4: Implementation Plan

### Phase 1: Foundation (Week 1-2)

**1.1 Docker Containerization**
- Create Docker image for Jarvis voice server
- Create Docker image for socket.io-computer
- Configure Docker Compose orchestration
- Set up networking between containers

**1.2 WakelAI Integration**
- Update `src/lib/dokploy.ts` for workspace instances
- Add "AI Workspace" instance type
- Create workspace provisioning endpoint
- Add subdomain routing for services

**1.3 Authentication Bridge**
- Extend WakelAI auth to Jarvis
- Generate JARVIS_HUD_TOKEN per instance
- Share API keys between services
- Implement session management

### Phase 2: Jarvis Integration (Week 3-4)

**2.1 Hermes Agent Setup**
- Package Hermes Agent for container deployment
- Configure API server for each instance
- Set up persistent memory per workspace
- Enable required skills

**2.2 Voice Pipeline**
- Deploy Jarvis voice server
- Configure Whisper STT (local)
- Set up ElevenLabs TTS integration
- Implement WebSocket streaming

**2.3 HUD Integration**
- Embed Jarvis HUD in WakelAI dashboard
- Add authentication iframe
- Implement media panel system
- Style to match WakelAI theme

### Phase 3: socket.io-computer Integration (Week 5-6)

**3.1 VM Infrastructure**
- Set up QEMU environment
- Configure VM image storage
- Implement Redis for state
- Create VM templates (Windows, Linux)

**3.2 WebSocket Streaming**
- Deploy socket.io-computer server
- Configure VNC protocol bridge
- Implement video streaming
- Add browser-based controls

**3.3 Collaborative Features**
- Implement turn system
- Add presence indicators
- Create user queue
- Implement session persistence

### Phase 4: Hermes Skill Development (Week 7)

**4.1 VM Control Skill**
- Create Hermes skill for VM operations
- Implement command parsing
- Add safety checks
- Create approval workflows

**4.2 Integration Commands**
- "Start VM [name]"
- "Stop VM [name]"
- "Show VM screen"
- "Install [software] on VM"
- "Open [website] in VM"

### Phase 5: Testing & Optimization (Week 8)

**5.1 Integration Testing**
- End-to-end voice-to-VM flow
- Multi-user collaboration
- Performance under load
- Error handling

**5.2 Optimization**
- Reduce voice latency
- Optimize video streaming
- Improve VM boot times
- Memory/CPU tuning

---

## Part 5: Configuration Requirements

### 5.1 Environment Variables

**WakelAI Instance:**
```bash
WORKSPACE_TYPE=jarvis_vm
JARVIS_HUD_TOKEN=<generated>
HERMES_API_KEY=<shared>
ELEVENLABS_API_KEY=<user-provided>
REDIS_URL=redis://localhost:6379
VM_STORAGE_PATH=/data/vms
```

**Jarvis Server:**
```bash
API_SERVER_KEY=<from-hermes>
HERMES_BASE_URL=http://hermes:8642
ELEVENLABS_API_KEY=<from-env>
ELEVENLABS_VOICE_ID=<user-preference>
STT_MODEL=small.en
TLS_CERT_PATH=/certs
```

**socket.io-computer:**
```bash
COMPUTER_VM_DIR=/vms
COMPUTER_ISO_DIR=/isos
REDIS_HOST=redis
TURN_DURATION_MINUTES=5
```

### 5.2 Docker Services

**docker-compose.yml structure:**
```yaml
services:
  jarvis-voice:
    image: wakelai/jarvis-voice
    ports: ["8765:8765", "443:443"]
    environment:
      - HERMES_URL=http://hermes:8642
    volumes:
      - vm-storage:/vms
      - certs:/certs

  hermes:
    image: wakelai/hermes-agent
    ports: ["8642:8642"]
    environment:
      - API_SERVER_ENABLED=true
      - API_SERVER_KEY=${API_KEY}
    volumes:
      - hermes-data:/data
      - hermes-plugins:/plugins

  socket-computer:
    image: wakelai/socket-computer
    ports: ["5000:5000"]
    environment:
      - REDIS_URL=redis://redis:6379
    volumes:
      - vm-storage:/vms
      - iso-storage:/isos

  redis:
    image: redis:alpine
    volumes:
      - redis-data:/data

volumes:
  vm-storage:
  iso-storage:
  hermes-data:
  hermes-plugins:
  redis-data:
  certs:
```

---

## Part 6: API Integration Points

### 6.1 Jarvis API Extensions

```typescript
// New endpoints for WakelAI integration
POST /api/workspaces/:id/jarvis/config
  - Update Jarvis configuration
  - Set voice preferences
  - Manage skill enablement

GET /api/workspaces/:id/jarvis/status
  - Voice server status
  - Active conversation info
  - Resource usage

POST /api/workspaces/:id/jarvis/command
  - Send text command to Jarvis
  - Execute skill directly
  - Return response
```

### 6.2 socket.io-computer API Extensions

```typescript
// VM management endpoints
POST /api/workspaces/:id/vm/create
  - Create new VM instance
  - Specify OS template
  - Allocate resources

POST /api/workspaces/:id/vm/:vm_id/control
  - Send keyboard/mouse input
  - Execute VM commands
  - Manage screen sharing

GET /api/workspaces/:id/vm/:vm_id/screen
  - Get current screen frame
  - Stream video data
  - WebSocket endpoint for real-time
```

### 6.3 Hermes Skill API

```typescript
// Skill for VM control
interface VMControlSkill {
  startVM(vmId: string): Promise<void>
  stopVM(vmId: string): Promise<void>
  getVMStatus(vmId: string): Promise<VMStatus>
  executeCommand(vmId: string, command: string): Promise<string>
  openURL(vmId: string, url: string): Promise<void>
}
```

---

## Part 7: Security Considerations

### 7.1 Authentication Flow

1. User authenticates with WakelAI
2. WakelAI generates JARVIS_HUD_TOKEN
3. Token passed to Jarvis and socket.io-computer
4. Each service validates token independently
5. Sessions tracked via Redis

### 7.2 Network Security

- All services on internal Docker network
- Only Traefik exposed externally
- TLS for all external connections
- API keys never exposed to browser

### 7.3 VM Isolation

- Each VM in separate network namespace
- No direct internet access (proxy through host)
- Disk I/O limited per workspace
- Resource quotas per user

---

## Part 8: Resource Planning

### 8.1 Per-Instance Requirements

**Minimum:**
- 2 vCPU
- 4GB RAM
- 20GB disk (VM + models)

**Recommended:**
- 4 vCPU
- 8GB RAM
- 50GB disk (multiple VMs + models)

**Maximum:**
- 8 vCPU
- 16GB RAM
- 100GB disk (enterprise workspace)

### 8.2 Platform Requirements

**Development:**
- Docker build environment
- QEMU for testing
- Hermes Agent setup
- ElevenLabs test account

**Production:**
- Docker orchestration (Traefik/Docker Compose)
- Redis cluster for scaling
- Load balancer for multi-instance
- Storage for VM images

---

## Part 9: Success Metrics

### 9.1 Technical Metrics

- Voice latency < 4s (end-to-end)
- VM boot time < 30s
- Video streaming FPS > 15
- Concurrent users per VM > 5
- Uptime > 99%

### 9.2 User Experience Metrics

- First-toke success rate > 95%
- Command recognition accuracy > 90%
- VM control responsiveness < 1s
- Session setup time < 2 minutes

---

## Part 10: Timeline & Milestones

### Week 1-2: Foundation
- Docker containers built
- Basic provisioning working
- Authentication bridge complete

### Week 3-4: Jarvis Integration
- Voice pipeline operational
- HUD embedded in dashboard
- Hermes Agent connected

### Week 5-6: socket.io-computer Integration
- VM infrastructure ready
- Video streaming working
- Collaborative features complete

### Week 7: Skills & Integration
- VM control skill developed
- End-to-end flow working
- Basic commands operational

### Week 8: Testing & Launch
- Performance optimization
- User acceptance testing
- Production deployment

---

## Part 11: Rollout Strategy

### 11.1 Beta Phase
- Limited to 10 users
- Single VM template (Linux)
- Reduced resource limits
- Active monitoring

### 11.2 General Availability
- All subscription tiers
- Multiple VM templates
- Full resource allocation
- Automated scaling

### 11.3 Enterprise Features
- GPU acceleration for voice
- Dedicated VM instances
- Custom branding
- API access

---

## Conclusion

This integration creates a unique "AI-powered workspace" where users can:

1. **Talk to their workspace** - Voice commands to control everything
2. **See their AI work** - Visual feedback through Jarvis HUD
3. **Collaborate in real-time** - Multi-user VM sessions
4. **Automate tasks** - Hermes Agent skills + VM control
5. **Work from anywhere** - Browser-based, mobile-ready

The combination of Jarvis AI's sophisticated voice interface with socket.io-computer's collaborative VM platform, powered by WakelAI's infrastructure and Hermes Agent's capabilities, creates a powerful new computing paradigm.

---

## Sources

- [Jarvis AI GitHub Repository](https://github.com/eadmin2/jarvis_ai)
- [socket.io-computer GitHub Repository](https://github.com/kevin-roark/socket.io-computer)
- [Hermes Agent Documentation](https://hermes-agent.nousresearch.com/docs/getting-started/quickstart)
- [Socket.IO Documentation](https://socket.io/docs/v4/)
- [Xuanwo's Blog - Socket VM Architecture](https://xuanwo.io/2015/07/06/socket-vm/)
- [SkillsLLM - Jarvis AI Information](https://skillsllm.com/skill/jarvis-ai)
