# Jarvis AI Resource Requirements (ElevenLabs/OpenAI Only)

## Per-Instance Resource Requirements

### With ElevenLabs Flash TTS

| Component | RAM | CPU | Disk | Notes |
|-----------|-----|-----|------|-------|
| **FastAPI Server** | 50-100MB | 0.1-0.2 vCPU | ~50MB | Core server |
| **Whisper STT (tiny.en)** | ~273MB | 0.3-0.5 vCPU | ~75MB | Live transcription |
| **Whisper STT (small.en)** | ~460MB | 0.5-1 vCPU | ~460MB | Better accuracy |
| **ElevenLabs TTS** | **0 MB** | **0 vCPU** | **0** | Cloud API |
| **WebSocket Handler** | 20-50MB | 0.1 vCPU | 0 | Per connection |
| **HUD Static Files** | <1MB | 0 | ~1MB | Served from disk |
| **Buffers & Processing** | 50-100MB | 0.1 vCPU | 0 | Audio buffers |
| **TOTAL (tiny.en)** | **~400-500MB** | **~0.6-0.8 vCPU** | **~500MB** | |
| **TOTAL (small.en)** | **~600-700MB** | **~0.8-1.2 vCPU** | **~500MB** | |

### With OpenAI TTS

| Component | RAM | CPU | Disk | Notes |
|-----------|-----|-----|------|-------|
| **FastAPI Server** | 50-100MB | 0.1-0.2 vCPU | ~50MB | Core server |
| **Whisper STT (tiny.en)** | ~273MB | 0.3-0.5 vCPU | ~75MB | Live transcription |
| **Whisper STT (small.en)** | ~460MB | 0.5-1 vCPU | ~460MB | Better accuracy |
| **OpenAI TTS** | **0 MB** | **0 vCPU** | **0** | Cloud API |
| **WebSocket Handler** | 20-50MB | 0.1 vCPU | 0 | Per connection |
| **HUD Static Files** | <1MB | 0 | ~1MB | Served from disk |
| **Buffers & Processing** | 50-100MB | 0.1 vCPU | 0 | Audio buffers |
| **TOTAL (tiny.en)** | **~400-500MB** | **~0.6-0.8 vCPU** | **~500MB** | |
| **TOTAL (small.en)** | **~600-700MB** | **~0.8-1.2 vCPU** | **~500MB** | |

---

## Key Insights

### ✅ TTS Uses Zero Local Resources
Both ElevenLabs and OpenAI TTS are **cloud APIs**:
- **0 MB RAM** - No local models needed
- **0 vCPU** - No local processing
- **0 MB disk** - No local storage
- Audio streams directly to client browser

### 🎯 STT is the Main Resource Consumer
Whisper transcription accounts for **60-70%** of total resources:
- **tiny.en**: 273MB RAM, 0.3-0.5 vCPU
- **small.en**: 460MB RAM, 0.5-1 vCPU

### 📊 Idle vs Active Resource Usage

**Idle (waiting for user):**
- RAM: 400-500MB (models loaded)
- CPU: ~0.1 vCPU (just keeping server running)

**Active (user speaking + AI responding):**
- RAM: 400-500MB (same)
- CPU: 0.6-1.2 vCPU (STT processing + streaming)

**After response (waiting again):**
- RAM: 400-500MB (same)
- CPU: ~0.1 vCPU (back to idle)

---

## Cost Comparison

### ElevenLabs Flash
- **Cost**: $0.05 per 1,000 characters
- **Latency**: ~75ms (very fast)
- **Quality**: Excellent, natural voices

**Monthly Estimate** (1000 turns, 1000 chars each):
- 1,000,000 characters = **$50/month**

### OpenAI TTS
- **Cost**: $0.015 per 1,000 characters (standard)
- **Cost**: $0.030 per 1,000 characters (HD)
- **Latency**: ~200-300ms
- **Quality**: Very good, slightly less natural than ElevenLabs

**Monthly Estimate** (1000 turns, 1000 chars each):
- 1,000,000 characters = **$15/month** (standard)
- 1,000,000 characters = **$30/month** (HD)

---

## Scalability Analysis

### Per-User Resources

| Configuration | RAM | CPU | Disk |
|--------------|-----|-----|------|
| **Tiny.en + ElevenLabs** | 400-500MB | 0.6-0.8 vCPU | 500MB |
| **Tiny.en + OpenAI** | 400-500MB | 0.6-0.8 vCPU | 500MB |
| **Small.en + ElevenLabs** | 600-700MB | 0.8-1.2 vCPU | 500MB |
| **Small.en + OpenAI** | 600-700MB | 0.8-1.2 vCPU | 500MB |

### Multi-User Scenarios

**Shared Model Server Architecture** (recommended for 5+ users):

```
Central Model Server (shared)
    ├── Whisper Model: 273-460MB
    ├── FastAPI: 100MB
    └── Total: ~400-600MB

Per-User Session (marginal cost)
    ├── WebSocket handler: 20-50MB
    ├── Audio buffers: 50-100MB
    └── Total: ~70-150MB per active user
```

**Resource Examples:**

| Users | Total RAM | Total CPU | Architecture |
|-------|-----------|-----------|--------------|
| **1 user** | 400-500MB | 0.6-0.8 vCPU | Single instance |
| **5 users** | 700-900MB | 1-1.5 vCPU | Shared model |
| **10 users** | 1-1.2GB | 1.5-2 vCPU | Shared model |
| **50 users** | 3-4GB | 4-6 vCPU | Shared model + load balancing |
| **100 users** | 5-7GB | 7-10 vCPU | Shared model + horizontal scaling |

---

## Optimization Strategies

### 1. Shared Model Server (75% RAM savings for 5+ users)
```python
# Instead of: 5 users × 500MB = 2.5GB
# Use: 1 shared model (500MB) + 5 × 100MB = 1GB
```

### 2. Use tiny.en Model (40% RAM reduction)
```python
# small.en: 460MB
# tiny.en: 273MB
# Trade-off: Slightly lower accuracy, much faster
```

### 3. Lazy Model Loading (Scale to zero when idle)
```python
# Only load Whisper when first user speaks
# Unload after N minutes of inactivity
# Saves 273-460MB when completely idle
```

### 4. Browser-Native STT (Zero local STT cost)
```typescript
// Use Web Speech API (Chrome/Safari)
// 0 MB RAM, 0 vCPU for STT
// Trade-off: Requires user's device power
```

---

## Recommended Configuration

### For WakelAI Platform

**Free Tier** (Voice Assistant):
- **Resources**: 0.1 vCPU, 64MB RAM (shared)
- **Model**: tiny.en (shared server)
- **TTS**: OpenAI (cheaper)
- **Perfect for**: Testing, light usage

**Standard Tier** (Voice Assistant):
- **Resources**: 0.25 vCPU, 128MB RAM (shared)
- **Model**: small.en (shared server)
- **TTS**: ElevenLabs (better quality)
- **Perfect for**: Daily use, better accuracy

**Pro Tier** (Voice Assistant):
- **Resources**: 0.5 vCPU, 256MB RAM (dedicated)
- **Model**: small.en (dedicated instance)
- **TTS**: ElevenLabs Flash
- **Perfect for**: Heavy usage, lowest latency

---

## Implementation Recommendations

### Phase 1: Start Small (Week 1)
- Use tiny.en model
- OpenAI TTS (cheaper)
- Single shared server
- Support 10-20 concurrent users
- **Resources needed**: 1-1.5GB RAM, 1.5-2 vCPU

### Phase 2: Scale Up (Week 2-3)
- Add small.en option for premium users
- ElevenLabs integration
- Optimize shared architecture
- Support 50-100 concurrent users
- **Resources needed**: 5-7GB RAM, 7-10 vCPU

### Phase 3: Optimize (Week 4+)
- Implement lazy loading
- Add browser-native STT option
- Auto-scaling based on demand
- Support 100+ concurrent users
- **Resources needed**: 8-12GB RAM, 10-15 vCPU

---

## Quick Reference

### Minimum Resources (per user, shared architecture)
- **RAM**: 70-150MB (when active)
- **CPU**: 0.1-0.2 vCPU (when active)
- **Disk**: 500MB (one-time for model + code)

### Typical Resources (per user, dedicated instance)
- **RAM**: 400-700MB (always allocated)
- **CPU**: 0.6-1.2 vCPU (peak when processing)
- **Disk**: 500MB (one-time)

### Monthly TTS Costs (per active user)
- **OpenAI**: $15-30/month (heavy usage)
- **ElevenLabs**: $50/month (heavy usage)

---

## Conclusion

**Jarvis AI with ElevenLabs or OpenAI TTS is very lightweight:**

- **Per user**: 400-500MB RAM, 0.6-0.8 vCPU (tiny.en)
- **Shared architecture**: 70-150MB per additional user
- **TTS cost**: $15-50/month per heavy user
- **No local TTS resources** (cloud APIs handle it)

**Key advantage**: You can serve 10-20 users with <2GB RAM using shared model architecture.

---

## Sources

- [ElevenLabs API Pricing](https://elevenlabs.io/pricing/api)
- [Whisper Model Sizes Explained](https://openwhispr.com/blog/whisper-model-sizes-explained)
- [OpenAI TTS Pricing](https://openai.com/pricing)
- [faster-whisper GitHub](https://github.com/SYSTRAN/faster-whisper)
