# Beyond SaaS: Introducing AOA (Agent-Offloaded Architecture)
*How to build zero-backend, privacy-first AI production platforms with static frontends and the coding agents your users already have.*

---

## The AI Compute Tax Problem

If you've built or used generative AI applications recently, you know the formula:
1. User types a prompt into a slick web interface.
2. A cloud server spins up expensive GPUs to generate media/code.
3. The platform charges the user a recurring $20–$50/month subscription (or token fees) to subsidize massive cloud bills.
4. The user's code, voice, or sensitive assets live in someone else's cloud.

We call this the **"AI Compute Tax"**. But look around your development machine:
- You have a machine that easily handles scraping, transcoding, and rendering — and if it's beefy enough, even local LLM / TTS models.
- You have battle-tested CLI tools: `ffmpeg`, `playwright`, local TTS engines.
- And increasingly, you have an AI **Coding Agent** (Claude Code, Codex, Cursor, Cline, Aider, Pi) with direct filesystem access — one you're already paying for.

**Why are we still running cloud backends for generative tasks when the user already has a world-class agent wired into their own machine?**

---

## Enter AOA: The Brain and the Engine

Today, we want to formalize a new architectural paradigm: **AOA (Agent-Offloaded Architecture)**.

The core idea is simple. Every AI product needs two things:
> **A brain — LLM inference. And an engine — the tokens that inference consumes.**

Cloud SaaS makes the vendor supply both and bill them back as subscriptions. In AOA, both belong to the agent the user already has. The application doesn't provide compute, API keys, or databases. It provides:
1. **A Living Protocol**: Strict JSON Schemas and workflow definitions.
2. **A Zero-Backend Workbench**: A purely static web UI (hosted for pennies or completely free on GitHub Pages).
3. **A State Visualizer**: A visual lens that looks directly into the local directory via the browser's **File System Access API**.

The heavy lifting—reasoning, web analysis, prompt iteration, narration generation, asset synthesis, and video rendering—is **offloaded to the user's own Coding Agent**:
- **Execution** (scraping, transcoding, compositing, rendering) runs on the user's machine.
- **Inference** runs wherever the user's agent runs it: by default on their agent provider's cloud, billed to the plan they already have — or, if they choose, fully offline on self-hosted models (Pi-style agents pointed at Ollama / llama.cpp, Piper / Kokoro for TTS).

Either way, the app developer pays nothing for inference, and the user's files never pass through the app developer's servers.

The application doesn't even need a frontend. A service can expose only an API, publish its specification at a URL (OpenAPI, a guide, or a skill), and let the user's agent call it directly. Whatever the shape, the agent's first step is the same: fetch the specification from a URL.

---

## How It Works in Practice

Here is the life of an AOA application:

```
[ Static Web UI (GitHub Pages) ]
      │
      │ 1. Pick Folder (File System Access API)
      │ 2. Write initial parameters (video.start.json)
      │
      ▼
[ Local Filesystem Directory (SSOT) ]
      ▲
      │ 3. Agent reads task instructions & protocol schemas
      │ 4. Agent runs the pipeline (e.g., Playwright + TTS + FFmpeg)
      │ 5. Agent updates activity status (video.activity.json)
      │
[ Coding Agent on the user's machine (e.g. Claude Code) ]
      │
      └── LLM inference: agent provider's cloud (default) or local model (optional)
```

### 1. Zero-Backend Presentation
The frontend is pure HTML/CSS/Vue/React. It contains zero LLM weights, zero WebGPU bloat, and zero backend servers. It's served straight from a CDN — no server cold starts.

### 2. File System Access as the Bus
Using `window.showDirectoryPicker()`, the user grants the static web app access to an empty working directory. From that point forward, the browser communicates with the agent **not through cloud HTTP requests, but through the local filesystem**.

### 3. Lightweight Fingerprint Polling
How does the static frontend know when the agent finishes generating a scene or compiling a video?
Every 2 seconds, the UI computes a cheap fingerprint from key files' timestamps (`lastModified`) and sizes — no file contents are read. When the agent produces an asset or writes an updated `scene.json`, the fingerprint changes and the UI reloads reactively.

### 4. Schema as the Contract
There is no ambiguity between the UI and the Agent. A strict `protocol.schema.json` dictates valid state transitions, mandatory fields, and asset formats. Both the frontend validator and the Agent conform to the same contract — and since agents are not deterministic, the frontend validates everything it reads back.

---

## Does AOA Require a Pure Static Frontend? Not At All!

While our reference implementation is a zero-backend static SPA, **AOA applies just as powerfully to applications with real backends**:

### The "Decoupled Control Plane" Paradigm
In traditional enterprise SaaS, the backend handles both **Control** (Auth, Billing, Team Sync) and **Compute** (Inference, Rendering, Heavy ETL). This is why scaling AI apps is financially exhausting.

Under **Backend-Enabled AOA**:
- **The Cloud Backend** serves as a thin Control Plane: authentication, team permissions, organization templates, and billing.
- **The User's Agent** handles the Compute Plane: private source code, internal databases, and heavy media processing stay on the employee's machine, and inference goes to the enterprise's chosen LLM provider (e.g., under an enterprise data agreement) or an in-house model.
- Only lightweight metadata, template definitions, and final sanitized exports flow to the cloud backend.

This gives SaaS builders **near-zero marginal compute costs** and gives enterprises **control over their own data boundary (BYOA - Bring Your Own Agent)**.

---

## Why AOA Matters

1. **For Indie Hackers & Creators**:
   - Zero infrastructure overhead. You can launch complex generative AI tools that scale to 100,000 users without paying a single dollar in cloud GPU bills — each user brings their own agent plan.
2. **For Enterprises & Privacy-Conscious Users**:
   - Data sovereignty by choice. Proprietary source code, company secrets, and unreleased designs never pass through the app vendor's servers. The only third party that sees context is the LLM provider *you* chose — and with local models, there is none.
3. **For Developers**:
   - Inspectable and hackable. Because all intermediate assets (scripts, JSON states, partial audio files) sit in the user's project folder, anyone can modify or debug any frame with standard tools.

### The Honest Trade-offs
- File System Access API (`showDirectoryPicker`) currently works only in desktop Chrome / Edge and other Chromium browsers; Brave disables it by default, and Firefox and Safari don't support it.
- Users need a working agent setup — and pay for it themselves.
- In pure static mode, the user still triggers the agent by pasting a command into their terminal (an optional local companion can remove that step).

---

## Read the Full Specification

We have detailed the complete architectural blueprints, communication patterns, and security considerations.

Check out the full documentation at:
- Architecture Whitepaper: [docs/architecture.md](../docs/architecture.md)
- Reference Implementation: [Agent Video Producer](https://github.com/tigernaxojr/index-url-director)

We invite the community to explore, implement, and challenge the AOA pattern. The future of AI software is not in centralizing compute—it is in unleashing the agent your users already have.
