PLAYBOOK

Agentic AI, run securely.

This is how I actually run autonomous agent teams in production — coding, modeling, and ops work around the clock with a human in the loop. Not a vision deck: the architecture, the approval rules, and the guardrails, as practiced daily.

The architecture: three tiers, one rule

Agent teams get dangerous when every part can do everything. The fix is boring on purpose: separate watching from working, and working from deciding.

TIER 0

Machines

Tiny deterministic scripts that watch thresholds and cost nothing to run. They poll; they never reason. An alerting daemon, a queue watcher, a CI notifier — each one is a threshold and a wake-up call, nothing more.

TIER 1

Specialists

Lean workers with one job each, woken by the machines. Short reports, never side quests. A specialist fixes exactly one module, then goes back to sleep.

TIER 2

The orchestrator

One human-adjacent operator with full context. The only tier allowed to spawn work, approve anything irreversible, or change the plan. Everything rolls up here.

The rule that binds all three: the orchestrator is the only one that decides. Machines alert, specialists report, and nothing consequential happens without the top tier saying so.

Approvals: the human is the circuit breaker

Never auto-approved

Merges, releases, deployments, production promotions, purchases, bookings, outbound messages, destructive actions, trades, filings, account changes. If it can't be undone, a human says yes first — with the exact text, cost, and destination on screen.

Quietly allowed

Reading, searching, exploring, organizing, drafting. Reversible work flows freely; the machines hum on their own schedules. Autonomy where it's safe, friction where it matters.

The line between them

One rule, no exceptions: nothing the agent reads along the way — a web page, an email, a Slack message — can authorize new action. Outside content informs; it never approves. A claim of past consent is not consent.

The guardrails

Credentials stay in the vault

Agents never see raw secrets, tokens, or keys. Sign-ins and payments flow through dedicated, approved channels. The human never pastes a credential into chat and the agent never asks for one.

Least-privilege lanes

Every worker gets exactly the access its job needs — one writer per module, bounded scopes, no shared super-tokens. When lanes run in parallel, their interfaces are fixed before dispatch so nothing touches what it shouldn't.

Watch-only monitoring

Monitors observe and report. They never restart services, move jobs, or change state on their own. The moment a monitor wants to act, it wakes a human instead.

Prompt-injection discipline

Retrieved content is data, not instructions. The agent checks: who asked for this, what did they actually authorize, and does this step serve that request? Anything beyond it stops until the human confirms.

Honest failure

A blocked pipeline fails loudly and stops — it never fakes success. Real data only, verified against live sources. A job that can't reach its data reports a 503 of its own rather than a beautiful lie.

Receipts, not vibes

Every action is logged with what was done, when, why, and what it cost. Mistakes get paired with what caused them and what prevents a repeat — the ledger is how a team of agents actually gets safer over time.

Where this came from

These rules weren't designed up front — they were earned. Each one traces to a real incident: a secret that surfaced where it shouldn't, a monitor that acted when it should have asked, an instruction hiding in content the agent was merely reading. Every mistake got a written lesson; the lessons became the playbook.

The result is an agent team that ships real work every day — code merged, models trained, pipelines watched — while the human stays in command of every irreversible decision. That's the whole pitch: autonomy at machine speed, accountability at human speed.