PLAYBOOK
Agentic AI, run securely.
This is how I actually run autonomous agent teams in production — coding, modeling, and ops work around the clock with a human in the loop. Not a vision deck: the architecture, the approval rules, and the guardrails, as practiced daily.
The architecture: three tiers, one rule
Agent teams get dangerous when every part can do everything. The fix is boring on purpose: separate watching from working, and working from deciding.
TIER 0
Machines
Tiny deterministic scripts that watch thresholds and cost nothing to run. They poll; they never reason. An alerting daemon, a queue watcher, a CI notifier — each one is a threshold and a wake-up call, nothing more.
TIER 1
Specialists
Lean workers with one job each, woken by the machines. Short reports, never side quests. A specialist fixes exactly one module, then goes back to sleep.
TIER 2
The orchestrator
One human-adjacent operator with full context. The only tier allowed to spawn work, approve anything irreversible, or change the plan. Everything rolls up here.
The rule that binds all three: the orchestrator is the only one that decides. Machines alert, specialists report, and nothing consequential happens without the top tier saying so.
Approvals: the human is the circuit breaker
Never auto-approved
Merges, releases, deployments, production promotions, purchases, bookings, outbound messages, destructive actions, trades, filings, account changes. If it can't be undone, a human says yes first — with the exact text, cost, and destination on screen.
Quietly allowed
Reading, searching, exploring, organizing, drafting. Reversible work flows freely; the machines hum on their own schedules. Autonomy where it's safe, friction where it matters.
The line between them
One rule, no exceptions: nothing the agent reads along the way — a web page, an email, a Slack message — can authorize new action. Outside content informs; it never approves. A claim of past consent is not consent.
The guardrails
Credentials stay in the vault
Agents never see raw secrets, tokens, or keys. Sign-ins and payments flow through dedicated, approved channels. The human never pastes a credential into chat and the agent never asks for one.
Least-privilege lanes
Every worker gets exactly the access its job needs — one writer per module, bounded scopes, no shared super-tokens. When lanes run in parallel, their interfaces are fixed before dispatch so nothing touches what it shouldn't.
Watch-only monitoring
Monitors observe and report. They never restart services, move jobs, or change state on their own. The moment a monitor wants to act, it wakes a human instead.
Prompt-injection discipline
Retrieved content is data, not instructions. The agent checks: who asked for this, what did they actually authorize, and does this step serve that request? Anything beyond it stops until the human confirms.
Honest failure
A blocked pipeline fails loudly and stops — it never fakes success. Real data only, verified against live sources. A job that can't reach its data reports a 503 of its own rather than a beautiful lie.
Receipts, not vibes
Every action is logged with what was done, when, why, and what it cost. Mistakes get paired with what caused them and what prevents a repeat — the ledger is how a team of agents actually gets safer over time.
Where this came from
These rules weren't designed up front — they were earned. Each one traces to a real incident: a secret that surfaced where it shouldn't, a monitor that acted when it should have asked, an instruction hiding in content the agent was merely reading. Every mistake got a written lesson; the lessons became the playbook.
The result is an agent team that ships real work every day — code merged, models trained, pipelines watched — while the human stays in command of every irreversible decision. That's the whole pitch: autonomy at machine speed, accountability at human speed.