How every inbox and domain in the fleet gets measured, tagged, ramped, hospitalized, rotated, reassigned, and retired, automatically, every night. Money follows one rule: spend goes to assets that deliver, and burnt assets come off the books before they bill again.
Draft for Mitchell's sign-off · Aug 26The system already measures health nightly (DNS auth, blocklists, landing pages, bounces, warmup scores, reply rates). This doctrine adds the layer that acts on those measurements: every inbox carries exactly one status tag, every tag has defined entry and exit rules, and every dead asset has a scheduled, billing-aware exit. Humans approve destructive steps with one click; nothing else waits on a human.
The recurring cost is inboxes, not domains. The whole doctrine exists to move inbox spend off burnt domains and onto newly warmed ones, fast, and never to pay a second month for an asset we already know is dead.
Applied nightly in Bison and the database. Worst signal wins. One bad night is enough to demote; it takes three clean nights to promote, so nothing flip-flops.
| Tag | Meaning | Sending |
|---|---|---|
| 🟢 sending | Healthy, in rotation | Normal ramp schedule |
| 🟡 watch | One degraded signal last night, under observation | Unchanged, no upward pushes |
| 🔴 hospital | Active damage or cannot send | Limit forced to zero. Stays in its campaigns; warmup keeps running |
| 🔵 warming | Inside its warmup window, not yet graduated | Warmup only, zero cold email |
| 🟣 resting | Healthy, deliberately rotated out this period | Zero cold, warmup maintains |
| ⚪ reserve | Warmed and healthy, benched, awaiting deployment | Zero cold, warmup maintains |
| Signal | Threshold | Sends you to |
|---|---|---|
| SMTP/IMAP connection | Auth failure: the inbox physically cannot send | 🔴 hospital |
| DNS auth (SPF/DKIM/DMARC) | Any of the three failing. Every send while broken deepens the damage | 🔴 hospital |
| Bounce rate | >3% on meaningful volume | 🔴 hospital |
| Bounce rate | 1.5–3%, creeping | 🟡 watch |
| Reply rate vs own baseline | Dropped ≥75% week-over-week (min 30 sends both windows) | 🔴 hospital |
| Reply rate vs own baseline | Dropped 50–75% | 🟡 watch |
| Reply rate vs client peers (new) | ≥50% below the median of the client's other domains (volume gate scaled by domain class, §7) | 🟡 watch; at 75% below → placement test |
| Bounce text names a block | Receivers explicitly citing a reputation policy | 🔴 hospital |
| Blocklist listing alone | New SURBL/DBL listing, no downstream evidence | 🟡 watch only |
Why a listing alone is only watch: we live-tested it. Sends from SURBL + DBL-listed domains with production copy landed Gmail Primary 9 of 10 times. 238 of 448 live domains are DBL-listed today; treating listings as burn signals would torch half a working fleet. Listing plus failed placement or reply collapse upgrades to hospital.
Ramp rules already live: upward pushes at most double the current limit, one step per 7 days per inbox, down-snaps immediate. This doctrine adds the missing interlock:
No inbox below graduation score ever receives an upward push. Found live on Aug 26: the age-based ramp had stepped inboxes with warmup scores of 10–27 up to 8 cold sends per day. Only their lack of campaign attachment prevented cold email from provably burnt domains. The ramp must read the warmup score before every push.
Graduation gate: 90+. A 90 score means 9 of 10 warmup emails are landing in the inbox, which is graduation-worthy. Below 90 is watched, never pushed, and handled by the checkpoint schedule below. There is no fixed "extension week": what earns more warming time is trajectory, not a calendar band.
Which score window are we watching? Today: the wrong one. Verified live against all 606 P1 warmup inboxes: Bison's displayed warmup score is lifetime cumulative: total kept-in-inbox ÷ total ever sent since day one. It lags badly (a rough first week drags the score for months) and it cannot show direction. So the orchestrator snapshots each inbox's warmup counters nightly and computes its own rolling 3-day and 7-day scores plus the trend between them. Every gate and checkpoint in this doctrine reads the 7-day rolling score; trend means the 7-day score vs itself ~4 days earlier (up = +3 points or more, flat = within ±3, down = −3 or worse).
Context: a domain warms ~14 days before inboxes go on. Inboxes then warm ~21 days. Inbox subscriptions are prepaid monthly and locked to their domain, so an early cut buys nothing: the month is already paid. The play is to decide early, queue the replacement early, and execute the cut at the renewal boundary so a failing domain never bills twice.
Inboxes provisioned on the warmed domain. Warmup starts. Billing clock starts: renewal date recorded per inbox.
Score means nothing under ~20 warmup sends. First meaningful read lands here. Under 50 → 🟡 watch. No cut: the month is paid, warming continues either way, and low starts sometimes recover.
7-day rolling score under 80 → queue the replacement now. Buy or pull the new domain and start its 14-day domain warming immediately; the failing domain keeps warming on the already-paid month while its successor gets ready. 80–90 → trend check: trending up, keep warming; flat or down, queue the replacement anyway. Direction tells you where this is heading before the calendar does.
90+ → graduate, cold email begins. 80–90 → trend decides: trending up, keep warming toward the renewal boundary; flat or down, it's replaced. Under 80 → replaced, no trend check needed; the replacement should already be mid-warm.
Replacement domain finishes its 14-day warm right as the failing domain's inbox month ends. Cancel the old inboxes on the eve of renewal. Zero second-month spend, zero coverage gap. Domain gets a death certificate (cause: probable prior abuse) and lapses at the registrar.
The math is the argument: a day-14 decision plus a 14-day domain warm means the replacement is ready at day 28, exactly when the monthly inbox subscription renews. The queue-early rule turns the billing calendar from a leak into the natural cut schedule.
Every inbox on the domain uniformly low → the domain is burnt (typically bought with unseen abuse history): retire the domain, replace it. One straggler among healthy siblings → replace the inbox only; the domain is fine.
Every inbox carries a renews_at date, sourced from the provider at provision time (icemail order date, tenant billing anchor, workspace billing date). This powers:
Reply rate is the only deliverability measure that survives contact with reality, but it mixes two causes: bad copy and bad infrastructure. Comparing a domain against the other domains sending the same client's copy untangles them, because the copy is identical across all of them:
| Picture | Diagnosis | Action |
|---|---|---|
| ALL of a client's domains are getting weak replies, roughly evenly | The copy, offer, or targeting is bad. The infrastructure is fine | These are healthy domains stuck on a dead campaign. When the client churns or the copy is chronically dead, move those good domains over to another client's campaigns (new redirect, new signature, new campaigns) instead of retiring working assets. That is all "reassignment" means |
| One domain doing clearly worse than its siblings on the same copy | Copy proven innocent by the siblings. The domain itself is suspect | Two-arm placement test (§10). Both arms clean → keep watching. Neutral arm spams too → hospital → retire path. A lagging domain never gets moved to another client: its reputation travels with it |
No fixed 1,000-send minimum: a Google/MS domain runs 2 inboxes and might not see 1,000 sends in its first two months. The gate is statistical instead: enough sends that the client's own median reply rate predicts about 5 replies. For a client whose domains median 2%, that's ~250 sends; sitting at 0–1 replies on that volume is signal, not luck. A ramped Google domain reaches it in ~2 weeks; an azure domain (25 inboxes) in ~2 days. Below the gate the verdict is "insufficient data," never a fake judgment. Google/MS domains are measured over a rolling 30 days, azure over 14, so slow senders still accumulate a real sample.
Healthy inboxes rotate through sending and resting periods by batch, so no domain sends every week forever; the reserve bench holds warmed, healthy inboxes ready to deploy or to backfill a replacement. Rotation depends on the batch system being populated (it currently is not: 0 of 569 domains are in a batch), so 🟣 and ⚪ ship after batches exist. Everything else in this doctrine stands alone.
Retirement requires a confirmed burnt verdict, never idleness. Idle but clean goes to reserve. Qualifying verdicts:
EmailGuard Pro (~$49/mo) replaces the homemade seed-account build. Their seed bank beats anything we'd maintain, we already integrate with their API for blocklist checks, and the doctrine only ever uses placement tests sparingly, as signals: confirming a peer-outlier, confirming a hospital recovery. Trigger logic stays ours; execution is theirs. Placement tests never decide alone; they confirm.
Every diagnostic placement test runs two arms from the same inbox: the real campaign copy, and a neutral control so short it cannot trigger a content filter (subject "tomorrow meeting", body "see you tomorrow - Bob"). The pair separates what a single test never can:
| Real copy | Neutral copy | Verdict | Action |
|---|---|---|---|
| Inbox | Inbox | Domain clean, copy delivering | If replies are still low, it's the offer or targeting, not deliverability. No infra action |
| Spam | Inbox | The copy is fingerprinted. The domain is innocent | Copy surgery: strip sections (phone number, links, calendar URL) and retest until the trigger is found, then rotate the trigger. Phone numbers are the most common. The domain never goes to hospital for this |
| Spam | Spam | The inbox/domain is burned. Even "see you tomorrow" can't land | Hospital → retire path |
| Inbox | Spam | Noise | Retest before concluding anything |
This also splits "copy problem" into two different diseases: fingerprinted copy (filters block it, fix with copy surgery) vs weak copy (delivered fine, ignored by humans, fix with a better offer). Reply-rate data alone cannot tell them apart; the neutral arm can.
The doctrine, condensed to eight positions. Most are decided; a thumbs-up on the set and Charles builds. Anything Mitchell would veto from deliverability experience, flag by number.