LeadGrow · infra-orchestrator

Fleet Health Doctrine

How every inbox and domain in the fleet gets measured, tagged, ramped, hospitalized, rotated, reassigned, and retired, automatically, every night. Money follows one rule: spend goes to assets that deliver, and burnt assets come off the books before they bill again.

Draft for Mitchell's sign-off · Aug 26

0The one-screen version

Nightly sensors One tag per inbox Automatic action Exit: recover, reassign, or retire

The system already measures health nightly (DNS auth, blocklists, landing pages, bounces, warmup scores, reply rates). This doctrine adds the layer that acts on those measurements: every inbox carries exactly one status tag, every tag has defined entry and exit rules, and every dead asset has a scheduled, billing-aware exit. Humans approve destructive steps with one click; nothing else waits on a human.

The recurring cost is inboxes, not domains. The whole doctrine exists to move inbox spend off burnt domains and onto newly warmed ones, fast, and never to pay a second month for an asset we already know is dead.

1The six tags

Applied nightly in Bison and the database. Worst signal wins. One bad night is enough to demote; it takes three clean nights to promote, so nothing flip-flops.

TagMeaningSending
🟢 sendingHealthy, in rotationNormal ramp schedule
🟡 watchOne degraded signal last night, under observationUnchanged, no upward pushes
🔴 hospitalActive damage or cannot sendLimit forced to zero. Stays in its campaigns; warmup keeps running
🔵 warmingInside its warmup window, not yet graduatedWarmup only, zero cold email
🟣 restingHealthy, deliberately rotated out this periodZero cold, warmup maintains
⚪ reserveWarmed and healthy, benched, awaiting deploymentZero cold, warmup maintains

2The sensors

SignalThresholdSends you to
SMTP/IMAP connectionAuth failure: the inbox physically cannot send🔴 hospital
DNS auth (SPF/DKIM/DMARC)Any of the three failing. Every send while broken deepens the damage🔴 hospital
Bounce rate>3% on meaningful volume🔴 hospital
Bounce rate1.5–3%, creeping🟡 watch
Reply rate vs own baselineDropped ≥75% week-over-week (min 30 sends both windows)🔴 hospital
Reply rate vs own baselineDropped 50–75%🟡 watch
Reply rate vs client peers (new)≥50% below the median of the client's other domains (volume gate scaled by domain class, §7)🟡 watch; at 75% below → placement test
Bounce text names a blockReceivers explicitly citing a reputation policy🔴 hospital
Blocklist listing aloneNew SURBL/DBL listing, no downstream evidence🟡 watch only

Why a listing alone is only watch: we live-tested it. Sends from SURBL + DBL-listed domains with production copy landed Gmail Primary 9 of 10 times. 238 of 448 live domains are DBL-listed today; treating listings as burn signals would torch half a working fleet. Listing plus failed placement or reply collapse upgrades to hospital.

3Ramping and the graduation gate

Ramp rules already live: upward pushes at most double the current limit, one step per 7 days per inbox, down-snaps immediate. This doctrine adds the missing interlock:

No inbox below graduation score ever receives an upward push. Found live on Aug 26: the age-based ramp had stepped inboxes with warmup scores of 10–27 up to 8 cold sends per day. Only their lack of campaign attachment prevented cold email from provably burnt domains. The ramp must read the warmup score before every push.

Graduation gate: 90+. A 90 score means 9 of 10 warmup emails are landing in the inbox, which is graduation-worthy. Below 90 is watched, never pushed, and handled by the checkpoint schedule below. There is no fixed "extension week": what earns more warming time is trajectory, not a calendar band.

Which score window are we watching? Today: the wrong one. Verified live against all 606 P1 warmup inboxes: Bison's displayed warmup score is lifetime cumulative: total kept-in-inbox ÷ total ever sent since day one. It lags badly (a rough first week drags the score for months) and it cannot show direction. So the orchestrator snapshots each inbox's warmup counters nightly and computes its own rolling 3-day and 7-day scores plus the trend between them. Every gate and checkpoint in this doctrine reads the 7-day rolling score; trend means the 7-day score vs itself ~4 days earlier (up = +3 points or more, flat = within ±3, down = −3 or worse).

4Warmup checkpoints and the billing-aware cut

Context: a domain warms ~14 days before inboxes go on. Inboxes then warm ~21 days. Inbox subscriptions are prepaid monthly and locked to their domain, so an early cut buys nothing: the month is already paid. The play is to decide early, queue the replacement early, and execute the cut at the renewal boundary so a failing domain never bills twice.

The math is the argument: a day-14 decision plus a 14-day domain warm means the replacement is ready at day 28, exactly when the monthly inbox subscription renews. The queue-early rule turns the billing calendar from a leak into the natural cut schedule.

Domain problem vs inbox problem

Every inbox on the domain uniformly low → the domain is burnt (typically bought with unseen abuse history): retire the domain, replace it. One straggler among healthy siblings → replace the inbox only; the domain is fine.

5Billing awareness (new data point)

Every inbox carries a renews_at date, sourced from the provider at provision time (icemail order date, tenant billing anchor, workspace billing date). This powers:

6The hospital

Admission (any one signal)

  • Connection dead (auth failure)
  • SPF/DKIM/DMARC failing
  • Bounce >3%
  • Reply rate collapsed ≥75% from own baseline
  • Bounce texts naming a block

The stay

  • Sending limit forced to zero
  • Stays attached to its campaigns
  • Warmup keeps running (that is the rehab)
  • Blocklists and score watched nightly

Exits: trajectory decides, the cap backstops

  • Auth-type cases: fixed + 3 clean nights → out
  • 7-day rolling score trending up → keep rehabbing, up to the cap
  • Flat or down through day 14 of the stay → out early, retirement queue. Don't wait for the calendar
  • Recovered → paid placement test to confirm → re-ramp from watch
  • Outer cap 30 days regardless of trend

Skip the test when it's pointless

  • The warmup network is a free, continuous placement panel
  • A flat or falling rolling score = a failed placement test we didn't pay for
  • Straight to retirement queue; paid tests confirm recovery only

7Copy problem or domain problem: the routing rule

Reply rate is the only deliverability measure that survives contact with reality, but it mixes two causes: bad copy and bad infrastructure. Comparing a domain against the other domains sending the same client's copy untangles them, because the copy is identical across all of them:

PictureDiagnosisAction
ALL of a client's domains are getting weak replies, roughly evenlyThe copy, offer, or targeting is bad. The infrastructure is fineThese are healthy domains stuck on a dead campaign. When the client churns or the copy is chronically dead, move those good domains over to another client's campaigns (new redirect, new signature, new campaigns) instead of retiring working assets. That is all "reassignment" means
One domain doing clearly worse than its siblings on the same copyCopy proven innocent by the siblings. The domain itself is suspectTwo-arm placement test (§10). Both arms clean → keep watching. Neutral arm spams too → hospital → retire path. A lagging domain never gets moved to another client: its reputation travels with it

The volume gate, scaled honestly

No fixed 1,000-send minimum: a Google/MS domain runs 2 inboxes and might not see 1,000 sends in its first two months. The gate is statistical instead: enough sends that the client's own median reply rate predicts about 5 replies. For a client whose domains median 2%, that's ~250 sends; sitting at 0–1 replies on that volume is signal, not luck. A ramped Google domain reaches it in ~2 weeks; an azure domain (25 inboxes) in ~2 days. Below the gate the verdict is "insufficient data," never a fake judgment. Google/MS domains are measured over a rolling 30 days, azure over 14, so slow senders still accumulate a real sample.

8Rotation: resting and reserve

Healthy inboxes rotate through sending and resting periods by batch, so no domain sends every week forever; the reserve bench holds warmed, healthy inboxes ready to deploy or to backfill a replacement. Rotation depends on the batch system being populated (it currently is not: 0 of 569 domains are in a batch), so 🟣 and ⚪ ship after batches exist. Everything else in this doctrine stands alone.

9The exit door: retirement

Verify dead first

Retirement requires a confirmed burnt verdict, never idleness. Idle but clean goes to reserve. Qualifying verdicts:

The mechanics

10Placement testing: buy, not build

EmailGuard Pro (~$49/mo) replaces the homemade seed-account build. Their seed bank beats anything we'd maintain, we already integrate with their API for blocklist checks, and the doctrine only ever uses placement tests sparingly, as signals: confirming a peer-outlier, confirming a hospital recovery. Trigger logic stays ours; execution is theirs. Placement tests never decide alone; they confirm.

The two-arm protocol (Mitchell's)

Every diagnostic placement test runs two arms from the same inbox: the real campaign copy, and a neutral control so short it cannot trigger a content filter (subject "tomorrow meeting", body "see you tomorrow - Bob"). The pair separates what a single test never can:

Real copyNeutral copyVerdictAction
InboxInboxDomain clean, copy deliveringIf replies are still low, it's the offer or targeting, not deliverability. No infra action
SpamInboxThe copy is fingerprinted. The domain is innocentCopy surgery: strip sections (phone number, links, calendar URL) and retest until the trigger is found, then rotate the trigger. Phone numbers are the most common. The domain never goes to hospital for this
SpamSpamThe inbox/domain is burned. Even "see you tomorrow" can't landHospital → retire path
InboxSpamNoiseRetest before concluding anything

This also splits "copy problem" into two different diseases: fingerprinted copy (filters block it, fix with copy surgery) vs weak copy (delivered fine, ignored by humans, fix with a better offer). Reply-rate data alone cannot tell them apart; the neutral arm can.

11What this catches in today's fleet

12The sign-off

The doctrine, condensed to eight positions. Most are decided; a thumbs-up on the set and Charles builds. Anything Mitchell would veto from deliverability experience, flag by number.

  1. Six tags, worst signal wins, one bad night demotes, three clean nights promote. Hospital = limit zero, stays in campaigns.
    Decided.
  2. Graduation: 90+ on the 7-day rolling warmup score before any cold email. All checkpoints read rolling scores + trend, never Bison's lifetime number.
    Decided. Open sub-question: same 90 bar for azure-mode inboxes, or different?
  3. Warmup checkpoints: day-7 first read (under 50 = watch, never an early cut, the month is prepaid); day-14 under 80 = queue the replacement domain, 80–90 goes by trend; day-21: 90+ graduates, 80–90 by trend, under 80 replaced. Cuts execute on the eve of the inbox renewal.
    Decided. Whole domain uniformly low = retire the domain (prior abuse); one straggler = swap the inbox only.
  4. Ramp interlock: no upward limit push on any inbox below graduation score.
    Decided. Live bug today: score-17 inboxes were ramped to 8/day by age alone.
  5. Hospital: trajectory-driven exits (flat or falling rolling score through day 14 = out early to the retirement queue; rising = rehab continues), 30-day outer cap.
    Decided.
  6. Reply routing: whole client low = copy problem, good domains move to another client; one domain lagging siblings = domain problem, placement test then treat or retire, never moved. Volume gate scaled per class (~5 expected replies by the client's own median), no flat 1,000-send rule.
    Watch at 50% below the client median, act at 75% below.
  7. EmailGuard Pro (~$49/mo) instead of building seed-account placement testing. Tests confirm, never decide.
    First job after subscribing: our own Microsoft placement test from DBL-listed domains. The Gmail side is already proven (9/10 Primary); this is testable, not a matter of opinion.
  8. Retirement: verified-burnt only (never idleness), immediate-burn list for the provably dead, inbox-first cancellation at renewal eve, permanent tombstone, lapse-by-default at registrars, renews_at tracked on every inbox from provision day.
    Decided. Nightly queue, one-click human approval.