LeadGrow · infra-orchestrator

Fleet Health Doctrine

How every inbox and domain in the fleet gets measured, tagged, ramped, hospitalized, rotated, reassigned, and retired, automatically, every night. Money follows one rule: spend goes to assets that deliver, and burnt assets come off the books before they bill again.

Approved · Aug 27

0The one-screen version

Nightly sensors One tag per inbox Automatic action Exit: recover, reassign, or retire

The system already measures health nightly (DNS auth, blocklists, landing pages, bounces, warmup scores, reply rates). This doctrine adds the layer that acts on those measurements: every inbox carries exactly one status tag, every tag has defined entry and exit rules, and every dead asset has a scheduled, billing-aware exit. Humans approve destructive steps with one click; nothing else waits on a human.

The recurring cost is inboxes, not domains. The whole doctrine exists to move inbox spend off burnt domains and onto newly warmed ones, fast, and never to pay a second month for an asset we already know is dead.

1The six states

Applied nightly in Bison and the database. Worst signal wins. One bad night is enough to demote; it takes three clean nights to promote, so nothing flip-flops. Each state's official name is a life-arc status phrase; the full vocabulary system is §12.

StateMeaningSending
🔵 in schoolInside its warmup window, not yet graduatedWarmup only, zero cold email
🟢 on the jobHealthy, in rotation, producingNormal ramp schedule
🟡 under observationOne degraded signal last night, being watchedUnchanged, no upward pushes
🔴 in the hospitalActive damage or cannot sendLimit forced to zero. Stays in its campaigns; warmup keeps running (that's the physio)
🟣 on leaveHealthy, deliberately rotated out this period, returns on scheduleZero cold, warmup maintains
⚪ on callGraduated and healthy, benched, deployable on demandZero cold, warmup maintains
⚫ in the graveyardPronounced dead; sender awaiting burial (purge from Bison)Frozen. Excluded from all automation and all health math; no exit except burial

The grammar is load-bearing: "in" = held by an institution with entry/exit rituals (enrolled→graduated, admitted→discharged), not freely deployable. "under" = suspended, deployable but frozen. "on" = free and healthy. The preposition alone tells you the tier. ⚫ is not a seventh health state but a coffin lid: entered only by pronounced, left only by burial, immune to hysteresis, invisible to every average.

2The sensors

SignalThresholdSends you to
SMTP/IMAP connectionAuth failure: the inbox physically cannot send🔴 hospital
DNS auth (SPF/DKIM/DMARC)Any of the three failing. Every send while broken deepens the damage🔴 hospital
Bounce rate>3% on meaningful volume🔴 hospital
Bounce rate1.5–3%, creeping🟡 watch
Reply rate vs own baselineDropped ≥75% week-over-week (min 30 sends both windows)🔴 hospital
Reply rate vs own baselineDropped 50–75%🟡 watch
Reply rate vs client peers (new)≥50% below the median of the client's other domains (volume gate scaled by domain class, §7)🟡 watch; at 75% below → placement test
Bounce text names a blockReceivers explicitly citing a reputation policy🔴 hospital
Blocklist listing aloneNew SURBL/DBL listing, no downstream evidence🟡 watch only

Why a listing alone is only watch: we live-tested it. Sends from SURBL + DBL-listed domains with production copy landed Gmail Primary 9 of 10 times. 238 of 448 live domains are DBL-listed today; treating listings as burn signals would torch half a working fleet. Listing plus failed placement or reply collapse upgrades to hospital.

3Ramping and the graduation gate

Ramp rules already live: upward pushes at most double the current limit, one step per 7 days per inbox, down-snaps immediate. This doctrine adds the missing interlock:

No inbox below graduation score ever receives an upward push. Found live on Aug 26: the age-based ramp had stepped inboxes with warmup scores of 10–27 up to 8 cold sends per day. Only their lack of campaign attachment prevented cold email from provably burnt domains. The ramp must read the warmup score before every push.

Graduation gate: 90+. A 90 score means 9 of 10 warmup emails are landing in the inbox, which is graduation-worthy. Below 90 is watched, never pushed, and handled by the checkpoint schedule below. There is no fixed "extension week": what earns more warming time is trajectory, not a calendar band. Azure modes set the calendar floor, not the bar: azure-fast may graduate at day 21 like everyone else, azure-aged holds until day 90; both need 90+ at release.

Which score window are we watching? Today: the wrong one. Verified live against all 606 P1 warmup inboxes: Bison's displayed warmup score is lifetime cumulative: total kept-in-inbox ÷ total ever sent since day one. It lags badly (a rough first week drags the score for months) and it cannot show direction. So the orchestrator snapshots each inbox's warmup counters nightly and computes its own rolling 3-day and 7-day scores plus the trend between them. Every gate and checkpoint in this doctrine reads the 7-day rolling score; trend means the 7-day score vs itself ~4 days earlier (up = +3 points or more, flat = within ±3, down = −3 or worse).

4Warmup checkpoints and the billing-aware cut

Context: a domain warms ~14 days before inboxes go on. Inboxes then warm ~21 days. Inbox subscriptions are prepaid monthly and locked to their domain, so an early cut buys nothing: the month is already paid. The play is to decide early, queue the replacement early, and execute the cut at the renewal boundary so a failing domain never bills twice.

The math is the argument: a day-14 decision plus a 14-day domain warm means the replacement is ready at day 28, exactly when the monthly inbox subscription renews. The queue-early rule turns the billing calendar from a leak into the natural cut schedule.

Domain problem vs inbox problem

Every inbox on the domain uniformly low → the domain is burnt (typically bought with unseen abuse history): retire the domain, replace it. One straggler among healthy siblings → replace the inbox only; the domain is fine.

5Billing awareness (new data point)

Every inbox carries a renews_at date, sourced from the provider at provision time (icemail order date, tenant billing anchor, workspace billing date). This powers:

6The hospital

Admission (any one signal)

  • Connection dead (auth failure)
  • SPF/DKIM/DMARC failing
  • Bounce >3%
  • Reply rate collapsed ≥75% from own baseline
  • Bounce texts naming a block

The stay

  • Sending limit forced to zero
  • Stays attached to its campaigns
  • Warmup keeps running (that is the rehab)
  • Blocklists and score watched nightly

Exits: trajectory decides, the cap backstops

  • Auth-type cases: fixed + 3 clean nights → out
  • 7-day rolling score trending up → keep rehabbing, up to the cap
  • Flat or down through day 14 of the stay → out early, retirement queue. Don't wait for the calendar
  • Recovered → paid placement test to confirm → re-ramp from watch
  • Outer cap 30 days regardless of trend

Skip the test when it's pointless

  • The warmup network is a free, continuous placement panel
  • A flat or falling rolling score = a failed placement test we didn't pay for
  • Straight to retirement queue; paid tests confirm recovery only

7Copy problem or domain problem: the routing rule

Reply rate is the only deliverability measure that survives contact with reality, but it mixes two causes: bad copy and bad infrastructure. Comparing a domain against the client's other domains untangles them. Not because every domain sends identical copy (it doesn't: inboxes sit in several campaigns at once and assignment is mixed), but because random mixing means every domain sends roughly the same blend of the client's campaigns over a rolling window. When nobody's blend is special, a domain that still lags the family is lagging for domain reasons:

PictureDiagnosisAction
ALL of a client's domains are getting weak replies, roughly evenlyThe copy, offer, or targeting is bad. The infrastructure is fineThese are healthy domains stuck on a dead campaign. When the client churns or the copy is chronically dead, move those good domains over to another client's campaigns (new redirect, new signature, new campaigns) instead of retiring working assets. That is all "reassignment" means
One domain doing clearly worse than its siblings on a comparable campaign blendCopy proven innocent by the siblings. The domain itself is suspectTwo-arm placement test (§10). Both arms clean → keep watching. Neutral arm spams too → hospital → retire path. A lagging domain never gets moved to another client: its reputation travels with it

The like-for-like guard

Random mixing is the control, so the sensor must verify the mix before trusting a verdict. Before any black-sheep call, check the domain's campaign blend against its siblings'. If its sends are concentrated in one campaign unlike the rest of the family, compare within campaigns instead: below its siblings inside every campaign they share = a true black sheep. Below only in the overall number = it is carrying a weak campaign, and that is a copy finding, not a domain finding. The same guard applies to the week-over-week drop detector: a domain whose campaign mix shifts into a weaker campaign can fake a deliverability drop. And no domain is ever judged on reply data alone: the two-arm test (§10) hits the domain directly and is immune to campaign mixing entirely. The reply sensors screen; the lab diagnoses.

The volume gate, scaled honestly

No fixed 1,000-send minimum: a Google/MS domain runs 2 inboxes and might not see 1,000 sends in its first two months. The gate is statistical instead: enough sends that the client's own median reply rate predicts about 5 replies. For a client whose domains median 2%, that's ~250 sends; sitting at 0–1 replies on that volume is signal, not luck. A ramped Google domain reaches it in ~2 weeks; an azure domain (25 inboxes) in ~2 days. Below the gate the verdict is "insufficient data," never a fake judgment. Google/MS domains are measured over a rolling 30 days, azure over 14, so slow senders still accumulate a real sample.

8Rotation: resting and reserve

Healthy inboxes rotate through sending and resting periods by batch, so no domain sends every week forever; the reserve bench holds warmed, healthy inboxes ready to deploy or to backfill a replacement. Rotation depends on the batch system being populated (it currently is not: 0 of 569 domains are in a batch), so 🟣 and ⚪ ship after batches exist. Everything else in this doctrine stands alone.

9The exit door: retirement

Verify dead first

Retirement requires a confirmed burnt verdict, never idleness. Idle but clean goes to reserve. Qualifying verdicts:

The mechanics

10Placement testing: buy, not build

EmailGuard Pro (~$49/mo) replaces the homemade seed-account build. Their seed bank beats anything we'd maintain, we already integrate with their API for blocklist checks, and the doctrine only ever uses placement tests sparingly, as signals: confirming a peer-outlier, confirming a hospital recovery. Trigger logic stays ours; execution is theirs. Placement tests never decide alone; they confirm.

The two-arm protocol (Mitchell's)

Every diagnostic placement test runs two arms from the same inbox: the real campaign copy, and a neutral control so short it cannot trigger a content filter (subject "tomorrow meeting", body "see you tomorrow - Bob"). The pair separates what a single test never can:

Real copyNeutral copyVerdictAction
InboxInboxDomain clean, copy deliveringIf replies are still low, it's the offer or targeting, not deliverability. No infra action
SpamInboxThe copy is fingerprinted. The domain is innocentCopy surgery: strip sections (phone number, links, calendar URL) and retest until the trigger is found, then rotate the trigger. Phone numbers are the most common. The domain never goes to hospital for this
SpamSpamThe inbox/domain is burned. Even "see you tomorrow" can't landHospital → retire path
InboxSpamNoiseRetest before concluding anything

This also splits "copy problem" into two different diseases: fingerprinted copy (filters block it, fix with copy surgery) vs weak copy (delivered fine, ignored by humans, fix with a better offer). Reply-rate data alone cannot tell them apart; the neutral arm can.

11What this catches in today's fleet

12The vocabulary

The whole system speaks one language: a life, from birth to burial. Three layers, one law: states are prepositional phrases ("is ___"), events are past-tense verbs ("was ___"), instruments are nouns that pass the stranger test: a new hire hears the term once in a night-rounds report and knows what it is, no glossary. Never mix layers, and anything added later has a rule to follow.

Events, every transition has exactly one verb

What happenedVerbWhat happenedVerb
ProvisionedbornRotated out / brought backgranted leave / recalled
Day 5–7 first readscreenedMoved to another clientrelocated
Failed a day-14/21 gateflunkedRetirement approved by a humanwithdrawn (care withdrawn)
Passed at 90+graduatedDeath confirmed by the verify-dead gatepronounced
Deployed from on callhiredInbox subscriptions cancelledsettled
Sent to observationreferredTombstone written, sender purged from Bison, domain lapsesburied
3 clean nights, releasedclearedInto / out of the hospitaladmitted / discharged

Instruments, the named things

School

Night rounds (the nightly monitor run) · the report card (rolling 7-day warmup score) vs the GPA (Bison's lifetime score, never judge by it, read the latest report card) · milestones (day 7/14/21 gates) · newborn screening · failure to thrive · growth spurts (ramp pushes, milestone-gated) · the successor (the replacement domain, named at day 14, warmed to take over at the renewal boundary)

Family & clinic

The family (a client's domains, all on the same copy) · the black sheep (the one domain lagging its siblings) · the lab (EmailGuard) · the two-arm test (real copy + neutral control, §10) · copy surgery (strip sections until the trigger is found) · prognosis (hospital trend: improving / stable / declining) · physio (warmup during the stay) · clean bill of health (3 clean nights)

End of life

Life support (the retirement queue) and the machines (its monthly carry cost) · DOA (the immediate-retire class: found dead, straight to the coroner) · the coroner (verify-dead gate) · the death certificate (cause + evidence) · the estate (a dead domain's subscriptions) · the bills (the weekly renewal forecast: who renews, when, for how much) · the tombstone, the graveyard, no resurrections

Night rounds, as it will actually read: "Night rounds complete. 12 born. In school: 41 on track, 3 failures to thrive, successors named. 2 graduated to on call. Opus family: one black sheep referred, sent to the lab. In the hospital: 3 improving, 1 flatlined, moved to life support. The bills: 27 renew Friday while on life support, decide now. Coroner pronounced 2 DOA, estates settled, tombstones written."

13The sign-off

The doctrine, condensed to eight positions. Approved by Aydan and Mitchell, Aug 27. Build order: infra-orchestrator issue #29.

  1. Six tags, worst signal wins, one bad night demotes, three clean nights promote. Hospital = limit zero, stays in campaigns.
    Decided.
  2. Graduation: 90+ on the 7-day rolling warmup score before any cold email. All checkpoints read rolling scores + trend, never Bison's lifetime number.
    Decided. Open sub-question: same 90 bar for azure-mode inboxes, or different?
  3. Warmup checkpoints: day-7 first read (under 50 = watch, never an early cut, the month is prepaid); day-14 under 80 = queue the replacement domain, 80–90 goes by trend; day-21: 90+ graduates, 80–90 by trend, under 80 replaced. Cuts execute on the eve of the inbox renewal.
    Decided. Whole domain uniformly low = retire the domain (prior abuse); one straggler = swap the inbox only.
  4. Ramp interlock: no upward limit push on any inbox below graduation score.
    Decided. Live bug today: score-17 inboxes were ramped to 8/day by age alone.
  5. Hospital: trajectory-driven exits (flat or falling rolling score through day 14 = out early to the retirement queue; rising = rehab continues), 30-day outer cap.
    Decided.
  6. Reply routing: whole client low = copy problem, good domains move to another client; one domain lagging siblings = domain problem, placement test then treat or retire, never moved. Volume gate scaled per class (~5 expected replies by the client's own median), no flat 1,000-send rule.
    Watch at 50% below the client median, act at 75% below.
  7. EmailGuard Pro (~$49/mo) instead of building seed-account placement testing. Tests confirm, never decide.
    First job after subscribing: our own Microsoft placement test from DBL-listed domains. The Gmail side is already proven (9/10 Primary); this is testable, not a matter of opinion.
  8. Retirement: verified-burnt only (never idleness), immediate-burn list for the provably dead, inbox-first cancellation at renewal eve, permanent tombstone, lapse-by-default at registrars, renews_at tracked on every inbox from provision day.
    Decided. Nightly queue, one-click human approval.