A modern guide to building responsible and effective outbound marketing systems with Claude Code
A guide to making people superhuman at outbound, not automating them out of a job. 65 research reports, 209 practitioner videos, 604 claims checked by hand. Twenty-two chapters, three priced stacks, and 37 actions in order. Every number graded, dated and sourced.
This is a guide to making people superhuman at outbound, not automating them out of a job. Every agent in it extends an operator's reach, and nothing reaches a buyer without passing through that operator's judgment. The limit is not headcount; it is how many agents one person can keep doing real work every day, rather than installed and forgotten.
Built from 65 commissioned deep-research reports and 209 recorded videos by working practitioners, sifted from 1,821 candidates and current as of August 2026. Fifty-one reports cover what practitioners do and what it produces for them. The other fourteen, commissioned in August 2026, cover what the tools themselves can do, which is why the second half names components and prices. A "builds" count means true walkthroughs, the builder's own system with tools on screen: thirty-two show enough to judge, fourteen name every tool in the chain. Every number carries a grade for how far to trust it, a date, and a named source reachable in one step; numbers missing any were cut.
What this guide is, and how to read it
This guide has two halves: Part One the strategy, Part Two the implementation, with components, controls, costs, and three priced stacks to copy. Eleven decisions immediately below carry the short version, so you can disagree with the argument before spending an hour on the evidence. Part One documents the volume machine failing in public: a hundred domains bought, warmed, and sprayed across a whole addressable market inside a quarter. The replacement is a narrow list nobody else can build, an offer worth answering, and a few hundred researched, human-read messages a month. The system is an exoskeleton: the person makes every judgment, and the machine multiplies the speed and reach of each one. If you want reach without research, or a system that runs while nobody watches, Chapter II says plainly this is the wrong document.
Eleven decisions carry the whole argument.
Every one of these eleven contradicts something the outbound industry currently sells as best practice; the evidence for each comes later. Three closing sections follow the eleven moves: what one campaign actually returns, which builds did not survive, and why the responsible build is also the profitable one. One definition first: in this guide a campaign means one list paired with one offer, worked for a period of weeks. It is the unit everything here is counted in.
1 Sort tactics by what survives everyone copying them.
Before you spend on any tactic, ask two things: will it survive everyone else doing it, and do you control what it depends on? Nick Saraev, who sells automation and cold-email education, says his highest-reply campaigns all ran on conference attendee lists bought or requested from the organizers. Those contacts are not on the internet at all, and the ask is easy to try, impossible to mass-produce, and unmeasured. Saraev's claims are self-reported and he argues both sides, so Chapter VII weighs his list ranking against his own course's copy advice. Now compare either claim to a 300-mailbox setup that needs four to eight weeks of warmup before a single live send. Or to Growthly, whose setup and backups died in launch week because every domain shared the mass-blacklisted .info ending.
2 Pick a list nobody else can build.
One agency reviewed its own 214 campaigns across roughly 250,000 prospects. Meetings per prospect varied about a hundredfold depending on which list you sent to (RevPack, self-published July 2026). Copy changes inside a single list move results by tens of percent. That includes a +28% lift from a compelling offer measured across 85 million emails at Gong, the sales-conversation platform. A hundredfold beats tens of percent, so spend your time on the list.
3 Test your offer, not your wording.
Changing what you offer can move results two to three times over, a swing big enough to see at a few hundred sends a month. Proving one wording beats another by 50% on a 5% reply rate takes 1,469 sends of each version, about 2,938 in all. That is roughly a year of your entire volume for one clean comparison. And that 5% is the best case, the rate strong researched campaigns reach, while high-volume agency data comes in far lower.
4 Start publishing under a real name on day one.
Audiences take time to build and a content flywheel takes time to gather momentum, so the right time to start is now. Publishing under a named human is the one asset here a late start can never buy back.
5 Write the message yourself and cap the AI steps at one.
No mailbox provider can detect whether a machine wrote your email, because what filters catch is sameness, not authorship. AI output nobody reviewed is templated by construction, so every generative step you leave unread makes your emails look more like everyone else's. Kyle Norton, the CRO of Owner.com, walks through five such steps running back to back and calls the result slop. Use as few AI steps as possible before a human reviews the output. One is the right number.
6 Own memory and judgment, rent execution.
Own four things: your suppression list of who must not be contacted, the record of what you already did, and the offer library. The fourth is the check that a meeting actually happened. Rent the rest, meaning contact data, enrichment or filling in missing facts, address verification and sending. Vendors in those categories change products and prices monthly, and a sending mistake damages a reputation that Google and Microsoft keep, not you.
7 A human approves every message before it sends.
Every message needs a human in the loop: a person who approves it before it sends. The field almost never builds that on purpose. Across every recorded build, only three kinds of gate appear in the send path: a billing prompt, a rate limit, and an approval step. The approval step disappears the moment the job gets scheduled, and none of the three was designed as a control. So stop writing rules into prompts, which a model can ignore, and put the rule and the sending credential itself into code the model cannot touch.
8 Judge the program on meetings held.
Count replies too, because sorting them by reason is your best early diagnostic. Then rest the verdict on the numbers that pay: meetings that appear in a calendar and actually happen, and what they convert into afterward.
9 Start the clock at domain purchase.
How long you wait before your first email depends on which sending build you chose, and the answers are weeks apart. A fleet of hundreds of mailboxes needs four to eight weeks of warmup, putting the honest decision-to-pipeline gap at 110 to 130 days. One or two mailboxes on a domain you own need only the one to three days DNS propagation takes, and researched sending is itself the ramp. Either way the clock starts at the domain purchase, not at your first email.
10 Count the agents that do real work.
Count the agents that do real work every day, not the ones you have installed. The HR software company Personio built 400 internal assistants and reports that its top ten carry roughly 80% of the value, so a headline count means nothing.
11 Fear the mailbox providers before the regulators.
Two risks have documented, dated damage behind them: mailbox providers cutting you off, and platforms enforcing their rules. The privacy fine everyone talks about has the least evidence behind it, so it ranks last.
One campaign will not fill a pipeline on its own.
One campaign produces 250 to 300 sends a month and books about 2.25 meetings, roughly 1.5 of them held once people drop off between booking and calendar. Over a quarter that is 6.75 booked and 4.5 held. Those figures are arithmetic rather than observation: Chapter III works them out from a reply rate and a meeting-conversion rate and grades how far to trust them. If your plan needs more, add campaigns: one operator has checking room for a second campaign with its own list and offer, and past two campaigns the plan needs another person.
Eight documented builds did not survive, and most of their builders said so themselves.
This research base contains eight specific builds and tactics that stopped working or were shut down, each attached to a named operator, a date and a stated reason. Most of those operators repudiated their own work in public; at least one never did, which is why the count is eight builds rather than eight confessions. Reading that list is the fastest way to check whether you are about to build the ninth.
The responsible controls pay for themselves.
Three controls each pay three ways at once: the suppression list, a complete audit log, and a cap on how many AI generation steps run back to back. Each one keeps you legal, protects deliverability, which is whether your mail reaches inboxes at all, and makes the system measurable.
You are building an exoskeleton, not an assembly line.
An assembly line takes the person out: a list goes in, emails come out, and the operator is demoted to monitoring a dashboard. An exoskeleton keeps the person in and multiplies them: agents fetch, research, draft, check and file, while every judgment that touches a buyer stays human. The eight documented builds that died were assembly lines, and nearly every item on the durable list, the tactics that survive copying, is something only a person can hold.
This guide is for deep-research selling to a narrow list.
If you sell by researching a small number of accounts deeply, this guide is for you at any company size. Scale changes how many operators you run, not the architecture.
Deep research lowers your volume, and that is the point.
A campaign built on deep research into a narrow list produces a few hundred touches a month. That is because the limit is how much research a human can check before it goes out.
Software picks the targets, a person starts the campaign.
That is the autonomy line, and it sits in the same place whether you are one person or Snowflake's 700-person marketing organization.
Capacity is per person, volume is per campaign.
Every capacity figure counts one person and every volume figure one campaign: one person can properly check 5 to 10 buyer-touching outputs a day. One campaign spends 4 to 5 of those on newly researched contacts, while replies, the watch list, warm accounts awaiting a trigger, and agent upkeep draw down the rest. So one campaign per operator is comfortable, two is a full load, and a third degrades checking into clicking through. Research never counts against the budget, because agents multiply it freely; the ceiling counts only what leaves the building. Building, vetting and correcting an agent took about two weeks in the one measured account, so Chapter IX counts agents doing checked work, not installations.
Build inbound and outbound on one system.
Give both the same four things: one list of people, one library of offers, one suppression list, and one audit log. The only real difference between them is who moved first, the buyer or you. Split them into two teams with two sets of tools and you will maintain two of everything that matters. That includes two answers to who you are allowed to email.
Enterprises run more campaigns, not bigger ones.
Scale multiplies the number of campaigns without softening the statistics inside any one of them. Take SaaStr, aiming at $15M in revenue, whose chief AI officer Amelia Lerutte published as Amelia Ibarra until 2026. In February 2026 she counted twenty to thirty agents and vibe-coded apps, meaning apps built by prompting, in production. They set a target of 200 emails, asked one of their own agents to count actual sends, and got 87. At 87 sends you cannot prove anything at all, at any effect size.
Stop reading if you want reach without research.
Every recommendation here assumes each outgoing email has been researched first, and that when the two collide you cut the send count rather than the research. If you want more researched volume, run more campaigns and add people. At 250 to 300 sends a campaign a month, 2,000 researched emails is seven or eight campaigns carried by more than one person, on this same architecture. What breaks is the other trade, keeping the reach and dropping the research, and if that is the one you want then nothing downstream will help you.
Daily checks and working agents set your ceiling.
Your ceiling is how many outputs one person can actually check in a day, together with how many agents that person can keep doing real work every day.
Read as a builder even if you buy.
Some layers of the stack get more valuable to you every month you own them, while others are work you rent by the month. You cannot buy well until you can tell the two apart, which is what reading as a builder makes visible.
This guide's scope ends at the held meeting.
This guide measures up to the held meeting: a prospect who actually showed up. Everything after that is your own sales process, which none of this research measures. One warning, stated once: every meeting figure ahead is built from vendor-reported reply and show rates, in the two regimes, high-volume and researched sending, that Chapter I separates, and nobody has independently measured a working system end to end.
This guide covers US senders writing to US recipients.
Everything here assumes both ends of the message sit in the United States, where the rule is opt-out. Compliance there is the four chores in Chapter X, yet a US ideal customer profile usually collects European names without anyone deciding it should. Several European countries require prior opt-in consent for business-to-business email, and one enrichment vendor was fined two million euros over how it built lists. So filter non-US recipients at the list-building step, the only place the filter is cheap, or take advice from somebody working that jurisdiction daily. Everything else applies wherever you sell, because architecture, components, costs and controls do not change at a border.
Give this guide to your agents, whatever harness they run in.
Every node here is written to stand alone, which makes the guide an operating manual your agents can read as well as you can. Author's note: the first reader to put a draft to work did exactly that, feeding it to his agents to help them help him, in a different harness entirely, Cursor driving Smartlead, and building three 100-contact lists with a distinct offer for each. The architecture is component-based, so the central control system can be Claude Code or Claude Cowork, or another agent harness such as Cursor or OpenAI's equivalents, with two Claude-specific parts to swap: the licence analysis in Chapter XX, which reads Anthropic's terms, and the named controls in Chapter XIV, which need equivalents in whatever harness you run. The components, the controls, the arithmetic and the checklist do not change.
Strategy
What the evidence says, and what it means for how you spend your attention. Twelve chapters establish which levers actually move results, in what order, on what clock, and what one campaign honestly returns. Nothing here names a product, because everything here outlives the products. Read this half to decide whether the motion is worth building at all, and to know what to expect from it before you spend a dollar.
Email changed in 2024, so the old benchmarks lie.
Providers rewrote their rules on published dates, and buyers changed how they shop over the same years. So every baseline you inherited was measured on a channel that no longer exists. Most numbers you have been quoted cannot be traced to a source; every number below can.
1.1 Providers and buyers changed the rules on purpose.
Both sides of the channel moved deliberately while outbound tactics stayed put. Read the result as a soft quarter and you will push harder on a channel whose rules changed underneath you.
1.1.1 Every channel dies the same five deaths.
Every channel dies in the same order: tooling gets cheap, volume explodes, attention saturates, the platform cracks down, and buyers adapt. The easier a tactic is to automate toward unlimited volume, the shorter its life. That puts the working estimate for a current outbound tactic at roughly 18 to 36 months. Price any new channel against that clock before you commit, and treat the estimate as reasoning from the pattern, not a measurement.
1.1.2 Two dates ended the volume era.
On February 1, 2024, Google and Yahoo began enforcing authentication on any sender above 5,000 messages a day. They also set a spam-complaint target of 0.1% with a hard cliff at 0.3%, and made RFC 8058 one-click unsubscribe mandatory from June 2024. On May 5, 2025, Microsoft skipped the junk-folder grace period and moved straight to rejection with code 550 5.7.515. That code turns mail away at the door instead of filing it into spam. Both are published rules, so you can read them, comply, and date the moment your inherited benchmarks stopped applying.
1.1.3 Authentication only buys you a ticket.
SPF, DKIM and DMARC are the three DNS records that prove you are who you claim to be. Google reports 265 billion fewer unauthenticated messages in 2024. Passing all three only gets you considered, because Google still decides placement on what recipients do with your messages and whether they consented. Set the records up once, then spend your attention on consent and engagement, which are the inputs that move placement.
1.1.4 The best deliverability data contradicts itself.
GlockApps plants seed addresses across providers to measure where mail lands, on a methodology built for newsletters, not one-to-one cold mail. Its Q1 2025 reading showed inbox placement falling year on year by 26.73% at Office 365, 22.56% at Outlook and 10.49% at Google Workspace. Senders above a million messages a month sat below 28% while moderate senders rose; three quarters later, Q4 2025 pointed the other way. Senders at 1,000 to 10,000 a month lost 15% of Gmail and 19% of Workspace placement, while the million-plus group gained 20% on Gmail. The stated mechanism is filters increasingly rewarding established long-reputation senders; Growthly adds a third variable, abandoning .info for .co after a top-level-domain blacklisting. Read GlockApps quarterly; two more quarters of low-volume decay would undercut this guide's design and signal moving weight to LinkedIn, video and inbound-led work.
1.1.5 Pipeline dollars hid the conversation collapse.
Watch conversations per rep, not pipeline dollars, because the dollar figure can rise while conversations collapse underneath it. The Bridge Group's sales-development survey, the one long-running dataset here from a consultancy with no outbound software to sell, shows exactly that. Pipeline per rep rose from $2.83M in 2022 to $3.78M in 2025 on bigger deals, while live conversations fell roughly 40% over the decade. The share of reps hitting quota dropped to 60%, the lowest in ten editions. The same 2025 edition records 4.1 quality conversations per rep per day, the survey's first rise, in the year attainment bottomed. The widely quoted "seven-plus a day in 2014" has no traceable primary Bridge Group source, so do not repeat it.
1.2 Most celebrated AI sales wins were inbound.
Almost every famous AI-GTM result in circulation is inbound recovery, expansion into existing customers, or self-serve funnel design, yet each one gets an outbound headline. Check which of the three a case study describes before you let it set your outbound plan.
1.2.1 Salesforce's 3,200 opportunities were unworked inbound.
Salesforce, source of the most-quoted agent result in enterprise software, pointed agents at its own inbound leads, specifically the 75% nobody had ever touched. It reports over 3,200 opportunities in a few months but never publishes how many leads it started with, so the efficiency is unknown. Read it as a backlog-clearing result, and find your own untouched inbound before you copy anything else from it.
1.2.2 Personio's 140 meetings a week come from website chat.
Personio, an HR software company, books about 140 meetings a week from an agent that talks to visitors already on its website. Those same seven days bring roughly 200,000 sessions, a 0.07% conversion rate on traffic that arrived on its own. The result is routinely quoted as evidence for outbound agents, so ask any vendor citing it where the visitors came from.
1.2.3 The autonomous AI-SDR thesis collapsed in public.
An AI-SDR is software sold to do the job of a junior sales development rep, and 11x was the category's loudest example. It displayed customer logos for companies that said they were not customers, including ZoomInfo and Airtable. Reporting plus ex-employee accounts put roughly $14M of claimed recurring revenue against about $3M that survived once customers reached the three-month break clause. The founder stepped down in May 2025, and 11x disputes the framing. A whole category thesis failed in the open, so put any autonomy claim on a short break clause of your own.
1.2.4 Believe the buyer who shut down his own outbound.
Ramp, a spend-management company, shut down its own five-year-old outbound automation team in late 2025, even though that team produced roughly 30% of pipeline. Co-founder Gene Lee cited cold-outbound fatigue and, in a widely reproduced line from his own post, called it slop in, slop out. Ramp redeployed the capability into growth engineering rather than dropping AI.
1.2.5 The "stop hiring humans" vendor hired a human.
Artisan ran "Stop Hiring Humans" billboards from October 2024, then retired the slogan and hired its first human business development rep in August 2026. Its chief executive told TechCrunch in April 2025 that the billboards were "mostly just for attention." The company was hiring 22 more people of its own at the time, including sales roles.
1.3 Ask what the benchmark was divided by.
Most published benchmarks divide replies by opens, by delivered messages, or by contacts, and almost none by messages attempted. Change that base and a 5% headline becomes 0.45% on identical activity, an elevenfold difference produced by nothing but math.
1.3.1 Divide replies by every send you attempted.
Demand replies divided by all messages attempted, and throw away any rate divided by anything else. The denominator is a choice, not a fact: the same work reads as 5% when divided by opens and 0.45% when divided by attempts. That 0.45% comes from the outbound agency Belkins across 7.5 million of its own 2025 sends with no open-tracking inflation. It is an agency average at volume rather than a small researched campaign. Hold every benchmark you meet to this denominator before you compare it to your own.
1.3.2 At volume it takes about 157 sends to find one interested buyer.
Two measurement regimes produce two true numbers. Competent researched campaigns get replies on 3 to 6% of sends, counting refusals as replies. The top tenth get 8 to 12%, so a hundred sends returns about three replies. At volume, positive replies run near 0.64% of sends, one in 157, which puts a hundred sends well under one truly interested person. That 0.64% is an agency and platform average measured at volume, well below the 3 to 5% positive reply good researched campaigns reach (Saleshandy, across 53 million emails). Both are real, measured under different regimes, and the planning arithmetic in this guide runs on the researched band. Sends per booked meeting is contested: Gong's dataset averages one meeting per 344 emails. Platform funnels at volume run from about 1% of sends (Smartlead, across 14.3 billion sends) down to about 0.2% (Instantly). Plan against that whole range, not the hundred sends practitioners quote.
1.3.3 Open rate no longer measures anything.
Apple's Mail Privacy Protection loads tracking pixels whether or not anyone opens the message. One publisher panel of roughly two billion emails across about 80,000 deployments saw its total open rate move from 22.6% to 40.5% across the rollout. There was no change to audience or content. Apple now accounts for between half and 58% of all reported opens, so turn open tracking off in week one and report replies instead.
1.3.4 Phone and LinkedIn rates are just as thin as email.
On generic purchased data the phone runs about 370 dials per booked meeting, measured across a full year of tracked rep activity. On LinkedIn, replies to personalized connection notes fell 37% between May 2025 and April 2026, from 3.5% to 2.2%, while connection acceptance held at 26 to 30%. That is measured across 13.2 million requests by Expandi and more than 20 million by the outbound agency Belkins. Acceptance holding while replies fall says the note is what is decaying, so rewrite the note before you abandon the channel.
1.3.5 High-reply campaigns share four things you cannot scale.
High-reply case studies that survive inspection share four ingredients, the first a list of under 50 named recipients justifiable one at a time. Lists under 50 average 5.8% replies against 2.1% above 500 in Belkins' 2024 study, which is most of the lift. The second is a genuine recent-weeks trigger a person at the account would recognize, not the bought intent surge whose miss rate this chapter prices below. The third is contact data verified at send time rather than bought and trusted, because addresses decay around 2.1% a month, faster in software. Fourth is the counting: a generous definition counts every response including refusals and divides by openers rather than attempts, turning 5% into 0.45%. Part of any 40% headline is a counting choice, and no source shows 40% sustained at volume, because the first ingredient vanishes at scale.
1.3.6 Founders reply most and VPs least.
Across 7.5 million of its own 2025 sends, Belkins found founders and owners replying at 0.57%, C-level at 0.42%, and VPs at 0.32%. Reply rate falls the further you get from the founder. That inverts the standard advice to start a rung below the C-suite. Write the founder first.
1.4 Low send volume makes A/B testing useless.
A few hundred sends a month cannot settle a wording question, and that holds at any headcount. Spend your statistical budget on the decisions with large effects, and settle the small ones by craft.
1.4.1 One clean test costs more sends than you have.
Detecting a 50% improvement on a 5% reply rate at the usual statistical standards takes 1,469 sends in each version, about 2,938 sends for one clean two-way test. At 250 sends, a measured 5% reply rate carries a confidence interval from 2.9% to 8.5%. That band would hold the true rate in 95 runs out of 100. So the same result would look identical whether the truth were 3% or 8%.
1.4.2 Only the biggest decisions are worth testing.
Four choices move results by two or three times: who you target, what you offer, which channel you use, and what kind of hook you lead with. The offer is the one you can test, because a new offer runs against the list you already built, and an effect that large resolves within a few thousand sends. The other three get decided by judgement and watched over quarters. Subject lines, opening sentences and send times are craft decisions, and testing them consumes your entire volume while returning noise.
1.4.3 Meeting rate is permanently untestable at low volume.
Resolving a 50% difference in a 0.75% meeting rate requires tens of thousands of sends in each version, which at 250 sends a month is decades of waiting. That 0.75% is the researched-campaign meeting rate Chapter III derives, and the figure this guide plans on. Ten operators each running 250 sends work different lists and offers, so their results never pool, and headcount buys no relief.
1.4.4 Ten weak tests manufacture a winner.
Run ten comparisons where nothing is truly different and there is a 40.1% chance at least one comes back significant, rising to 64.2% at twenty. Check results ten times along the way and a 1% significance threshold behaves like 5%. Bayesian statistics break the same way under repeated peeking: on a 6-in-200 result the Bayesian interval runs about 1.1 to 5.6%, exactly as wide as the standard one. Fix the comparison and the send count before you start, then read the result once.
1.4.5 Big companies get no relief from the statistics.
Each campaign in a large program runs a different list, offer and segment, so running many in parallel gives you many small datasets, not one large one, and never the send counts one clean statistical test requires.
1.5 The most-repeated numbers often have no locatable source.
Trace the standard outbound statistics and most dead-end in blogs citing blogs, motivational books, or vendor-funded studies from another decade. Those numbers are setting your cadence, your urgency and your expectations right now, so delete them deliberately rather than leaving them in place.
1.5.1 The two numbers behind every cadence have no study.
A cadence is the fixed run of touches a rep works through before giving up, and two numbers set its length everywhere. One is an average of eight attempts to reach a prospect, and the other is 2% of cold calls ending in an appointment. Both carry named attributions, TeleNet and Ovation for the eight attempts and Leap Job for the 2%. Both traces end in blogs citing other blogs with no locatable study. Every cadence length built on the eight-attempt number rests on nothing. Set your sequence length from your own record of how many attempts came before your last fifty replies.
1.5.2 The five-follow-ups rule came from a motivational book.
The claim that 80% of sales require five follow-up contacts, and its companion "44% give up after one," are the standard argument for persistence. They are the reason the default sequence runs five steps or more. Both trace to Go for No!, a sales-motivation book published in 2000, and neither traces to any research. A quarter century of cadence design rests on that book. Persistence may still be right for your market, so keep it if your own reply curve earns it, but stop offering the 80% as the evidence.
1.5.3 The bar teams judge themselves against is folklore.
The claim that a competent cold-email campaign replies at 8.5% comes from Backlinko and Pitchbox's 2019 study of 12 million outreach emails. That study was published without a methodology worth the name, and every serious trace ends in folklore. The current average measured across sending platforms at volume is 3.43%, with the top quartile at 5.5% and the top decile at 10.7%. Anchoring on 8.5% makes an ordinary result look like personal failure, so reset the bar to 3.43% and treat anything above 5.5% as strong work.
1.5.4 The 57% buying-journey figure was never a constant.
CEB's famous finding that 57% of the buying journey completes before a seller is contacted came from just 22 companies. The range across them ran as high as 70%, and a decade of strategy decks has quoted the reported midpoint as a law. The direction is sound even though the number is not, so use the better-sourced version instead. 6sense finds first seller contact 61% of the way through a 10.1-month journey; treat it as a range, not a constant.
1.5.5 The five-minute rule is 2007 inbound web-form data.
The rule says a lead contacted within five minutes is 21 times likelier to qualify than at thirty minutes, and 100 times likelier to connect. It traces to James Oldroyd's 2007 study of dialer-vendor InsideSales data across six companies and more than 15,000 leads. A 2011 Harvard Business Review follow-up drew its seven-times-within-the-hour figure from 1.25 million leads across 2,241 audited US firms. Both looked only at inbound web-form submissions and measured reaching and qualifying rather than closing; the work is often mislabeled an MIT study. The companion 78% buy-from-first-responder claim loops through aggregator blogs to no source, and no independent study ties response time to conversion for cold-outbound replies. This guide's standard is minutes for a form-fill and same-business-hour to same-business-day for every other raised hand, answered by a person, not a link.
1.5.6 The 287% multichannel lift is B2C e-commerce data.
The claim that multichannel campaigns outperform single-channel ones by 287% gets quoted whenever someone wants to justify a fourth and fifth channel. It comes from a consumer e-commerce marketing platform with no control group and no causal design. That makes it unusable for B2B outbound no matter how many calling vendors re-cite it. Add a channel when you can staff it to the same standard as the first one, the only test that has ever governed this decision.
1.5.7 The 70.3% data-decay figure is not Gartner.
The claim that 70.3% of your contact database goes stale in a year is the figure data vendors use to sell continuous re-verification subscriptions. It is attributed to Gartner but traces to a data broker's blog. The underlying record shows an InsideView and Biznology field-level worst case, one bad field in one bad segment, stretched across every contact. The defensible aggregate is roughly 2.1% a month, with software and technology contacts at the bad end decaying 35 to 70% a year. That range overlaps the debunked headline, but manufacturing and government contacts decay far more slowly, so set re-verification schedules by industry, not one interval.
1.5.8 Fake multiples get wrapped around real findings.
The pattern is a real but modest finding dressed in a multiple nobody measured, and it survives because a multiple is easier to repeat than a percentage. A March 2026 study presented as an independent Gartner audit ranked the website visitor-identification vendor Leadpipe first at 82%. Leadpipe commissioned and paid for the study and distributed it through a paid PR wire. The release carried a fabricated media contact at "1000 Fiction Ave." whose listed website was gartner.com. The claim that offer-based calls to action beat interest-based ones fourfold runs the same way, tracing to a single third-party agency blog. The verifiable number is +28%, published by the sales trainers Jason Bay and 30MPC from Gong's 85-million-email dataset. The advice survives at +28%, so keep the offer-based ask and check who paid before you repeat the multiple.
1.6 Mail gets filtered long before a human sees it.
Seven layers now sit between your send button and a reply coming back, and almost all outbound effort goes into the message, the only part you control. Each layer takes its own cut, the cuts multiply, and the product is the sub-1% reply rate you keep measuring.
1.6.1 A bad send is worth less than no send.
Gartner's buyer surveys find 73% of buyers actively avoid suppliers who send irrelevant outreach, and its 1,464-buyer study on personalization found 53% reporting a negative experience. Those buyers showed 3.2 times the purchase regret and were 44% less likely to buy again.
1.6.2 Seven layers each take their own cut.
The layers run in order: authentication, sender reputation, tab and priority routing, AI summarization and batching, and delegation to an assistant or an agent. Then come the human's own attention and, last, the decision to reply. Each removes a fraction of what reaches it and the fractions multiply. No single source measures the chain end to end, so the size of each cut is an estimate. But the compounded shape matches the sub-1% strict reply rates actually observed.
1.6.3 Assume a machine reads your email first.
Gmail's Gemini-powered AI inbox and Outlook's Prioritize My Inbox summarize and pre-rank mail before a person sees it. So your verifiable trigger, the checkable recent event justifying your email, has to land in the subject line and the first sentence. The rollout moves in fits, and Microsoft pulled the mobile priority view in February 2026 citing user feedback and cost. But the direction holds: the body of your email is increasingly never rendered for a human at all.
1.6.4 The new gatekeeper costs $7 to $40 a month.
Buyer-side email assistants auto-archive anything that fails a relevance test, and the person paying for the subscription never sees it. Eight of them carried published prices in 2026: SaneBox at about $7 a month, Fyxer and Serif near $30, and Superhuman Business about $40. Cora takes a different shape and turns the inbox into a twice-daily digest. Buy one yourself, because it is the cheapest way to watch how your own outreach gets triaged.
1.6.5 Buyers rank the vendors before they contact anyone.
6sense surveyed more than 4,000 B2B buyers about purchases above $25K and published the study in November 2025 with a named research lead. It finds 94% of buying groups rank vendors before contacting any of them, and the preliminary favorite wins 77% of the time. 6sense sells into this market and published the study itself, but even discounted for that, your message arrives into a decision largely already made.
1.6.6 Nothing detects AI, yet AI mail gets flagged.
Gmail and Outlook score authentication, volume, reputation, complaints and engagement, and nothing in that stack scores whether a machine wrote the text. AI-only email is still flagged as spam at roughly 2.7 times the human rate. One 12,000-email test put it at 7.8% against 2.9%, and a second across 100,000 paired sends put it at about 8% against 3%. What the filters catch is identical structure repeated at volume, so have a human read and change the copy before it sends.
1.6.7 Free writing stops being a signal.
A message is evidence of effort only while writing it costs something, which is Spence's 1973 signaling argument applied to a mailbox. Once generation is free, outreach collapses into a pool the buyer cannot sort. Credibility migrates to the signals that stay expensive: a warm introduction, a referral, standing in a community, and work published under a real name.
1.7 Outbound's job is getting on tomorrow's shortlist.
The Ehrenberg-Bass Institute's 95-5 rule puts roughly 5% of a B2B market in-market in any given quarter. So outbound's realistic job is entering the consideration set before the buying window opens. That is a different objective from converting at decision time, and it is measured differently.
1.7.1 The 90-day window is a leftover.
6sense puts the B2B buying journey at 10.1 months, down from 11.3, with first seller contact happening 61% of the way through. That leaves about 3.9 months of journey after your first touch lands on an in-market account, and the famous 90-day window is that remainder. It measures how much journey was left, so stop reading it as a verdict on how fast your prospecting works.
1.7.2 Being remembered beats being persistent.
The eventual winner was on the buyer's first-day shortlist 95% of the time, and buyers start the conversation in more than 80% of deals. Persistence works on the small share already looking, while memory reaches everyone else, so spend on being recalled at the moment the shortlist gets written.
1.7.3 Two-thirds of buyers now prefer no rep at all.
Gartner tracks how many buyers would rather buy with no rep involved. The figures run 33% in 2020, 61% on 2024 fieldwork, and 67% on 2025 fieldwork published in March 2026, on samples of 632 and 646 buyers. Each point is a buyer who would rather research alone than talk to you.
1.7.4 Budget to be remembered in month six.
Treat the job as getting into the consideration set rather than converting this quarter, and the budget, the metrics and the patience windows all change shape.
1.8 Weight every teacher by what they sell.
Most outbound teaching is a product wearing the clothes of advice. Rank a voice by how tightly the advice couples to something the teacher sells, and by whether they publish the base numbers behind their rates. Audience size answers neither question.
1.8.1 Rank by disclosed numbers per hour.
The useful measure of a source is how many dated, fully explained numbers they hand you per hour of content. By that measure, "here is what we stopped doing and why" is the highest-value content class in this space, because publishing it costs the speaker something.
1.8.2 Every practitioner reply rate here is self-reported.
No headline practitioner reply rate in these sources is confirmed by a named client. That includes the 11 to 32% positive-reply range published by Jordan Crawford, the most-respected consultant in this field. It also includes the roughly 20.5% house average published by the email-writing tool Lavender. That figure measures the vendor's own users against independent baselines of 3.43 to 5.8% and conveniently supports buying that vendor.
1.8.3 Take the loudest teacher's method, quarantine his numbers.
Nick Saraev runs the largest automation channel in this space, about 490,000 subscribers on a third-party tracker in August 2026, and his productization ladder is worth taking. He also publishes his own school's monthly revenue as $250K, $290K, $330K and $335K across his own surfaces, four figures that cannot all be true. And he teaches a 500-emails-a-day sending architecture that contradicts the entire deliverability record, which puts the working range at 15 to 50 sends per inbox. The Trojan Horse Loom playbook widely credited to him traces in primary sources to Christian Bonnier of ListKit. Adopt the packaging and distribution logic, and discount every dollar figure he prints.
1.8.4 The best mental model has no attested results.
Jordan Crawford's framework for high-value vertical outbound, from the field's most-respected consultant, is the strongest one here, and his reputation verifies cleanly as a Clay advisor and investor since January 2021. His flagship cases are demonstrations and explicitly labeled hypotheticals, not delivered campaigns. His own honest calibration is about 2% positive reply to start, rising to 4 to 5% with a well-built value proposition. So plan at that number and treat the 11 to 32% as best-case survivorship.
1.8.5 Signal selling is 2010 trigger selling upgraded.
Watching for funding rounds, job changes and hiring is trigger-event selling, which Craig Elias and Tibor Shanto published in 2010; what is new is faster data pipes. The predictive lead-scoring cohort of 2013 to 2019 made structurally identical claims, and almost all of them were acquired or absorbed, with 6sense the lone independent survivor. Price the current wave against that outcome.
1.8.6 The 37%-versus-19% proof is the vendor's own report.
The headline proof for signal-based selling, a 37% win rate on signal-triggered accounts against 19% cold, comes from the signal vendor Champify's own impact report. No one else has replicated it. Topic-intent data is the weakest layer underneath, meaning signals bought from a co-op that watches what accounts read across the web. A semi-independent test puts the co-op provider Bombora's precision at 81%, roughly one false positive in five. And about a quarter of flagged surges produce nothing within six months.
1.8.7 Five sources carry the weight, and two are vendors.
Five sources carry the reliable weight: the Bridge Group's decade-long sales-development survey, Gong's 85-million-email dataset, and Gartner's independent buyer panels. The other two are 6sense's study of more than 4,000 buyers, and the statutes and court records, which cannot be spun. Two of the five are published by companies selling into this market, survivable only because they disclose sample size, fielding dates and method. Disclosure is what earns a source its weight, so read the method before you read the headline.
1.9 Honest math ends message testing and reply reporting.
Set what one clean test costs against a 250-send month and testing your offer rather than your wording becomes math, not taste. Counting meetings held rather than replies received follows from the same arithmetic, because a 5% claim and a 0.45% reality, counted with different denominators, can describe one campaign. And Apple's Mail Privacy Protection corrupts the open rate most teams still report.
Judge every tactic by what survives everyone copying it.
Ask two questions about any tactic: does it still work once everyone is running it, and do you own the ground it stands on? Effort answers neither, which is why the hardest builds keep showing up on the failure list. Every tactic in this guide gets sorted on those two questions.
2.1 Two questions decide whether a tactic lasts.
A tactic lasts when it cannot be mass-produced and does not sit on a platform that can change the rules at will. Effort is a separate matter, because the hardest builds documented here failed while the cheapest move in the set produced the best reply rate.
2.1.1 What automates cheaply dies fastest.
Anything that automates cheaply to unlimited volume saturates first, because everyone gets the same tool in the same month while attention stays the same size. Its runway is set by how fast your competitors adopt the tool, not by how well you run it.
2.1.2 The cheapest thing to try got the best reply rate.
Asking a conference organizer for the attendee list costs almost nothing. Campaigns run on those lists produced the only credible 15 to 20% reply figure in this research base. Nick Saraev reported it against his own interest, since he sells scraping and automation education. But it is self-reported off narrow researched campaigns, so read it as one operator's account, not a benchmark. Nobody has measured what each ask costs or how often organizers say yes; the same operator guesses roughly one in ten for smaller conferences.
2.1.3 The heaviest builds in this guide are the ones that failed.
Four builds documented here took more effort than anything else in the set. They were a 300-mailbox sending estate, a multi-vendor AI-SDR stack, custom scaffolding patching a model weakness, and a home-built marketing automation platform. All four stopped working quickly, and all four sit inside the eight failed cases rolled up later in this chapter. So read them as the heavy end of the eight rather than a separate list.
2.1.4 Good writing cannot save a rented domain.
The outbound agency Growthly built its whole sending estate on .info domains and sent roughly 100,000 emails from them. In August 2025, during its launch week, that entire top-level domain, the suffix at the end of the address, was blacklisted wholesale. And because the backup domains were also .info, the spare set died with the primary one. Growthly rebuilt on .co and now treats a spare set on the same suffix as no spare at all.
2.1.5 Cheap and owned beats expensive and rented.
A cheap asset you own outright and nobody can copy beats an expensive one you rent on someone else's platform, over every time horizon. Score only one of the two questions and you will keep mistaking expensive rented infrastructure for a moat.
2.2 Classify every tactic before you fund it.
Every tactic falls into one of three buckets. Ephemeral ones work now and are measurably decaying, evolving ones work while the ground under them keeps moving, and durable ones compound and cannot be mass-produced. Sort first and budget second, because paying asset prices for a consumable destroys build time.
2.2.1 Take the money from a decaying tactic, then walk away.
Ephemeral tactics pay real money right now, whether that is bidding in a marketplace, reactivating a warm list or borrowing someone else's audience. Take the proceeds and spend them on something that compounds, because building deeper into one of these only buys a bigger position in a known expiry date.
2.2.2 LinkedIn connection notes are measurably decaying.
The LinkedIn automation vendor Expandi tracked replies to personalized connection notes across 13.2 million requests, a whole-platform measurement at volume. Replies fell from 3.5% in May 2025 to 2.2% in April 2026 as templates saturated. Notes barely change acceptance, 26.42% with a note against 26.37% without in the Cleverly and Expandi data. What a note buys is the reply after acceptance, 9.36% against 5.44%.
2.2.3 The autonomous sales rep already failed outright.
An SDR is a sales development rep, the person whose whole job is prospecting for meetings, and the AI-SDR category sold software to replace one. The vendor 11x lost 70 to 80% of its customers as soon as they reached the three-month break clause. Ramp, a buyer with no vendor stake, shut its own AI-outbound unit in late 2025 even though it drove roughly 30% of pipeline. The vendor Artisan hired a human business development rep in August 2026. The tooling still works as a copilot, so buy it as one, because the self-running SDR was not left to fade, it was disproved.
2.2.4 Expect one email rule change every year.
Google and Yahoo began enforcing bulk-sender rules in February 2024, and Microsoft went straight to hard rejection in May 2025. Google replaced graded sender reputation with a binary pass or fail in October 2025. It then began locking down whole Google Workspace tenants instead of single inboxes in late 2025. That pace is the steady state, so budget a standing maintenance line instead of a one-time build project.
2.3 Sixteen things last, and almost nobody builds them.
Sixteen items pass both tests: they cannot be mass-produced and they do not sit on ground someone else owns. Most look boring in a demo, so almost nobody builds them, and that neglect is a large part of why they stay durable.
2.3.1 One pass down the durable list shows what your plan is missing.
Four are about people: a named accountable human, your sender brand, human reply handling, and warm introductions and referrals. Six are assets you own outright: an audience whose consent travels with you, exclusive lists nobody can scrape, a privately manufactured signal, proprietary proof, community standing, and your primary domain's age. Four are craft: define exactly who you sell to, and build a real offer while getting good at building offers. The other two craft items are deep research at low volume, and doing the work instead of asking for a call. The last two are plumbing: a calibrated instrument set, and suppression and audit infrastructure, the enforced record of who must never be contacted again. Ninety seconds is enough to check your plan against all sixteen.
A named human accountable for the claim.
69% of buyers turn to reps specifically to validate AI-generated insights (Gartner, 645 buyers, May 2026), so accountability is the thing they are buying rather than a courtesy. An agent cannot take responsibility for a decision, which is why this survives any amount of copying. The honest gap: the corpus, the 65-report evidence base behind this guide, contains no documented incident of an AI SDR sending a hallucinated commitment. So the risk this protects against is heavily guarded and never yet reported.
Your sender brand.
The CEO of Monaco, who sells message automation, names sender brand as a determinant of reply rate that is not the message. That is a witness testifying against his own product. The cleanest observation in the corpus is an 11% reply campaign into a community the sender visibly belonged to. Nobody has isolated brand in a controlled test, and it quietly contaminates every reply-rate benchmark you will ever read.
Human reply handling.
The operator running four to five AI SDRs declines the vendor's own checkbox recommending auto-reply: "I'm like, no, we'll reply." Reply triage is trivially buildable, and four of four inbox builds do it. So its absence from fourteen of fourteen recorded outbound builds is a choice rather than a tooling gap.
Warm introductions and referrals.
An introducer stakes their own reputation, which is what makes the signal costly to fake. The corpus carries 14.6% conversion against 1.7% cold with sales cycles half as long. It marks the magnitudes as varying while the direction is consistent, so carry the direction. Buyers initiate more than 80% of purchases themselves, which is why being introduced beats interrupting. The cost is a network built over years, and no spend accelerates it.
An audience whose consent travels with you.
An audience you own is the one asset that survives every future the corpus, this guide's research base, models. The buyers you will never hear from are reading: 63% of buyers who never talk to sales spend more than an hour a week on thought leadership, and 95% say it makes them more receptive (Edelman and LinkedIn, 1,934 buyers). It is a stock rather than a flow, so it keeps working through a pause, and every delta behind it is vendor-reported, so trust the direction and not the sizes.
Exclusive lists nobody can scrape.
The operator who sells scraping education states, against his own interest, that his best campaigns all ran on conference attendee lists. Those lists' contacts are not on the internet at all. A list that is not on the market cannot saturate, which is exactly the axis this chapter sorts on. The reply figure attached to the story is one man's self-report, Grade C, the guide's lowest evidence grade, so copy the method and quarantine the number.
A signal you manufacture privately.
Sign up for your prospect's own lead form or newsletter and watch what happens: how fast the response comes, what it says, and what follows it. Nobody else has that observation, so it cannot saturate the way bought signals did. And every documented signal failure in the corpus was a cheap public signal everyone bought at once. Nobody has run this play, including the operator who proposed it, so treat it as an experiment you could be the first to try.
Proprietary proof.
A competitor's agent can research your prospect, and it cannot have done your work. Named results, original research and a case-study corpus are the outbound version of what survived Google's March 2024 move against unoriginal content. That was when value migrated to original research and named brands. That supporting evidence is a proxy from search rather than an outbound measurement, and the cost is months to years of real work before the first quotable case.
Community standing.
Being a known member of the communities your buyers trust probably helps and has never been measured, which makes this the thinnest-evidenced item of the sixteen durable assets in this chapter. It earns its place because standing is granted by other people, so it cannot be bought.
The age of your sending domains.
A domain earns trust with age, and age is the one input money cannot buy back. Google and Microsoft treat a domain's first 30 days as probation. That is why Chapter XXI says to buy your sending domains in the next hour under either build. On a fleet they age through warmup, and on one or two mailboxes they age while you send. Your primary domain's clean history sits in the same ledger, which is one more reason cold mail never touches it.
A dense definition of who you sell to.
The list-quality spread is the largest measured effect in the corpus at roughly a hundredfold, which Chapter IV measures. And segmentation judgment is the least automatable step in the stack. Amelia Lerutte, running twenty to thirty agents in production, keeps it manual. The segments are "just all in my brain right now", and "no one can create these types of segments autonomously". Agents starve without this judgment, and idle agents are a documented running cost.
A real offer, and the skill of building offers.
Offer changes are a 2-to-3x effect, which makes this the one lever a low-volume sender can genuinely test. Gong's 85 million emails put a compelling offer at +28%, and pitching cuts replies by as much as 57%. The buyer-side proof is a single episode: about a hundred vendors reached one buyer during a purchasing freeze. The sole deployment went to the CEO who offered to do the deployment himself. An agent can compose an offer, and it cannot decide what you will guarantee or bear the cost when the guarantee is called.
Deep research at low volume to a narrow list.
Durable precisely because it does not scale, and scale is what everyone else wants. Belkins' 2024 study puts replies at 5.8% for lists under 50 recipients against 2.1% above 500. One condition: the research has to be something the recipient's own agent could not produce in ninety seconds, or it is table stakes.
Doing the work instead of asking for a call.
The thing offered is your attention, which cannot be automated, and as free AI-generated "value" floods the channel, delivered work becomes the signal that separates. The evidence is the same single purchasing-freeze episode as the offer row above, a freeze won by delivered work, so this rests on the mechanism plus one documented event rather than on a sample.
A calibrated instrument set.
Proving a 50% lift on a 5% reply rate takes 1,469 sends per arm, about a year of a solo operator's whole volume, which Chapter I prices. So the durable skill is knowing what you cannot measure. The defensible chain is reply, positive reply, meeting held, opportunity, and cost per opportunity, with opens and raw send counts ignored. Three separate builds run fake A/B tests on camera, and that discipline gap is the advantage this row buys you.
Suppression, audit and consent infrastructure.
The FTC settled with Verkada for $2.95M in 2024 over unhonored opt-outs, and the statutory exposure runs to $53,088 per email. Every one of the fourteen recorded builds skipped this layer because it does not demo, and one $50K-a-year incumbent's suppression failed silently for ten days. It takes about a day to build and it is permanent, which makes it the cheapest insurance in the guide.
2.3.2 Low-volume researched outreach lasts because of the list, not the research.
Deep research is already mass-produced: agents run it at volume, and the 65 commissioned research reports underneath this guide were produced exactly that way. Research depth a competitor's agent reproduces in ninety seconds is table stakes with a cost attached. What does not mass-produce is the list, because exclusivity is a property of the asset rather than of the effort spent getting it. Run the mass-production test on the list first, and treat research as what makes a good list land.
2.3.3 Who sends matters, and nobody has measured it.
Sam Blond, CEO of the message-automation company Monaco, says sender brand and message-market fit matter more to reply rate than the sequence or the structure of the message. That is a claim that costs him money to make. Nobody has run the clean test of the same list and the same message from a different sender. So until someone does, any benchmark that does not name the sender is contaminated.
2.3.4 Choosing the list beats everything else you do.
A dense, specific definition of your ideal customer produces the largest measured effect in this research base. Will Cyniak of RevPack found roughly a hundredfold spread in meetings per prospect driven entirely by which list was chosen. He measured it at volume across the 214 campaigns and about 250,000 prospects Chapter IV works through. It is also the one step Amelia Lerutte of SaaStr still does by hand while her team runs twenty to thirty agents and vibe-coded apps in production. She does so for the reasons the durable list above records.
2.3.5 A new offer reuses the list you already paid for.
Change the class of offer you make and results move two to three times, one of the four levers Chapter I names: target, offer, channel, hook. The offer is the cheap one to change, because a new offer runs against the list you already built. A new list or a new channel starts a campaign from zero, and the hook is just the offer's opening line. Wording effects are real but smaller, the +28% and minus 57% Gong measured across 85 million emails, and they sit below what your send volume can reliably detect.
2.3.6 Copy the conference-list method, but not its reply rate.
Two items on the durable list, the sixteen assets above, rest on argument rather than measurement. The conference-list reply figure earlier in this chapter is one operator's self-report. And Chapter VII shows the same operator ranking list exclusivity top in one place while blaming copy elsewhere. Community standing has no measurement behind it at all. Take the method, which is to get a list nobody else can get, and leave the percentage out of your forecast.
2.4 Eight builds and tactics died, and most of their makers said so in public.
Eight documented cases stopped working, each with a date and a stated reason, and they are the roll call the rest of this chapter keeps pointing back at. Most were repudiated in public by the named operator who built the thing, and where nobody repudiated one the arithmetic does it instead. Nobody talks down their own work for fun, which makes this the best-evidenced part of the picture. The eight run in order below, followed by the one thing they share.
2.4.1 The sending estate died in launch week, and the rebuild beat its numbers.
Michael Sarouja prices this shape at about 330 mailboxes across 83 domains, roughly $1,162 to set up plus $1,000 a month. Jordan Platten shows 300 warmed mailboxes carrying the same 3,000 sends a day, after two to four weeks of warmup on the clock Chapter III sets. SaaStr never builds sending itself, because its vendors run four to six weeks of warmup first. Growthly ran exactly this shape, published a 0.36% reply rate across about 100,000 emails, then lost the whole estate when one domain suffix was blacklisted. The rebuild reached 0.66%, nearly double yet still under the 0.8 to 1% the same operator calls a bad campaign, and it paid a real second warmup bill. It booked about 20 meetings in a week at a 70% show rate and a 20% close, so book an estate like this as a consumable you will buy twice.
2.4.2 The multi-vendor AI-SDR stack ended in regret.
The bill starts with two AI-SDR platforms at $30,000 to $70,000 or more each, plus a deployment engineer from the vendor and a go-to-market engineer of your own. Each agent costs the two-week onboarding and 10 to 60 minutes of daily attention that Chapter IX prices, while the agents you already have degrade. Jason Lemkin and Amelia Lerutte of SaaStr, who published as Amelia Ibarra until 2026, built the most-cited stack of this kind. They now say they only ever needed one. They would dump both for an off-the-shelf tool even at $50,000 a year, and the honest realized outcome is about the same results with four fewer humans.
2.4.3 The self-optimizing loop learns nothing.
Every two weeks, Jordan Platten's system swaps losing copy for winning copy on its own, with nobody supervising. The swaps run across roughly 90,000 sends a month spread over dozens of variants. It is promoting noise, because proving a 50% lift on a 5% baseline needs 1,469 sends per version, which Chapter I prices. And 3 to 5% is the band researched campaigns, those built on per-prospect research, hit. Platten has not disowned this one, so the arithmetic does it for him: real engineering effort produced a machine that confidently learns nothing.
2.4.4 Agents built on a founder's face backfired.
SaaStr built named agents on captured likenesses of Jason Lemkin and Amelia Lerutte, and they ran 1.5 million sessions of brand exposure tied to two individuals. SaaStr now advises against it, because likeness rights are legally unresolved and the knowledge of how to run it sits with a single employee. Its own sales lead logged into the platform twice in ten months.
2.4.5 Mass personalization stopped being an advantage.
Deep research separated serious senders only while it was expensive, and agents have driven that cost to zero, so the market now pools on the minimum signal. The penalty is already live: Chapter VII prices unreviewed AI mail at roughly 2.7 times the human spam-flag rate, convergent across the Saleshandy and Digital Applied tests. In practice, mass personalization collapses to one template plus merge fields anyway.
2.4.6 Build what encodes your knowledge, not what babysits the model.
Your customer-profile files, offer library, voice profile and suppression store encode what you know, and they compound under whatever model runs them. Prompt chains and code that compensate for what today's model does badly are an expense with a three-to-nine-month life, because a model release deletes their reason to exist. Anthropic did this to itself: it built three such compensating layers, including forced context resets for models that rushed work as the window filled. It deleted all three in March 2026 when Opus 4.6 stopped doing that, which turned weeks of its own engineering into maintenance nobody wanted. Boris Cherny, who runs Claude Code there, puts the gain from hand-built scaffolding at 10 to 20% and says the next model wipes it out.
2.4.7 LinkedIn deletes vendors rather than slowing them down.
In March 2026 LinkedIn removed the automation vendor HeyReach's company page and its founders' personal profiles while the software kept running. It had already removed the pages of the data vendors Apollo.io and Seamless.AI in March 2025. And in July 2025 it litigated the scraping API Proxycurl out of existence at roughly $10 million in annual recurring revenue. You do not lose here slowly, you get deleted, and where LinkedIn is also your distribution it takes the audience with it.
2.4.8 SaaStr wishes it had bought instead of built.
SaaStr built its own AI VP of Marketing, an agent it named 10K, ran it for five months, and concluded it wished it hadn't. It set a rule of buying 90% and hand-building at most 10%. Asked whether it had replaced a VP of marketing, the agent gave its own answer. It was a dashboard, a database, some scheduled jobs and a cheap model, glued together with six weeks of code.
2.4.9 Every failure here comes from scoring effort, not ownership.
Each of the eight cases above failed in one of two ways, and the four heaviest builds are no exception. One way is sitting on ground someone else could change whenever they liked. The other is betting on the volume half of the business at the exact moment volume stopped being scarce. Both are failures of scoring rather than of effort, which is why the two questions come first.
2.5 Find out why results fell before you react.
Five different things look the same from your dashboard: saturation, platform policy, buyers adapting, slop, and the economics flipping. Each calls for the opposite response: saturation says move, policy says comply, and buyers adapting says the ground is permanently gone because it runs on memory. That last one is the only change you cannot reverse, so the diagnosis is worth more than the speed of your response.
2.5.1 Saturation follows five steps you can forecast.
The tooling gets cheap and reaches everyone, volume explodes, the channel saturates because attention never grows to match it, the platform responds, and buyers adapt last. Run those five steps forward on any new channel and you get an estimate of its runway before you commit build time.
2.5.2 Three signs tell you the edge is gone.
The first sign is a $100 million funding round landing in the tactic's vendor category, which prices the opportunity in rather than proving it. The second is a paid curriculum, which means the method is now common knowledge. The third is the gap between top-quartile and average results widening, which marks the shift from works-by-default to works-only-with-craft. All three are cheap to watch, so watch them continuously.
2.5.3 Better execution buys weeks against a policy change.
Once a platform decides a behavior is unwanted, doing it better buys you weeks and nothing more. The tell that enforcement has moved from punishing users to punishing suppliers is a vendor's own page disappearing while its software keeps running. That is what happened across the LinkedIn automation category in March 2026.
2.5.4 Check the primary source before repeating a ban.
The claim that Google banned inbox warmup in February 2024 is widely repeated and simply wrong. The actual event was an API policy notice in November 2022 with a February 2023 compliance deadline. That deadline is what shut down the warmup network run by the email tool GMass, in February 2023, at 1.3 billion emails across 236,084 accounts.
2.5.5 A buyer who learned to ignore you stays gone.
Platforms can loosen their rules again and a saturated channel can empty out, but a buyer who learned to ignore you does not unlearn it. Chapter I tracks the Gartner series on buyers who would rather buy with no rep at all, and that share has doubled since 2020 without once moving back.
2.5.6 Put a human in the loop on the step that sends.
A human in the loop is a person who reads the output before it goes anywhere, and the slop cases all lack one. An agent crawls a website, infers the value proposition, then the ideal customer, then the positioning, then picks the competitors and writes the email. That is five generative steps chained with nobody reading in between. Kyle Norton, CRO of Owner.com, supplies the mechanism: each step compounds the loss from the last, so the output is slop by construction however good the model is. Count the steps in your own pipeline that nobody reads, and put a person on the last one.
2.5.7 Pay-per-meeting outbound has almost no margin left.
A booked meeting costs $2 to $10 of sending infrastructure at the funnel rates Chapter I reports, roughly 100 to 500 sends at about $0.02 each. A qualified booked call sells for $150 to $300, a price defending the research, reply handling and scheduling behind every delivered meeting, not servers. Never price your own offer per meeting, because you would be racing sellers who keep margin by automating those human steps out. Treat a cheap bought meeting as priced to be a bad one, since the seller's margin comes from whichever human step they cut. That incentive reading is an inference from the arithmetic, not a measured audit.
2.6 Trust dated reversals more than published benchmarks.
Almost every benchmark in this field is funded by a vendor. So the best evidence available is someone saying on a specific date that they stopped doing something, and why. Saying that costs the speaker something, and that price is what makes it worth more than a benchmark.
2.6.1 Nearly every published benchmark is sold by a vendor.
Almost every published outbound benchmark comes from a company selling the thing it measures, which leaves dated changes of position as the best decay evidence anyone has. Log them with the date and the speaker's exact words, because averaging them into a trend line throws away the only reliable part.
2.6.2 Jason Lemkin retired his own data-moat claim.
Jason Lemkin of SaaStr used to call a twenty-million-word archive a competitive moat, and he abandoned the claim in public with the verbatim words "I got that wrong." His reasons: the value decays sharply after a few months, and thirteen years of contacts are nearly useless because people change jobs. What replaced it is about 20,000 recent words. Separately, he ran more than twenty agents for thirteen months. And the first agent-written email he judged better than a human's came only two weeks before he said so.
2.6.3 A vendor relationship downgrades every number the operator publishes.
When an operator sells access to the vendors whose tools produced their results, every performance figure they publish is self-reported with nothing independent to check it against. Apply exactly the same discount to your own vendor relationships before you quote your own numbers to anyone.
2.6.4 Clay bans AI writing internally while selling it.
Varun, a co-founder of the go-to-market data platform Clay, says language models revert to the mean. And he bans them from Clay's own marketing writing while the product sells AI-written outreach at scale. Both positions are coherent, because averaging is fine for research and fatal for differentiation. The internal rule is the more honest signal, so weigh it above the product page.
2.6.5 The human step Anthropic removed was on buying, not on sending.
Anthropic dropped the human approval step from enterprise buying in January 2026, and by May self-serve was taking 54% of new enterprise logos. That is the sharpest documented removal of a human checkpoint in this research base, and it came off the path buyers take to reach you. The send gate on the path your messages take to reach them stayed exactly where it was.
2.6.6 Two unsettled questions could flip the conclusions in this chapter.
The first is whether the human approval step before sending becomes a shipped platform feature or gets quietly optimized away. Chapter XII names the only three gate types found anywhere in the send path, and none was put there as a guardrail, a control whose job is stopping harm. Each is an accident of billing, rate limiting or scheduling, which is why the question is open. The second is whether inbox placement for low-volume senders degrades for two more quarters. If it does, the low-volume design recommended here stops working, and the weight shifts to LinkedIn, video and inbound-led campaigns.
2.7 Move 1 sorts on two questions, not difficulty.
Move 1, the first of this guide's standing moves, keeps what still works when everyone does it, sorted on two questions: can this be mass-produced, and who owns the ground it stands on. The eight failed builds above show what scoring effort instead costs, and the four heaviest sit inside that eight. The same sort earns the move that says pick a list nobody else can build and the one that says test the offer rather than the wording. That is why the conference-list method and the two-to-threefold offer effect are on the durable list. Skip it and both moves become preferences, and a rented 300-mailbox estate still looks like a moat.
Count from your decision, not from your first email.
Every base rate you will read is measured from the first email you send, but your decision came months earlier. Hold every stakeholder to the clock that matches your build: 120 days from the decision on a fleet, many domains sending at volume. On one or two mailboxes it is the buyer's own window plus a setup measured in days. A campaign is one person running one list against one offer through one channel, and every volume figure here is per campaign. An operator is the named human accountable for that campaign and for checking what its agents produce, and every capacity figure is per operator.
3.1 Month three is really week eight.
This trap belongs to the fleet build, many domains sending at volume. There, warmup, the slow ramp to a trusted sending domain, plus list building and setup eats the first month. So a program killed at "90 days" has usually only been sending for sixty of them, and the verdict lands against the wrong start line. On one or two mailboxes you are live the week DNS propagates, and the clock that governs you is the buyer's window, not your ramp.
3.1.1 The 90-day rule counts from your first email, not from your decision.
The 90-day figure measures how long buyers take to decide once your first email reaches them. Count it from the day you decided to do outbound instead, and day 90 arrives while the buyer's clock has barely started.
3.1.2 Quote the clock that matches your build.
The reports put the honest decision-to-pipeline figure at 110 to 130 days, with 120 as the round number worth quoting. The gap between 120 and 90 is the month a fleet spends on warmup, list building and setup before anything lands. On one or two mailboxes that month collapses to the one to three days DNS takes to propagate, because your real sending is the ramp. Commit stakeholders to the figure that matches your build, or you will spend day 90 defending a program that has been live for sixty days.
3.1.3 The 90 days belongs to the buyer.
The 90 days is what remains of the buyer's own journey once a seller makes contact, not a rule about your cadence. Chapter I sources that journey at 10.1 months with first seller contact about 61% of the way through, leaving roughly 3.9 months, about 118 days. The quoted 90 is that remainder rounded down by nearly a month, and the rounding runs against you. Nothing you do to your sequence shortens either version.
3.1.4 Warm outreach does not get 90 days.
Working an existing network of 50 to 100 names produces a first client in about 30 days. And one sweep of a first-party list, meaning contacts you collected yourself, pays in days. Give either one a 90-day patience budget and you will sit on a visible failure for two extra months.
3.1.5 A 90-day decision date kills the slow tactics on schedule.
Publishing takes 12 to 18 months to move customer acquisition cost. A champion register, meaning past buyers you track across job moves, takes roughly a year to start firing. That figure is an estimate rather than a measurement, while partner-sourced introductions have no measured payback time anywhere in the research. At day 90 all three look exactly like failure, so the review date has to match the tactic rather than the calendar.
3.1.6 Day 90 is a written diagnosis.
Ninety days of one campaign buys roughly seven meetings booked and four or five held, a floor this chapter derives below. No rate computed on four or five events is a verdict about anything. The day-90 decision is therefore a written diagnosis of the list, the offer, and deliverability, meaning whether your mail reaches inboxes at all.
3.1.7 Two clocks run at once, one quarterly and one yearly.
The tactics that pay inside a quarter are the ones you already have access to. Those are your warm network, your first-party list, and a deep-and-narrow cold campaign, all of which cost judgment rather than budget. Publishing under a named human takes 9 to 12 weeks to produce sourced pipeline and 12 to 18 months to move customer acquisition cost. Chapter VI carries both figures with the warning that every source behind them is self-reported vendor or blog material, so anyone promising faster is overselling.
3.2 Start publishing on day 1, because you cannot buy back a lost year.
Nobody in the 65 research reports or 209 practitioner transcripts has isolated how much publishing lifts anything, which makes it the softest evidence in the whole book. The case for starting rests on cost, not on effect size. It is an hour a morning to start now, against twelve months of audience that no money buys back later.
3.2.1 Publishing is the one thing you cannot start later.
Every other asset here can be bought or rebuilt on demand: a domain purchased in October is warm by November. An audience you did not start in October is still twelve months behind in October 2027. That gap is why publishing under one named person's real byline, not the company's, cannot be postponed.
3.2.2 The publishing bet costs one hour a morning.
The worst case is one named person spending an hour each morning. That writing feeds the proposals and cold emails the person was going to produce anyway, so the hour is not lost even if the audience never arrives.
3.2.3 The shortlist is written before you speak.
Buying groups rank vendors before they speak to any of them, and the eventual winner is nearly always already on that first-day list; Chapter I sources the finding. Talking your way onto the list afterward almost never works. Publishing is the only thing you control that writes you onto the list before it gets written, which is why both start on day 1.
3.2.4 An audience you own keeps working.
An audience you own is the only asset that keeps working whichever way the next few years go. Platform-average reply rates have roughly halved since 2019 as the channel saturated, a decline Chapter I dates. The owned audience survives a market that splits into a careful tier and a spam tier, and survives buyers whose agents screen the inbox for them. It also survives the filters winning outright, because nobody else can switch it off.
3.2.5 Nothing compresses the trust clock on a domain.
Buy your sending domains today, because the trust clock starts at purchase and nothing shortens it. It runs 2 to 4 weeks on an aged domain, 4 to 8 on a fresh one. On one or two mailboxes your real mail at five to ten a day is the ramp. On a fleet the same weeks run as artificial warmup before any live send.
3.2.6 Buy domains and start publishing on the same day.
Buy the domains today to start their trust clock, and start publishing today because an audience is an asset you cannot buy back. The two reasons are unrelated, and confusing them is how operators end up warming domains while deferring the writing. Both start Monday morning.
3.2.7 Treating publishing as the thing you get to later is exactly backwards.
Warmup feels urgent and publishing feels eventual, which rushes the thing you could always redo and defers the thing you cannot. The bill arrives about twelve months out, when the cold channel has decayed further and there is still no owned audience to fall back on.
3.2.8 Launch one new agent at a time per operator.
Every new agent costs at least two weeks of configuring, checking and correcting, during which the operator's existing agents go unwatched. Chapter IX sources that cost and turns it into the capacity ceiling. Launch one at a time, because launching several at once quietly degrades a working campaign while the dashboard says you are expanding.
3.2.9 One warm conversion in month one is correct.
Track effort and results on two separate lines rather than one. For a single campaign, the honest month-one results line holds one warm conversion, meaning one person who comes back and actually wants to talk. It also holds some comment activity on other people's posts.
3.3 Give every bridge tactic an expiry date.
A bridge is any revenue tactic you run to survive the lag before the durable machine works. Every bridge either consumes a finite asset you own or rents distribution somebody else controls, and both run out. Write the date and the exit condition before the first day, because a freelance marketplace and a partner channel fail in exactly the same way.
3.3.1 Nobody fails by picking the wrong bridge.
In the practitioner accounts collected here, the documented failure is never leaving a bridge, not choosing a bad one.
3.3.2 Discount for proof, then stop at three.
Every name in your warm network converts exactly once, so spend them deliberately. Cut price in exchange for a testimonial until you have three delivered case studies, then stop.
3.3.3 Put a date on borrowed distribution.
A solo operator's version is a freelance marketplace such as Upwork, where the job-success ranking that earns you work belongs to the platform, not to you. Bogdan, a builder named only by first name in the source, sets his on-platform rate at $44 an hour. He also sells a paid AI training program, AI First Academy. A larger organisation's version is partner and channel pipeline. Both vanish the day you stop feeding them, which is why both need a written exit month.
3.3.4 Reactivate your own list once, then stop.
A first-party list you already own pays fast, and one sweep can produce revenue in days. But warm lists want 2.5 to 3 months between contacts and the second sweep yields materially less. So treat reactivation as a one-time spike rather than a campaign you can run on.
3.3.5 One cold campaign has a fixed meeting floor, and volume never lifts it.
A tightly targeted cold campaign books 2.25 meetings a month and holds about 1.5, which is 6.75 booked and 4.5 held across a quarter. The arithmetic starts from one campaign sustainably producing 250 to 300 sends a month at any headcount. Take 300 of those sends at a 3% positive reply rate, counting only replies that want to continue, and convert them to meetings at 25%. That lands on 2.25 meetings booked per campaign per month, with cancellations and no-shows producing the lower held number. The 3% is the band good researched campaigns hit, well above the pooled rate at volume, the average mass blasts achieve, and the whole chain is one report's arithmetic on vendor bands, not a dashboard reading. Any plan that needs the booked number to grow is a plan to buy more domains, the wrong purchase; the next section names the right one.
3.3.6 Grow by campaigns up to two per person, then by people.
2.25 meetings booked a month is a constant per campaign rather than per company, and pushing more volume through the same list and offer does not move it. The floor rises campaign by campaign. One operator comfortably runs one campaign and can carry a second, because a campaign spends about half of the daily checking budget. Past two, add a person with their own lists and offers. Stacking more agents behind the same operator leaves the number where it is, because the limit is the human clearing drafts.
3.3.7 Pay for reach until month 9, then stop.
Paying to put your published work in front of a named ideal customer profile (ICP) does one job. It buys back part of the 9-to-12-week wait between publishing and sourced pipeline. Once your own posts reach those same people at the volume you were paying for, the spend buys something you already have. Stop at month 9, or earlier if organic reach among your target titles has matched the paid volume for two months running.
3.3.8 Cap subcontracted delivery at half your revenue.
Selling delivery hours into another firm's existing pipeline gives you revenue at zero acquisition cost, which makes it the highest-quality bridge available and also the least evidenced. Every report sizes it as a market, and not one measures whether it brings you customers of your own. Cap it at 50% of revenue, because past that the dependency is exactly the risk you took the bridge to avoid.
3.4 Pre-commit the kill date and the trigger.
Panicking at week six and nursing a corpse at month twelve are the same failure. Both decide by feel instead of by a guardrail, meaning a limit you set in advance while you were calm. Set the date and the trigger before the first send.
3.4.1 Deliverability is the kill switch that overrides every other rule.
Bounce rates above 2 to 3% mean stop sending immediately: this is a kill switch, not a tuning dial. Spam complaints above the 0.3% provider cliff Chapter I dates carry the same instruction. It is the only trigger that fires regardless of what the rest of the dashboard says, because past it you are destroying a domain, not testing a message.
3.4.2 Under 1% reply after 200 sends condemns one list and offer, not outbound.
A sequence still under a 1% reply rate after 200 sends is finished, but 200 sends only condemns that one combination of list and offer. The 1% counts every response, not just the positive replies the meeting floor is built on. The wide gap between the numbers in this guide comes from the two measurement regimes Chapter I separates. Change the list slice or the promise and run it again. The two wrong moves are concluding outbound is dead and scaling the losing version in the hope volume rescues it.
3.4.3 Zero qualified meetings at day 60 is a list verdict.
At the floor this chapter derives, day 60 should show four or five bookings and about three conversations that actually happened. Arriving with zero is a gap too large to argue about at the edges, so read it as a targeting failure and repair the list rather than the copy.
3.4.4 Judge the LinkedIn motion after the first cohort completes it.
The manual motion warms each prospect with two or three genuine comments over one to two weeks before any direct message. The earliest honest read therefore comes once your first cohort has completed that loop and had time to reply. The claimed 2-to-3x reply lift from commenting first is vendor A/B data, directional rather than proven. Ignore the four-week account-warmup clock: it comes from automation vendors whose tools breach LinkedIn's terms, and LinkedIn banned HeyReach's company page in March 2026. An aged personal account worked by hand has no warmup clock; a brand-new one is throttled to 20 to 50 invitations a week, which the follow-first design in Chapter XVI never touches.
3.4.5 At week six content looks dead and is on schedule.
First inbound signals are documented at three to six weeks, so week six sits inside that band and still reads as nothing. Kill it there and you pay the full twelve-month cost for a program performing exactly to plan, which makes this the most expensive error available.
3.4.6 Judge content at month 9, and judge the topic.
Month 9 sits after sourced pipeline from publishing should have appeared and before acquisition cost could have moved, which makes it the earliest honest review point. When you review, change the topic and who you are writing for, and leave the cadence alone, because killing the habit is almost never the right response.
3.4.7 Do not buy third-party intent data.
Buy none by default: across the 65 reports there is no independent evidence that a purchased intent score predicts a B2B purchase. The one head-to-head in this base had signal-only targeting lose to plain fit, and Chapter IV ranks it the weakest-evidenced spend in its category. That is absence of proof, not proof of uselessness, so a feed you already pay for gets a test rather than a termination. It must beat your untargeted cold results at month 2 and every two-month checkpoint after, and it never gets the annual contract that removes your ability to cancel. Chapter XVI covers the signal worth having instead: what a named buyer said in public.
3.4.8 Never sign an annual contract for an AI sales rep.
The 50 to 70% buyer-churn figure attached to AI-SDR products, meaning software sold as an automated sales development rep, has never been audited. Chapter IX traces its attributions and the two time windows that cannot both be right. A better-sourced measurement sits beside it: about half of AI-SDR pilots are shut down within 90 days. Keep the break clause at day 90 and exercise it, because you have already paid roughly two weeks of operator onboarding for every agent you bought.
3.5 Name the tactics with no honest early indicator.
For a few tactics nothing observable early predicts the outcome, and inventing a proxy is worse than admitting the gap. For those, patience plus a date you committed to in advance is the method itself.
3.5.1 Day 7: read fifty rows of your list by hand.
Open your list, read fifty rows, and count how many companies you would refuse to sell to: thirty minutes of work for a direct read on list quality. The list moves meetings per prospect about a hundredfold, which Chapter IV prices, so in a week when nothing else is knowable this is the read worth having.
3.5.2 Day 30: sort your replies by reason.
Sort your twelve to twenty-four replies into categories: wrong person, no budget, no pain, bad timing, already have a solution, and the "what is this?" reply. Watch the wrong-person rate, because it separates a list problem from an offer problem, and those two have opposite fixes.
3.5.3 Day 60: the number that counts is meetings held.
A booked meeting is a forecast and a held meeting is an event, so rest the decision on held; Chapter XII carries the counting rule. The floor derived earlier loses about a third of its bookings between the calendar invite and the conversation. Inbound loses far less, about 6.5% against the 25 to 35% cold outbound loses (RevenueHero and Ziellab). So a booked count read on its own flatters a cold program more than a warm one. Report both counts, and let held carry the verdict.
3.5.4 Beyond those three reads, the first 90 days are blind in both directions.
You do not have enough sends for statistics and you have not waited long enough for outcomes, so everything else on the dashboard in that window, past the day-7 list, day-30 reply and day-60 held-meeting checks, is a proxy. Decide which proxies are honest before you start reading one, because you will read whatever is put in front of you.
3.5.5 Three tactics leave no early trace you can read.
Publishing in months one through six produces nothing readable. The referral engine has no base rate anywhere in the 65 reports or 209 transcripts, so there is no normal to compare yourself against. Brand memory, whether the buyer thinks of you in month six, leaves no trace until they contact you. By then it is an outcome rather than an indicator.
3.6 The clock you defend at day 90 is written before day 1.
Two of the eleven decisions at the front of this guide, the pre-launch calls made before any sending, are settled here. The clock starts when you buy the domain, and publishing under one person's real name starts on day one. Ignore them and, on a fleet build, many domains sending at volume, you will defend a program at day 90 that has been sending for sixty. You will judge slow tactics against the fast tactics' calendar. And you will treat publishing as the thing to start once cold email is working, when it is the one asset that cannot be bought back later.
Fix things in this order: list, offer, who it comes from, wording.
The four levers are not equal, and you do not work them at the same time. Most of the pages here go to the list, because until the list is right every improvement to the other three is rounding error. The offer, the sender brand and the wording each get their own part after this one.
4.1 "List beats message" names the winner and hides the runner-up.
The slogan gets the winner right but hides the full ranking: list, then offer, then sender brand, then the words. Nearly every published playbook spends its pages on the words, and the study that made the slogan famous never gave the offer a line of its own.
4.1.1 The "list beats message" study buries the offer inside "message".
Meetings per prospect differ about a hundredfold between the best and worst lists, measured by the outbound agency RevPack across 214 of its own campaigns. Reply rate swings only 28% between the best and worst wording sent to the same list. The source's own categories put the offer inside "message," so the 28% measures wording with the offer held constant, and nobody measured the offer separately at all.
4.1.2 Rank the levers by which effects your volume can actually detect.
List and offer produce effects big enough to see at a few hundred sends a month. Sender brand and message move the numbers by less than your own month-to-month swing, so a real improvement there is indistinguishable from a good week.
4.1.3 When practitioners blame copy, they mean the offer.
Practitioners who blame the copy are, on inspection, changing the offer. Nick Saraev, who sells a paid cold-email and automation course, blames copy for 50 to 60% of campaign failure. Yet across four hours of his course material, every teardown fixes the ask, the promise or the proof, which is the offer, and never the list.
4.1.4 Sender brand ranks third, and saying so costs Sam Blond sales.
Sam Blond, CEO of Monaco, which sells message automation, ranks who you are above what your sequence says. The ranking costs him sales, which makes it the best-incentivized claim in this order. It is also the least isolated, because nobody has run the experiment separating sender identity from everything else in the send.
4.1.5 Fix the four levers in order and stop at the first break.
Work down the order and stop at the first broken lever, because improving anything downstream of a broken list buys nothing. Payback differs about tenfold: a better offer shows in weeks, a better list at about 90 days, a published presence in 6 to 12 months. The list is done when 150 to 200 contacts in one segment return 4% or better positive reply, a researched-band number, from small hand-researched campaigns, rather than an agency-at-volume one. And it is done when wrong-person replies are no longer your largest reply category. Until both hold, offer changes are unreadable.
4.2 Fix the list first: it moves results a hundredfold.
The list gap is the hundredfold one; the writing gap is about double.
4.2.1 214 campaigns showed a hundredfold spread from the list.
RevPack's measurement covered the 214 campaigns across roughly 250,000 prospects in July 2026, and the list was the variable that moved. Booked, not held, meaning meetings scheduled rather than actually attended, is the unit RevPack reported, which flatters cold campaigns, so read the ratio and discount the level. Their own summary of writing: it moves reply from 0.2% to 0.4% and never takes a wrong audience to a booked meeting. Those are agency-at-volume levels you compare only with each other.
4.2.2 The list is the message.
Jordan Crawford's inversion names the mechanism behind the hundredfold spread: start from the customer's pain, find the data that proves it, and let the data write the message. Who receives the message decides what it can honestly say and which offer you can credibly make, hence "the list is the message" (Blueprint, 2026, quoted approvingly by Kyle Norton, CRO of Owner.com). Adopt the philosophy but discount the metrics: his honest calibration is about 2% positive reply to start, and 4 to 5% with a well-built value proposition. His published 11 to 32% is demo work Chapter I quarantines. Tibor Shanto argues the opposite school, targeting goals rather than pain to reach buyers before pain is felt, and the evidence does not resolve the two.
4.2.3 Nick Saraev sells personalization and still ranks the list first.
Nick Saraev makes his money teaching the scarily specific personalized emails he ranks second. Yet he puts list exclusivity above that personalization, above the offer and above follow-up cadence. That ordering costs him the product he sells, so it outweighs any benchmark. It also contradicts his claim that copy causes 50 to 60% of failure, quoted earlier: one man, two positions, and this is the one to follow.
4.2.4 Narrowing the profile moved reply from 2% to 11%.
One campaign documented by Builtforb2b tightened "all SaaS companies" into "Series B SaaS companies using Salesforce, 50 to 200 employees," and reply went from 2% to 11%. Both levels come from a single researched campaign, so read the swing and not the level. No wording variable anywhere in this evidence base swings that far, and this one is available before you write a single sentence.
4.2.5 The list is the only lever your volume can see.
The list is also the only lever your volume can measure, a second and separate reason to fix it first. Split 20 to 30 conversions across three quality bands and "the best band converts three times better" is hard to tell from random variation.
4.3 Build the segment from outcomes, then prove it by sending.
The rigorous method rewinds your closed deals for shared attributes and confirms the pattern on data you held back, which takes more closed deals than you will have. The honest fallback is to define the segment from outcomes and validate it by running outbound at one segment at a time.
4.3.1 A database filter is not a customer profile.
"50 to 1,000 employees in vertical X" describes how a contact database chose to index the world, because headcount and vertical are the vendor's filing system. They say nothing about which of those companies is in pain this quarter, and pain is the only thing that predicts a reply.
4.3.2 A usable segment adds up to a problem.
A usable segment stacks two to five conditions that together imply something is going wrong for the buyer right now. The blunt test is whether you can describe the same onboarding path for eight out of ten customers in it. If you cannot, the segment is too broad to write one offer against.
4.3.3 Test the pattern on deals you deliberately held back.
Fit the pattern on 80% of your deals, then test it on the 20% you never looked at, because the holdout is what makes the rewind honest. Jordan Crawford built a deliberately meaningless signal from the vowels in company names to show that a holdout catches nonsense. It showed 2.4x lift on the deals it was fitted to and nothing, 1.0x, on the held-back 20%. The real signal beside it, meanwhile, held up at 4.1x fitted and 3.6x on the holdout. Both figures are Crawford's own, from a constructed example rather than a client campaign, so take the method and treat the sizes as illustration.
4.3.4 The holdout needs 20 to 30 closed-won deals.
Below roughly 20 to 30 closed engagements the statistics stop working. A pattern found on a sample that small mostly describes the deals you happened to win, and it will not repeat on the next ten. Under that threshold, do not run the analysis, and do not treat its output as knowledge.
4.3.5 Run one segment at a time, 150 to 200 contacts.
Hold the message constant and change only the segment; two moving variables produce a result you cannot read. Budget 150 to 200 contacts before the response data means anything. Kill a segment below roughly 4% positive reply, the researched band, after 150 or more clean touches. Clean touches are delivered emails to verified addresses at companies that passed your disqualifiers. The 200-send version kill in Chapter III condemns the copy on all replies; this kill condemns the segment on positive replies only, and the segment verdict wins when they disagree.
4.3.6 Target founders first and VPs last.
The outbound agency Belkins measured reply rate by seniority across its 2025 send data, 7.5 million sends: founders 0.57%, C-level 0.42%, VPs 0.32%. Those are agency-at-volume levels, so take the ordering and not the levels. Response falls with distance from the founder rather than with rank. That inverts the folklore that the higher you aim the harder it gets, so re-point your targeting today.
4.3.7 Feed the model your losses too.
An attribute shared by all your customers is worthless if your disqualified leads share it too. So give the model both lists and make it hunt for the contrast.
4.3.8 The "68% higher win rate" figure is a 2019 survey.
The number behind most ICP sales pitches comes from TOPO's Account Based Benchmark Report of January 2019, a survey of roughly 150 self-selected practitioners. People credit it to Forrester or SiriusDecisions, and neither produced it; Gartner's credit is at least defensible by descent, since Gartner acquired TOPO in 2020. Narrow targeting survives on other evidence, so lean on that and stop quoting this number.
4.3.9 Discard any AI profile that cannot cite its attributes.
Ask a model to build your customer profile and it hands back your own assumptions in fluent, confident prose. Require a source for every attribute, meaning which deal, which page, which row, and discard anything that cannot produce one.
4.3.10 Nobody has automated the judgment part of segmentation.
Amelia Lerutte, chief AI officer at SaaStr, runs 20 to 30 agents and vibe-coded apps in production, counted in February 2026, and still does segmentation by hand. It "is just all in my brain right now," she says, and "no one can create these types of segments autonomously." That is one person speaking rather than two operators agreeing. She holds nine segments against an estimated need for about a hundred, and that gap is the clearest map of where the automation frontier sits today.
4.3.11 A tenth agent adds nothing; a second person does.
How many segments a campaign can hold is capped by the human who checks each agent output. The ceiling is the 5 to 10 checked outputs per operator per day that Chapter IX establishes. A tenth agent pointed at the same human buys nothing, so a second operator is the only way to raise the ceiling.
4.4 Source your list where competitors cannot reach.
Every credible high reply rate in this evidence base attaches to a list nobody else could buy or filter. Exclusivity, rather than accuracy or size, is the list property with real numbers behind it.
4.4.1 The only 15 to 20% campaigns were conference lists.
Nick Saraev, who sells scraping and automation education, says his 15 to 20% reply campaigns were all conference attendee lists he got from the organizers or bought. Those rates sit far above any agency-at-volume level, and those are lists of people who do not publish their contact details on the internet. Grade it n=1, Grade C: one operator, self-reported, nothing independent behind it, the weakest grade any claim here is allowed to carry. Take the mechanism and treat the percentage as illustration.
4.4.2 Emailing the conference organizer costs one email.
Asking a conference organizer for the attendee list, the highest-reply list source in this guide, costs you one email.
4.4.3 Anything you can filter, your competitor already filtered.
A filter available to you is available to everyone, and your reply rate is partly measuring how many of them already used it. Saraev reports a community member running the same campaign in English at 0.5% reply and in Czech translation at about 10% positive. Market saturation was the only variable that changed. The pair mixes metrics, all replies against positive-only, and it is one self-report with nothing independent behind it, though the mechanism itself is not in dispute.
4.4.4 A technology filter finds the companies already running what you complement.
Technographic data, the record of which products a company runs, is a fit filter inside the databases you already buy; Apollo exposes it as a search facet. It answers a question nothing else in this sourcing list answers directly: who has already bought the thing your offer attaches to. Author's note: I built a working US list this way, filtering for companies that run Claude, exactly the segment a Claude implementation service sells to. Two caveats: anything you can filter your competitor can filter too, and vendors differ in how they detect installed technology and how stale the reading is. So read a sample of rows by hand before trusting the column, use it as the segment seed, and earn exclusivity back through the research and the offer.
4.4.5 Trade row quality for scarcity on purpose.
Lists nobody else can reach come with messier, less complete rows than a polished database export. Take that trade deliberately, and cover it with harder verification before sending, because bounce is the one cost that damages the domain rather than the campaign.
4.4.6 Of eleven sourcing routes, one gives rows you own.
The eleven routes are job boards, podcast and conference speakers, communities, review sites, app marketplaces, technographics, regulatory filings, funding announcements, LinkedIn engagement, newsletter sponsors, and web-crawl signatures. The last means traits you detect by crawling that no database catalogs. Web-crawl signatures are the only route that produces a dataset that is yours rather than rented.
4.4.7 Sort sourcing routes by who owns the resulting row.
Sort sourcing routes by how many rows they yield and you reproduce the ranking every competitor already has. Sort them by whether the row is yours or licensed and the order changes completely, because ownership determines whether the list still works next quarter.
4.4.8 Mine your sent mail before buying a single row.
Your sent folder holds conversations that ended in a commitment nobody followed up on. That list costs nothing and carries no deliverability risk, because those addresses have already replied to you. It is first-party by construction, so exhaust it before you spend a dollar on data.
4.4.9 Small lists are a choice available at any size.
In Belkins' 2024 study across 16.5 million emails, campaigns under 50 recipients reply at 5.8% against 2.1% for campaigns over 500. Touching one or two contacts per company returns 7.8% against 3.8% for ten or more, though that second pair carries a weaker grade. All four are agency-at-volume rates from one corpus, so compare them only with each other. A 700-person marketing organisation can choose a list of 50 as easily as a solo operator can, so small lists are a deliberate setting, not a consolation prize.
4.5 Every enriched row is a guess with an expiry date.
An enrichment provider sells you a probability with a shelf life rather than a fact. Run the cheap, certain filters before you pay for guesses, and rent the multi-provider lookup instead of building one.
4.5.1 Never hold a standing enriched database.
SaaS and tech contact records go stale at 35 to 70% a year, against 10 to 25% for manufacturing and government. The blended 22.5% annual average everyone quotes describes a portfolio you do not have. Enrich at the moment you send, because a stored list rots while you plan.
4.5.2 An accuracy number tells you nothing about how many rows you get.
Accuracy grades only the rows a vendor returns; coverage counts how many rows you get. The contact database Apollo advertises 97% accuracy and measured 91.3% on a 5,000-contact benchmark in June 2026, so the accuracy claim is close to honest. The case against Apollo is coverage: 68.1% verified coverage, the share of requested contacts found and certified deliverable, with 6.5% false positives, addresses certified good that were not. The benchmark was run by Anymail Finder, an email-finding vendor that entered its own tool and published the full dataset. A rival vendor, Amplemarket, puts only about 96 million of Apollo's claimed 275 million contacts past Apollo's own verification filter, a C-grade figure from a company with a stake in the answer.
4.5.3 What a waterfall is worth depends on your alternative.
A waterfall, also called a cascade, queries provider after provider until one returns a contact, and its worth depends on what you would otherwise have bought. Against the best verify-first finder it buys one point, at four times the false positives, 4.0% against 0.9%, and a higher price per credit. On the one A-grade benchmark here, 5,000 triple-verified contacts in June 2026, the cascade product FullEnrich returned 87.1% verified coverage to Anymail Finder's 86.4%. A-grade means method, sample and full dataset published, with the conflict disclosed: Anymail Finder ran the test with its own tool as the single source compared. Against a contact database, a far lower bar, the gap is a whole category: cascades run roughly 80 to 95% coverage against 40 to 68% single-source, Apollo's 68.1% above supplying the 68% end. One 30-day test on 2,000 contacts put a waterfall at 78% against Apollo's 42%; the framings are unreconciled, but on hard rows FullEnrich reached 66% of rare contacts to Apollo's 29%.
4.5.4 Rent the multi-provider lookup, never build it.
Buy the cascade as a pay-on-success API and wrap it as a single tool that behaves identically every time your agent calls it. Stop at three or four providers, because the fourth recovers only 3 to 5% more contacts. And do not build your own until you pass 10,000 to 20,000 lookups a month, where the arithmetic flips.
4.5.5 Knock out bad rows before you pay to enrich them.
Run deterministic ICP knockouts first, meaning rules written in code that return the same answer every time rather than a model call. Make them fail closed, so a row with missing evidence is rejected rather than waved through. Stripping 30 to 50% of the noise before it reaches a paid lookup pays for everything else in this chapter.
4.5.6 Verification gives you a risk score, not a verdict.
The deliverability firm WizLeads put the same 1,000-address list through several email-verification tools in April 2026 and measured the real bounces. Results ran from 0.10% with Debounce to 3.70% with ZeroBounce, a 37-fold spread, against vendor accuracy claims that cluster at 97 to 99%. Treat the verdict as a risk score and suppress anything not returned as cleanly valid.
4.6 Score far less than the scoring vendors want.
Lead scoring is sold as a discipline and has never been shown to work. No published controlled study compares scored prioritization against simply working the list in order, so every lift claim in the category rests on the seller's own arithmetic.
4.6.1 No controlled study has ever tested whether lead scoring works.
No published controlled holdout study shows scored prioritization beating simple rules at any scale, and every lift number in the category comes from the vendor itself. The intent-data platform 6sense claims its top-scoring tenth of accounts converts six times better. The scoring tool MadKudu claims "92% accuracy, 50% lift," and the CRM vendor HubSpot claims "45% more qualified leads." None of the three says how it was measured, has anything unscored to compare against, or reports a margin of error.
4.6.2 You will never close enough deals to train a scoring model.
Microsoft's CRM, Dynamics, will not train its predictive lead score without at least 40 won and 40 lost opportunities from the past two years. Practitioners put the point where a model beats a good operator's intuition at 1,000 or more closed deals. A high-research campaign books about 2.25 meetings a month and holds roughly 1.5 of them, on Chapter III's arithmetic. Neither count reaches either threshold inside any planning horizon you have, so a model trained on your own outcomes is not available at this volume.
4.6.3 Never trust a model's 1-to-100 fit score.
Three findings stack against it. Ask the same model twice and the answer moves about 1.2 points on a ten-point scale, which is 12 points out of 100. Models never use the bottom 60% of a hundred-point range, clustering almost everything at 90 and 95. And attributes scored together rise and fall in lockstep: Stureborg and colleagues measured a 0.979 correlation, where 1.0 would mean identical. Human raters on the same attributes move nearly independently at 0.315.
4.6.4 Replace the score with knockouts and cited checks.
Use 5 to 15 hard disqualifiers that fail closed on missing evidence, then 5 to 10 yes/no questions that each require a cited source. Compute all weighting and addition in code, because any number a language model produces is a guess about what a number would look like.
4.6.5 Your nine scoring axes are really about four.
The go-to-market engineer Garrett Wolfe documented the failure directly in June 2026: scoring axes that look independent move together, so nine signals behave like about four. And 91% of his accounts landed in the same 10% band. A score that has stopped spreading accounts out has stopped prioritizing anything, and it fails without any visible signal.
4.6.6 A tier that changes nothing is decoration.
Define three tiers at most and attach a truly different approach to each, because a tier without its own approach is a colour-coding scheme. If your volume lets you research every account on the list anyway, skip prioritizing entirely. Predictive platforms such as 6sense and Demandbase list at $12,000 to $80,000 a year, so put that money into research compute instead.
4.7 Use signals as a tiebreaker, not as the filter.
A buying signal tells you when to contact someone who was already worth contacting, and used to decide who, it measurably does worse than plain fit. Almost none of the signals worth having need to be purchased at all.
4.7.1 Signal-only targeting lost to plain fit.
In RevPack's own comparison, campaigns selected purely on signals against loosely qualified lists replied at 0.66%, while campaigns selected purely on fit replied at 0.79%. Both are agency-at-volume rates from the same corpus, which is what makes the comparison fair. Signals helped only where they confirmed a fit that already existed, the rare data-backed dissent to signal enthusiasm in this evidence base.
4.7.2 Fix fit first if under 70% of wins score high.
Run your closed-won accounts back through your current fit model. If fewer than 70% of them land in the top two tiers, the model does not describe your business. No timing signal layered on top of it will rescue the campaign.
4.7.3 Let a champion's job change fire on its own.
Someone who bought from you before and has moved to a new company is the only signal here supported by more than one source. The standard advice requires two signals inside thirty days, which produces almost no triggers when you are watching only a few hundred accounts. So let this one fire alone.
4.7.4 Third-party intent data carries the weakest evidence in the category.
The data vendor DemandScience surveyed 750 B2B marketers in 2026: 91% use intent data, 24% report exceptional returns, and 87% call their own signals unreliable. That is the same population saying all three things at once.
4.7.5 Watch job changes, funding and hiring yourself, then drop the paid feed.
Executive job changes, funding events and hiring are the best-evidenced signals, and an agent can watch all three from public sources for close to nothing. Three others you structurally cannot replicate: a data co-op reporting spikes in what a company reads across a pooled publisher network. The second is review-site browsing that stays private to the platform. The third is bidstream data from the ad auctions a company's employees pass through. Those three carry the worst evidence in the category, so let 6sense's roughly $59,000-a-year median platform go.
4.7.6 The honest warm-over-cold lift is 2 to 4x.
Warm contacts outperform cold ones by roughly two to four times, not the ten to fifteen times printed on vendor marketing pages. The widely cited study claiming these tools name 82% of anonymous website visitors was commissioned and paid for by the vendor it ranked first. It went out on a paid PR wire with a fabricated media contact address.
4.8 A list nobody can build means unreachable, not cleaner.
Exclusivity is the list property with real numbers behind it: the hundredfold spread across RevPack's 214 campaigns is one set. The conference attendee lists behind the only 15 to 20% reply figures here are another. Accuracy and size are what the enrichment and scoring vendors sell instead. Spend the list budget on a cleaner database and you end up testing offers, rather than wording, on a segment still below 4% positive reply.
Your offer beats everything except your list.
Your offer is the second-biggest lever in outbound and the only big one your volume can honestly test. A structurally different offer moves results two to three times, and an effect that large shows up in a few thousand sends.
5.1 Spend your whole test budget on offers.
Every other variable at this volume is too small to detect or too slow to read. So the offer is the only place a limited send budget can buy a real answer.
5.1.1 Two different things get called the offer, and the difference decides what you test.
The company offer is what you sell and on what terms: promise, proof, price and guarantee. It changes over weeks, the construction rules later in this chapter build it, and it is what you test across campaigns. The ask is what one message requests, a reply, permission to send something, or fifteen minutes, and a buyer judges it in about a second. Cold-email practice calls both the offer, but they move independently: a strong company offer can die behind a greedy ask, and a modest ask cannot rescue a commodity offer. When this guide says offer with no qualifier it means the company offer; the ask has its own section here.
5.1.2 Only offer changes are big enough to see.
Swap one offer for a structurally different one and results move roughly two to three times, so compare versions on replies. On a 5% reply baseline, the top of the researched band, a doubling resolves in 434 sends per version. But the 50% lift a wording change might buy needs the 1,469 sends per version Chapter I prices. Booked-meeting rates need tens of thousands of sends per version, which no volume in this book reaches.
5.1.3 Run offer discovery on a rig you are willing to burn.
Nick Saraev, who sells a paid automation community and an agency, runs five to ten structurally different offers at once. Each runs on its own scraped list, so offer and list get tested together at volumes that contradict everything here about cutting sends. He prices the week at $200 to $300 for 5,000 sends: about $150 for fifty mailboxes, about $50 for leads, $50 to $100 for software. Run it only as a sacrificial rig on throwaway domains before your researched system exists, and retire it the day an offer wins. He also blames copy for failures, reversing what his rig varies and Chapter VII's ranking, so take his design and leave his diagnosis.
5.1.4 A booking floor culls the losers and cannot crown a winner.
Saraev kills any version that books under 0.25% of its sends: at a thousand sends per version, fewer than three meetings. That is a blast-volume floor near the bottom of the 0.2 to 1% band Chapter I measures for agencies. Chapter III's arithmetic puts a researched sender nearer 0.75% booked, three times higher. Use the floor to drop obvious losers and never report the survivor as a finding, because the threshold is his own and unsourced.
5.1.5 A different niche each weekday tests far too slowly.
Nick of Reprise AI, who sells an inner-circle program, teaches testing a different niche every weekday at five to ten sends a day: Monday orthodontists, Tuesday immigration attorneys. List and offer are the right things to vary, but five to ten sends per version sits one to three hundred times below what a reply-rate read requires.
5.1.6 Tag every version and count closed deals.
Give each offer version its own discount code and CRM tag, then judge versions on closed deals. Per-agent discount codes plus dedicated CRM tags are the one cheap tracking method anyone in this evidence base runs, at SaaStr, which sells conference tickets and sponsorships. Nobody has measured the method on offer versions. A few tenths of a point of reply is inside the noise at this volume, while a closed deal you can count on one hand and still trust.
5.2 Every specific offer dies; the construction rules last.
The free audit went from standard opener to documented spam pattern by May 2025, and Nick Saraev argued it down again in 2026 because its friction runs backwards. The rules underneath last: risk reversal, meaning you pay only if it happens, a countable unit and a stated deadline. Those have held across direct mail, email and DM for decades because they are not tricks tied to one channel.
5.2.1 One template covers every offer worth sending.
Write it as "X thing in Y time or Z risk reversal". The template is Saraev's, and the offer vocabulary here traces back to Alex Hormozi, whom he openly draws on. It stays the most reusable template anyone produced because it forces the three things most offers leave out: a countable deliverable, a deadline, and a risk reversal.
5.2.2 Run the offer checklist in order, or you guarantee the wrong thing.
Saraev's SOLVE checklist: spot the pain, outline a specific outcome, limit the buyer's risk with a guarantee, value-pack with urgency, execute with one clear call to action. Everything he teaches routes back to the community he sells, so treat it as a checklist rather than evidence. The order matters because running it out of order leaves you guaranteeing an outcome nobody wanted.
5.2.3 Cut friction while everyone else inflates the promise.
Alex Hormozi's $100M Offers writes perceived value as dream outcome times perceived likelihood of achievement, divided by time delay times effort and sacrifice. He claims the denominator is the moat: promises and testimonials can be faked, cutting the buyer's time and effort cannot. Saraev's version, stated return on investment times social proof over friction, restates Hormozi rather than confirming him. Buyers have stopped believing the numerator everyone competes on, so cut the time, effort and risk they spend before finding out whether you are real. Carry the equation as vocabulary, because the corpus, this guide's evidence base, contains no test of it in B2B outbound.
5.2.4 The two best-credited offers are opposites.
The first is pay-on-results with a countable unit: twenty booked sales appointments in the next sixty days or you do not pay. Saraev credits it with taking his agency Leftclick past $70,000 a month. The second gives a finished deliverable away first: send us a title and we will write you a free 500-word blog post. He credits it with scaling One Second Copy to $90,000 a month, but has told it elsewhere at $50,000 a month and roughly $1M a year. Both figures are self-reported with no time period or underlying numbers, so keep the structures and discard the numbers.
5.2.5 Price against the human you replace.
Offers that save time price at the floor; offers that make money price at the ceiling. A buyer pays a fraction of a cost saving and a multiple of a revenue gain, so bid against the salary of the human role your work displaces. Pricing your cost plus a margin makes you a tool instead of an outcome.
5.2.6 Guarantee the milestone, never the revenue.
A 90-day money-back guarantee tied to a defined delivery milestone is a refund promise. A guaranteed monthly revenue floor, the "$15,000 a month minimum" that shows up in cold DMs, is an earnings claim. That is the classic deceptive-advertising fact pattern, worse when the promise never reaches the signed contract. Nobody has measured which converts better, so choose on legal exposure.
5.2.7 The free audit died, and your offer will too.
Free audits worked until agents made them free to produce, and the category filled with near-identical scraped audits carrying a $100 to $200 payment link. Ritner Digital's May 2025 teardown documents this as a recognised scam pattern. Every specific offer decays the same way, so put a date in the calendar to retest yours.
5.2.8 When everyone claims the same thing, bring proof you can touch.
Every AI vendor pitching one small legal firm used the same words about security, compliance and certifications. The winner carried in a server, put it on the table, and said the data lives here and will not leave. The account is second-hand from an unnamed firm, nothing measured, but the mechanism travels.
5.3 Change the ask before you change the copy.
The call to action is the offer rather than the line that closes it, and a buyer judges that one part of a cold message in a second.
5.3.1 The one vendor who got in offered to do the work himself.
SaaStr had frozen new agent purchases, counted roughly 100 vendor approaches over 30 days, and let exactly one through. That one was Vector, whose CEO offered to do the deployment himself, live fifteen minutes later. That is a buyer-side conversion rate with a known denominator, which almost nothing in this field has. And the freeze makes one in a hundred a floor, because this buyer had decided to buy nothing at all.
5.3.2 Every message that failed asked for a call.
SaaStr summed up the other ninety-nine, of 100 vendor approaches during its purchase freeze, in one line: everyone else asked to get on a phone and talk about their product. Same freeze, same inbox, same month, so the only difference was the ask.
5.3.3 Ask for interest when cold, ask for a time once warm.
Gong Labs, the research arm of the sales-conversation platform Gong, scored 304,174 emails against whether a meeting was booked within ten days. The corpus is sales reps' own streams rather than cold lists, and its booking levels run roughly forty times this guide's cold arithmetic. So carry the ratios and leave the levels: cold, an interest-based ask ("worth exploring?") booked about twice as often as proposing a specific time. Once a deal is open the order flips, and the specific-time ask books about a sixth more often. The sample is large, but the study is the vendor's own and nobody independent has audited it.
5.3.4 Climb the ladder: finished thing, permission, then a time.
Lead with the finished thing, then ask permission to send more: "want the full version?" Propose a time only after a hand goes up, because each rung asks for less than the one above.
5.3.5 Executives with no time do better with a direct ask.
Sam McKenna, who sells a trademarked sales methodology, reports a direct ask doing about 50% better with qualified executives who have no time. The figure is hers and self-reported, so read it as a sign the ladder, finished thing to permission to a proposed time, depends on the person: pick your ask by who you are writing to.
5.3.6 You cannot spam an offer to do real work.
Offering to do the work yourself lasts for a structural reason, not a measured one. Real work costs real hours for every prospect, so nobody can mass-produce it.
5.4 Check the arithmetic before you promise the offer.
A great offer can be impossible to deliver at a profit, and pay-per-meeting is the standing example. At the send rates anyone has actually measured, the infrastructure alone eats most of the price.
5.4.1 Know your sends per meeting before you quote.
The published funnels put a booked meeting at roughly 100 to 500 sends at volume, the range Chapter I reports. Smartlead self-reports 1% across 14.3 billion sends, the widely relayed Gong figure is one per 344, and Instantly runs near 0.2%. This guide's researched floor implies about 133, from the 300 sends to 2.25 bookings Chapter III sets, making the comfortable 100 to 200 practitioners quote the optimistic, self-reported edge. The one on-camera demo implying better, Saraev's 48-hour burst, counted meeting requests rather than booked meetings, into community owners on the platform his own paid community runs on. At about $0.02 a send the range prices infrastructure at $2 to $10 per booked meeting. Your price also has to recover the research and review hours behind each send, which the next node counts.
5.4.2 Pay-per-meeting pricing leaves almost nothing to pay yourself with.
The AI agency Reprise publishes the going rate for outbound as a service at $150 to $300 per qualified booked call. Infrastructure claims $2 to $10 of that at the funnel rates above, so the price is really carrying data and operator hours. What is left is $50 to $200 of gross before any labour, data or sequencer cost.
5.4.3 Research is cheap now; your review time is not.
The scarce input is no longer money, which Chapter VII prices. Agents made research cheap while the 10 to 30 minutes of operator review per asset did not move. Ration custom assets against how many a human can check, not against the budget.
5.4.4 One person's review time caps how much custom work you can promise.
Chapter IX caps the human-in-the-loop check, a person reading each output before it ships, at five to ten per person per day, your whole budget for custom assets. Neither more compute nor more agents lifts it, so your offer cannot promise more custom work than one person can read.
5.5 The rate card is not the call to action.
Practitioners argue past each other because one group means the ask inside a cold email. The other means the pricing behind a services business, and the two questions share almost nothing.
5.5.1 Skip this section if you sell your own product.
Everything below prices delivered labour: consulting, agency work, deployment services. If you sell a product you already built, the offer in your email is the only offer you have, and none of this applies to you.
5.5.2 Price on outcomes only where money is countable.
Charge per meeting only after your own measured sends-per-meeting clears the arithmetic two sections back, or take 10 to 20% of closed revenue. Either works where the outcome is countable and traceable to money: cold outreach and LinkedIn qualify, SEO does not. That is because rankings move on a delay you do not control and the buyer can argue your work did not cause them.
5.5.3 A not-to-exceed cap gives safety without giving away margin.
Bill time and materials off a written requirements document, then cap the total at 10 to 35% over your estimate. That means 10% for work you have done before and 35% for work you have not. The buyer gets a fixed-price ceiling while you keep the upside of billing what the job costs, the cleanest risk-reversal device in the pricing material.
5.5.4 Sell a small fixed project first; nobody retains a stranger.
Land a $1,000 to $2,000 fixed-price project pitched at a specific dollar outcome, deliver it, then present the retainer. Saraev's ladder puts that retainer at about $5,000 a month for roughly six months, about $30,000 of lifetime value a cold retainer pitch would never open. The arithmetic is illustrative, with no client cohort behind it.
5.5.5 Real buyers start at tens of thousands, not nine dollars a seat.
Jason Lemkin of SaaStr says every AI go-to-market agent platform they evaluated starts at $30,000 to $70,000 and up. The deployment engineer, training and onboarding come on top, and none of the self-serve tiers worked. This is the organisation behind the one-in-a-hundred count above, buying rather than selling, so price against that floor. The $9-a-month seat rival vendors advertise is not what buyers at this end pay.
5.6 Almost nobody follows their own offer rules, which is your opening.
Practitioners publish offer advice constantly and almost never apply their own rules, and the one tactic they all recommend carries documented legal exposure.
5.6.1 Doing what everyone recommends is still enough to stand out.
Across eleven offer structures collected from practitioners, risk reversal appears in three and a stated timeframe in exactly one. These are the same people who teach risk reversal and deadlines, so the gap is application rather than knowledge. Clearing the bar costs one guarantee and one delivery date.
5.6.2 Leave out the delivery date and the guarantee cannot be enforced.
The field almost always leaves out when the thing arrives, which a vague consulting pitch survives and a 90-day outcome does not. The deadline is half of what the buyer is purchasing, and without a stated date a refund promise has no moment it can be called on.
5.6.3 Building the prospect an asset drew a cease-and-desist.
Building a prospect a personalized asset before pitching is the most-recommended tactic among the practitioners studied here. Reprise sells it as "25 times more effective than a generic cold email" with no stated baseline anywhere. It is also the tactic that produced a documented copyright cease-and-desist, because the asset carried the prospect's own logo and data and the sender had licensed neither.
5.6.4 Use their numbers, never their logo.
Build the asset from figures the prospect has publicly stated, and keep their logo, scraped reviews and competitor comparison tables out of a PDF carrying your branding. The arithmetic that makes the offer obvious needs no logo, and the logo is what turns a helpful asset into an infringement claim.
5.6.5 Free work needs a signed release, and your agent needs a guardrail.
Both are promises escaping your control, and both are cheap to close early. Free-work-for-a-testimonial offers assume a case-study release nobody ever writes up, so get it signed at kickoff while the client is keen. An agent can offer terms you never authorised, so give it a guardrail, a rule it cannot write around, listing what it must never offer. Nate Herk, who sells a paid automation community, supplies the failure. His agent misread its task list and emailed his whole list a discount code never supposed to ship, list size never stated.
5.7 Offer testing is affordable; the offer you pick may not be.
Test your offer rather than your wording. A structurally different offer moves results two to three times, while a wording change needs more sends per version than you have. Count meetings that happened rather than replies received, because a few tenths of a point of reply is inside the noise at this volume. Test well and you can still pick a winner that leaves $50 to $200 of gross per meeting.
Who the email comes from is the only thing that compounds.
Who you are sets your reply rate before anyone reads your message, and you build that by publishing under the names of real people. Published work is an input to outbound, so fund it from that budget and judge it on pipeline.
6.1 Even the company selling message automation puts the wording last.
Sam Blond runs the message-automation company Monaco and puts two things above the wording of any email. Those are sender brand, who the message comes from, and message-market fit, whether your offer lands with the market you picked. Sequence and message structure, everything a copywriter actually types, rank below both, an admission against the interest of what he sells, and no study isolates the effect.
6.1.1 Sender brand outranks anything a copywriter controls.
Both of Monaco message-automation CEO Sam Blond's drivers, sender brand and message-market fit, sit ahead of anything a copywriter touches. Sender brand is the reputation of the human name in the from line. The startup he calls "definitionally unknown as a brand" loses to the known company on that alone; message-market fit he treats as a precursor to product-market fit. His summary, "there's only so much we can do with sequence structure and message structure," caps the wording rather than what you are saying. Nothing published separates sender brand from the list or the copy, so this is the best-incentivized unmeasured claim in the field.
6.1.2 Company recognition and a personal byline are two different assets.
Two assets travel under the same phrase: company recognition is whether the recipient has heard of the business on the envelope. A new company starts the first touch at a disadvantage no writing repairs. A personal byline, whether they have heard of the human sending it, accumulates to the person and travels between jobs. It cannot be created quickly at any price, and this chapter is mostly about the byline because an operator can start building it on a Tuesday afternoon. While you are unknown on both counts, the list and the offer carry the first year.
6.1.3 Your tone is set by the sender and the list.
Nick Saraev, who sells cold-email and automation education, teaches a casual "hey man" salutation he says lifts replies with small-business owners, creators and agencies under $10M. Tim Ferriss, who sells books and a podcast and no outbound product, describes the same emails from the receiving end as an instant archive. That is because he is a high-status recipient behind a gatekeeper. Both are right about their own segment, so tone follows from who is sending and who is receiving. Weigh Saraev carefully wherever he appears, because Chapter VII shows him blaming copy in one place and ranking the list first in another.
6.1.4 Get on the shortlist before you send anything.
Forrester surveyed 11,352 buyers in 2024: 92% start with a vendor already in mind, and 41% with a single preferred vendor before formal evaluation. The account-intelligence vendor 6sense surveyed more than 4,000 buyers of $25K-plus purchases in November 2025. It found 94% pre-rank vendors, with the early favorite bought 77% of the time.
6.1.5 The preview text is a lever almost nobody uses.
Everything that decides the open renders before your body text does. In email that is a sender name of about 20 characters, a subject line, and a preheader, the teaser shown beside the subject. Subject and preheader truncate at roughly 148 to 150 characters combined and unused space fills with metadata, so write to the full 150. On LinkedIn it is your picture, first name, a teaser of 50 to 55 characters, a job title of 50 to 60 characters, and the premium badge. So a shorter first name buys a longer visible teaser.
6.2 Publish under real names and aim narrow.
The 2026 B2B creator research from Limelight and the Content Marketing Institute puts audience match ahead of audience size. A creator whose 1,000 to 5,000 followers are exactly your buyers outranks one with 100,000 general business followers. At enterprise size that means adding named publishers, not making the company page louder.
6.2.1 Sender brand cannot be pooled into a company account.
The case for the personal byline is reach mechanics plus portability, not an outcome study. A personal LinkedIn profile reaches multiples of what a company page reaches for the same content. Recognition attaches to the person and travels with them between jobs, and nobody has isolated person-versus-company publishing on pipeline. So a large company adds named publishers, one per operator, each carrying their own cost and audience because neither transfers.
6.2.2 Put publishing in the outbound budget.
The same creator research, the 2026 Limelight and Content Marketing Institute study, found 73% of buying committees check vendors through practitioner content before contacting sales. It is a single-source figure from people who study creator marketing, but it puts your published work upstream of the first sales conversation. So fund it from the outbound budget and judge it on pipeline.
6.2.3 Count the readers who match your buyer profile.
Content aimed narrowly at one buyer profile is reported at 15 to 22% engagement from people who match that profile; viral content is reported at under 1%. Both figures are self-reported, but the direction is the entire strategy: a hundred right readers beat ten thousand wrong ones.
6.2.4 Views are not the product.
Nick Saraev sells automation education, so this admission costs him: one video reached 43,000 views and produced 336 subscribers and zero paid conversions. That is reach-first publishing working exactly as designed and returning nothing, so watch engagement from people who fit your buyer profile instead.
6.2.5 Post 3 to 5 times a week into a shrinking platform.
Richard van der Blom's Algorithm Insights study covered 1.8 million LinkedIn posts in the twelve months to February 2025. It found views down 50%, engagement down 25%, follower growth down 59%. His October 2025 update, from a different sample of roughly 400,000 profiles, gives views down 47%, engagement down 39%, follower growth down 42%. The editions are not comparable, so read each on its own. Post three to five times a week anyway and expect reach per post to keep falling.
6.2.6 Comment more than you post.
Comment anyway: a comment costs a fraction of a post and borrows an audience somebody else assembled. Van der Blom's 2024 report put a comment at fifteen times the weight of a like, which AuthoredUp's later language-aware analysis revised to roughly two times. A separate report credits van der Blom and the vendor Botdog with a different claim: comments over fifteen words carry about twice the weight of shorter ones. One of those on a high-reach post returns 5,000 to 10,000 impressions. Both versions land on about 2x and both involve fifteen, suggesting one garbled claim rather than two findings, and LinkedIn has never published any comment multiplier.
6.2.7 Engagement bait and pods get you caught.
One founder ended posts with "agree or disagree?" and lost 74% of his impressions, falling from 4,200 to 1,100. Another account joined an engagement pod, a group that trades likes on cue. Its reach collapsed from 8,500 impressions to 340 overnight, with recovery taking 60 to 90 days. LinkedIn vice-president Gyanda Sachdeva confirmed the enforcement and the Lempod extension was pulled from the Chrome Web Store, so this is detection working rather than bad luck. The "97% pod-detection accuracy" figure circulating alongside has no attribution.
6.2.8 Budget fifteen minutes a week for each publisher.
Publishing under a real name costs about fifteen minutes a week of voice memos, which a person or an agent edits into posts. The shape is credited to Peter Kazanjy and the Atrium model. It is a practitioner estimate, so treat it as a floor, and it is per publisher. Five named publishers cost five times fifteen minutes, because the voice is the asset and voices do not merge.
6.2.9 Almost nobody in your buyer's feed is publishing against you.
Roughly 7.1% of LinkedIn's one billion or so users post regularly, a figure with no traceable primary source, so take it as an order of magnitude. It still explains why narrow publishing works while reach falls: competition for one specific buyer's attention is far thinner than the headline user count suggests.
6.3 Warm beats cold, and nobody has proven by how much.
Every published warm-versus-cold number comes from somebody who sells warm, and one confound could explain all of them. People who engage with your content have already picked themselves out as interested. Design for warm anyway, and plan against the conservative multiple, the 2-to-4-times reply lift.
6.3.1 Nobody has isolated content's causal lift.
Every warm-versus-cold gap on offer is vendor-reported or self-reported: Warmly's 12.8% "warmbound" close rate comes from "500+ B2B deals" on its own platform. The 18% against 3.4% reply figure is a blog number citing another blog. It never says whether those were researched campaigns or volume sends, two regimes an order of magnitude apart. Underneath sits the confound: engaged people had already selected themselves as interested, so the warm group was more likely to buy before you published a word. The experiment nobody has run is a holdout test, leaving a matched set of engaged accounts alone on purpose.
6.3.2 Budget against the low end of the warm advantage.
A synthesis of data from the cold-email platforms Woodpecker and Instantly, plus Forrester, puts the honest lift at roughly 2 to 4 times on reply. That was measured at volume across whole customer bases, so carry the multiple as a direction. Vendor pages claim 10 to 15 times, and the close-rate version is Warmly's 12.8% warm against its own 6.3% outbound. That is the vendor's own data on a few hundred deals, with "warm" defined by the vendor that sells warm. Budget against the low end, because if warm is not clearing 2 to 3 times cold in your own numbers, the warmth is illusory.
6.3.3 The strongest warm number comes from old relationships.
Champify's 2025 Impact Report looked at deals involving prior-experience contacts, people who already worked with you somewhere else. Those deals win at 36.8% against a 19% SaaS average, and the same report says CRMs miss 78% of former customers with qualified job changes. Champify sells champion-tracking software and this is its own report, but it points at your customer history rather than your posting cadence.
6.3.4 Three sellers with unrelated products all put cold last.
All three make cold outreach conditional on case studies that warm work produced first. Nate Herk, who sells community and course access, puts cold at step seven of seven, opening only once the first six produce proof. Nick of Reprise AI calls warm work method one and cold method two, the way to scale beyond it. Nick Saraev will not sell a retainer cold and inserts a cheap fixed-price project as the trust step. None reports a comparative number, and both Saraev and Herk cite Alex Hormozi, so this may be one idea circulating rather than three independent confirmations.
6.3.5 The strongest warm-first claim costs its author sales.
Nick of Reprise AI sells automation-consulting training into this exact market. He says it "doesn't work well selling to cold traffic, and that is the number one issue that we have seen with people inside of our community." He is describing his own customers failing at work his own product sits beside, which is why this outranks every reply-rate chart in the category.
6.4 Capture your own signal and skip de-anonymization.
Capturing signal from your own site is free and legally quiet. Person-level de-anonymization names the individual who visited, via a tracking pixel, a snippet of vendor code matched against identity data the vendor holds elsewhere. That gives the layer a different legal profile: US-only, not GDPR compliant by its own vendor's account, and inside a live wave of California wiretapping claims.
6.4.1 Person-level de-anonymization is out for any EU audience.
Adam Robinson runs RB2B, which sells exactly this product, and he says it "is not GDPR compliant. So it's US only." A vendor drawing that boundary around his own product is the most credible statement available about where the layer works. Treat it as usable for US recipients and out of scope the moment a campaign reaches anyone else.
6.4.2 Price the California wiretapping exposure before you install the pixel.
California's Invasion of Privacy Act, or CIPA, makes it illegal to install a pen register without consent, a device that records who contacted whom. Section 638.51 is now used against website tracking on that theory at $2,500 per violation. More than 800 CIPA claims were filed in 2025 alone, and two 2026 settlements show the scale. The Los Angeles Times paid $3.85M in June and the health system Sutter Health $21.5M in April. Do that multiplication before the pixel goes in, not after the demand letter arrives.
6.4.3 Defendants win some pixel cases, and the legal theory survives.
In Rounds v. Development Dimensions International, a California federal judge held that cookies do not satisfy section 638.51 and dismissed with prejudice on 11 March 2026. But Camplisson v. Adidas, filed November 2025, survived the pleading stage. California's proposed SB 690 exemption stalled in 2025 before a rehearing on 1 July 2026. The question is unresolved, so whoever installs the pixel is the test case.
6.4.4 Work the inbound you already have before buying more.
Salesforce's much-quoted 3,200 agent-sourced opportunities came out of the 75% of its own inbound leads nobody had ever touched. Salesforce never published what the 3,200 was out of, so read it as inbound recovery. Before installing a capture layer with litigation attached, count how much of your existing inbound goes unworked: the same asset with none of the exposure.
6.5 Judge inbound on its own clock and its own ceiling.
Inbound shows sourced pipeline at nine to twelve weeks and moves customer acquisition cost, the CAC number your board tracks, at twelve to eighteen months. And its conversion volume is tiny even when it works. Measure it against a send count and you will kill it while it is still on schedule.
6.5.1 Inbound pays on three clocks, the longest eighteen months.
Expect first signals at 3 to 6 weeks, measurable sourced pipeline at 9 to 12 weeks, and an effect on customer acquisition cost at 12 to 18 months. Jessica Schultz of Amplify sets her floor at "minimum 4 to 6 months." Every source behind the timeline is self-reported vendor or blog material, so trust the shape and treat anyone promising faster as overselling.
6.5.2 Inbound conversion is tiny even in the flagship case.
The HR software company Personio ran a website chat agent called NIA, which six weeks after launch produced 140 meetings in seven days against roughly 200,000 sessions: 0.07%. That counts meetings per website visit rather than per send, and it is one of the few inbound figures published alongside its denominator. It is also the case the category holds up as proof, so the part nobody quotes belongs beside it. NIA started giving legal advice and bashing competitors, and chief revenue officer Philipp Lohr names "four weeks where we didn't do enough" before it was trained properly. And he does not know "how many demos we wasted by not training NIA." Inbound earns its place because of who converts, never how many.
6.5.3 Add "how did you hear about us" on day one.
The question measures nothing useful in month one and matters a great deal by month nine, when content, cold sends and referrals all run at once. Without a self-reported source captured from the start, plus a source field in the CRM, you cannot tell which one produced the pipeline in front of you.
6.5.4 Route inbound to a human, not to a booking link.
The warm advantage gets spent between the raised hand and the held meeting. A bare booking link hands scheduling work back to the buyer at the moment you have their attention, so route the lead to a person. The field calls this speed to lead; the standard here is minutes for a form-fill and the same business day for everything else. Chapter I traces the famous five-minute rule to 2007 inbound web-form data and shows why it does not transfer to a cold account.
6.6 Publishing under a real name is the one bet you cannot place late.
Start publishing under a real name on day one, Move 4 of the eleven moves at the front of this guide, knowing the effect size is unproven. Every warm-versus-cold gap in this chapter is vendor-reported or self-reported, and the one hard count, Saraev's 43,000-view video, returned zero paid conversions. The honest case is the clock: sender brand attaches to a named human, compounds, and cannot be backdated once you are a year in. Skip it and you will judge publishing against a send count and kill it at week six, inside the three-to-six-week window before first signals appear. You will also drop person-level de-anonymization into the build without weighing its documented, dated damage: mailbox providers cutting you off and platforms enforcing their rules, not the privacy fine that gets talked about most.
Wording matters least, so work on it last.
Wording is the fourth of four variables, behind the list, the offer, and who the email comes from, yet it is the only one most published playbooks cover. You still have to write the thing, so work it in that order and the craft will pay.
7.1 Personalization stopped working once it became free.
Naming a prospect's alma mater or funding round used to cost time, and the time was the signal; AI made the research free, so it now signals nothing. What still works is situational relevance, meaning a specific problem people in their exact position have. Save the researched write-up for the few accounts senior enough to pay for the hours.
7.1.1 "Congrats on the funding" reads as a template now.
Buyers spot that opener instantly, and it wears out as you scale. The cold-email tool Woodpecker looked across roughly 20 million sends: the same merge-field campaign taken from 50 prospects to 1,000 dropped open rates from 49% to 38%. A merge field is the slot that drops each prospect's name or company into a fixed sentence. That measures decay, not absence, because Woodpecker ran no unpersonalized arm.
7.1.2 Above director level, talk about their company.
The sales-conversation platform Gong analyzed more than 85 million cold emails with the podcast 30 Minutes to President's Club. For directors and above, the company's situation moves the reply, not anything about the person. The alma-mater opener actively costs you, so lead with what is happening in their business.
7.1.3 The offer outweighs every wording choice.
In the same 85-million-email dataset, from the sales-conversation platform Gong, a compelling offer lifts replies 28% while pitching cuts them by as much as 57%. Both turn on what you ask for and what you give, decided before you write a word.
7.1.4 Save deep research for your most valuable accounts.
Your review time is the real cost of deep research, not the compute. One cost model puts a researched write-up at roughly $20 to $60 per account in 2026, against $150 to $400 of analyst hours in 2023. Most practitioners work on a flat monthly seat, where one more account costs nothing and the rate limit and your attention constrain you. Jordan Crawford runs a 53,000-company funnel on a $200 to $500 seat plus about $100 of data, while scheduled or chained research runs on the API, which Chapter XX prices. Spend the scarce 10 to 30 minutes of review per write-up on senior buyers, and send everyone else a hard offer with no personalization, which beats imitation research.
7.1.5 Automated research drops you back to average results.
Coldreach, an AI outbound tool built around research agents, reports 3.8% replies across more than 500,000 sends on its own platform corpus. That is the floor of the researched band rather than its ceiling. Jordan Crawford, who hand-builds these write-ups and sells the consulting around them, reports 11 to 32% on small hand-researched campaigns. Those numbers are self-reported with no named client and no control group. Automation at scale buys speed at the average, so keep the hand-built version for accounts that justify it.
7.1.6 Copy Nick Saraev's method, not his diagnosis of failure.
Nick Saraev's four-hour cold-email course blames copy for 50 to 60% of campaign failure and the mailbox or list for only 20 to 30%. Yet every teardown works the text and never shows the list. In a separate video he ranks list exclusivity as the single highest lever and teaches running three niches in parallel for one to two months. Then he doubles down on the best, calling list testing nearly free. He sells a paid automation community and both positions are his own recordings, so copy his method and ignore his diagnosis.
7.2 Spam filters catch sameness, not machine writing.
No spam filter in use today can tell that a machine wrote your email. What filters catch is nine thousand emails built on one sentence pattern, which is what one prompt run across a list produces. Spend the human hour on the reader, who deletes machine prose on sight, and nothing on fooling the filter, which never asks who held the pen.
7.2.1 AI vocabulary gets you deleted by humans.
A peer-reviewed study in Science Advances (Kobak and colleagues, July 2 2025) covered more than 15 million PubMed abstracts. "Delves" ran at 28 times its pre-LLM rate, "underscores" at 13.8, "showcasing" at 10.7. Humans read that vocabulary as machine-written and delete on sight while filters never look at it, so strip those words for the reader who scores them.
7.2.2 Gmail and Outlook score five things, none of them authorship.
They score authentication records, send volume, sending-domain reputation, spam complaints and recipient engagement, and nothing running in the filters today scores whether a machine wrote the text. Design against those five, because they decide whether you reach the inbox at all.
7.2.3 Stop counting em dashes and fix your DNS.
The em-dash hunt and the "not X but Y" hunt are forum folklore with no filter evidence. SPF, DKIM and DMARC are three DNS records proving your mail really came from your domain, and mailbox providers check all three before deciding where you land. Setup takes a morning, so move the hour off the copy and onto the records.
7.2.4 One prompt across the whole list is what earns the spam flag.
AI-written email gets spam-flagged at roughly 2.7 times the human rate, and two independent vendor tests agree. Saleshandy found 7.8% against 2.9% on 12,000 emails, Digital Applied 8% against 3% on 100,000 matched paired sends. Both sell into the answer they measured, so carry the multiple, not the decimals. The cause is fuzzy hashing, near-duplicate detection that fingerprints templated sameness whoever wrote the text, and unreviewed AI output is templated by construction. One prompt across 9,000 rows earns the 2.7x, not the machine drafting.
7.2.5 Merge fields do not break the pattern.
Merge fields change on every send and buy you nothing, because the hashing reads sentence structure and paragraph order, not the names slotted in. Two hundred emails sharing one shape are one email sent two hundred times, and the filter scores them that way. So vary the shape and split the list across two or three genuinely different structures.
7.2.6 Fake typos make the problem worse.
A typo inserted on purpose stays identical across the campaign, a rare token repeated at volume, which is exactly what fuzzy hashing catches. So does a fixed "Sent from my iPhone" line or a fixed stray double exclamation. Nobody has measured the penalty, so this is reasoning from the hashing, not evidence. Leave the typos out and vary the structure.
7.2.7 Let the model fill slots and finish it yourself.
Two independent vendor benchmarks, each selling into its own result, both rank AI-drafted, human-finished mail above pure human and pure AI, though their levels disagree. Saleshandy's 12,000-email test reads 14.7% reply for the hybrid against 10.4% human and 4.1% AI-only, while Lavender's 100-million-email platform corpus gives 5.1% against 3.8% and 2.4%. Carry the ordering; treat the levels as vendor-specific rather than planning numbers. Give the model the merge fields and a locked style guide, and keep the body, the ask and the final read yourself. Then run the first batch with sending switched off, reading every line before anything leaves.
7.3 Every channel needs its own writing.
Your angle, the reason this specific person should care, carries across email, LinkedIn and phone, while the wording has to be rebuilt for each one. Reuse the angle, rewrite the wording, and drop the fifteen-minute ask from all three.
7.3.1 Write email as plain text under 80 words.
Instantly, Lavender and the email-finding tool Hunter converge on a first touch under 80 words. Hunter's 34-million-email platform corpus puts the best reply rate in the 20-to-39-word band, at roughly 4.5%. Its separate 31-million-email read shows "quick question" opening at 28.7% against 32.9% for the alternatives, which makes the old safe subject line the below-average one.
7.3.2 A subject line only has to buy the open.
The outbound agency Belkins and the sequencer Reply.io measured 5.5 million emails: two-to-four-word subjects opened at 46% against 34% for ten-word subjects. Lowercase reads like internal mail, which is the whole job, because the subject only has to look like something a colleague could have sent.
7.3.3 Pick your subject-line rule by who reads it first.
Two to four lowercase words win when a human sorts the inbox, because a little ambiguity buys the click. A specific, verifiable trigger wins when an AI assistant pre-ranks the inbox, since the assistant scores relevance and archives anything vague. At director level and above, assume the assistant reads first and write the trigger into the subject.
7.3.4 Never ask for fifteen minutes in a cold email.
In the Gong and 30 Minutes to President's Club data, "do you have 15 minutes" cuts booking likelihood 44% on a first touch. A separate Gong Labs read of 304,174 emails found the same specific-time ask booking 15% of meetings while the contact is cold. It books 37% once the deal is live: a good move played three steps too early.
7.3.5 Send the blank LinkedIn connection request.
The automation vendor Expandi tracked 13.2 million connection requests. Replies to noted requests fell 37% in eleven months as templates saturated the channel, from 3.5% in May 2025 to 2.2% in April 2026. Acceptance is statistically flat, 26.42% with a note against 26.37% without. One post-acceptance measurement cuts the other way: Cleverly and Expandi jointly put reply after acceptance at 9.36% with a note against 5.44% without. Send blank because the noted channel is visibly decaying, not because blank replies better.
7.3.6 Comment to warm up one named prospect.
Leave one or two genuine comments on a named prospect's posts over about two weeks, then send the DM. Practitioners claim a 2-to-3x reply lift over a cold DM, and the strongest number behind it is the outbound tool Salesforge's own 1,000-prospect A/B test. That test is self-reported, so treat the multiple as directional. Use it only on named accounts, because commenting for broad reach has never been measured.
7.3.7 Auto-commenting destroys the signal it depends on.
The comment works because a person spent attention on it, so automating it removes the one thing the prospect was detecting. LinkedIn shipped a public "report AI slop" button on July 30 2026 that blocks hundreds of thousands of automated comments daily. It removed the vendor HeyReach's 16,400-follower page and its chief executive's profile in late March 2026. And it litigated the scraping API Proxycurl, roughly $10M in annual revenue, out of existence in July 2025. Comment by hand on the few accounts worth it.
7.3.8 Only name a mutual contact who will vouch.
Tim Ferriss receives this mail and sells no outbound product; describing his own filtering, he verifies a named mutual immediately. Nine times out of ten that person has no idea who the sender is or shook their hand once at a party. His stated consequence is "You're gone," so name a mutual only when you know they will vouch unprompted.
7.3.9 Follow up once, a week later, then stop.
Tim Ferriss again, describing his own inbox, from the demand side: display as little entitlement as possible, follow up once at least a week later, then stop. He sells books and a podcast, not outbound software, and he prescribes restraint exactly where the supply side prescribes persistence.
7.4 Cold calling at volume is a different business.
Cold calling at volume has its own economics, tooling and regulator, and none of that is covered below. What is covered is the single dial into a named account your research has already qualified.
7.4.1 Phone preference falls as you move down the org chart.
The sales-research firm RAIN Group puts stated phone preference at 57% for C-level and VP buyers, 51% for directors and 47% for managers. The survey circulates undated, and the declining ladder is the finding, not any single number. So flattening it into "VP and above" throws away the gradient that decides who is worth a dial.
7.4.2 Dial only the lists you built yourself.
The Bridge Group's independent B2B benchmark, 2024 edition, covers 351 to 406 companies: median connect 6.1%, top quartile 8.9%, against 46 dials a day. A separate full-year read by Belkins and Nooks of SDR activity, the reps who only book first meetings, puts one meeting at about 370 dials on generic data.
7.4.3 Never let an AI voice agent dial a cold prospect.
FCC Declaratory Ruling 24-17, released February 8 2024, classifies AI-generated voices as "artificial" under the Telephone Consumer Protection Act, effective immediately with no live-agent carve-out. Every AI marketing call to a US mobile requires prior express written consent, which a cold prospect has not given. Statutory damages run $500 to $1,500 per call with no cap.
7.5 Personalized video still beats text, on two conditions.
Video is one of the few wording-level levers with a defensible 2-to-3x over text, on two conditions: a human records it, and the recipient asked first. Break either and you are back at baseline with higher production cost and a new legal surface.
7.5.1 Discount the video reply rates that vendors publish.
For human-recorded video on a clean list, the defensible planning band is 10 to 22% reply against a text baseline of 3.4 to 5.8%. Both come from narrow researched campaigns rather than agency volume, which runs an order of magnitude lower. The 25 to 30% figures in circulation are vendor marketing or account-based-marketing ceilings on tiny hand-picked lists, self-reported with no control group, so plan on the band.
7.5.2 Ask permission in plain text, then record.
Send a plain-text touch offering the video and record only for the people who say yes. At a 1.4% positive-reply rate, arithmetic sitting between the volume and researched bands, a thousand candidates become roughly fourteen real recordings. That protects deliverability and keeps the work inside what one person can sustain: 15 to 25 videos a day at three to five minutes each.
7.5.3 Hit four beats and name yourself only at the end.
The recording runs 45 to 90 seconds and carries four beats. Open cold on the prospect's own screen, say why this is relevant to them, give two or three concrete ideas with one proof point, then one soft ask. Save your name for the final five seconds, because watch-to-completion drops about 40% past 90 seconds on the vendor's own data. And pin default playback at 1.2 to 1.5x, because the viewer will speed you up anyway.
7.5.4 Automate everything except the four minutes of recording.
Research briefs, script drafts, thumbnails, sending and follow-up are all agent work; the recording is the product, because it proves a person spent time on this. The published avatar pipelines automate exactly that step while leaving the rest manual, so build the automation around the recording rather than through it.
7.5.5 Personalize the email, reuse the same video.
Once an offer converts, the video stops needing to change and only the message carrying it does. State the price inside the video and say plainly what buying looks like. Someone who watched 90 seconds of you has already given more attention than a reply would have bought.
7.5.6 Voice clones and avatars lose you more than they win.
Roughly fifteen US states now have chatbot-disclosure laws, so a synthetic voice or face arrives carrying a disclosure obligation. Controlled experiments on identical content found people rated the AI-labeled version less natural and less useful.
7.6 Good sequencing means sending less.
Nearly every sequencing decision worth making is a decision to send less: fewer channels, fewer touches, a harder stop. The largest recoverable loss anywhere in outbound sits between a positive reply and a meeting that actually happens.
7.6.1 Hold the line at two channels and stop after four weeks.
Practitioners have converged on 8 to 14 multichannel touches across three to four weeks, then a hard stop and a 30-day cool-down, replacing the old 12-to-21-touch cadence. The "287% multichannel lift" used to justify channels four and five traces to Omnisend, a consumer e-commerce platform, with selection bias and no causal design. So hold the line at two channels.
7.6.2 Earn each extra email with performance.
Belkins measured spam complaints rising from 0.5% to 1.6% by the fourth email, roughly triple, and a third email cut replies by as much as 20%. Length also costs reach by plain arithmetic: at a fixed daily send capacity, a two-step sequence halves how many new prospects you can reach at all.
7.6.3 An unsubscribe costs you that person forever.
An opt-out removes that person from everything you will ever sell. So the blast radius of one careless send, how far the damage spreads, is your whole catalogue for good. No sequencer propagates an opt-out across email, LinkedIn, phone and video, because each tool knows only its own sends. So keep one do-not-contact store you own and check it on every path.
7.6.4 Answer inside the business day, never with a link.
The practitioner standard for a warm outbound reply is same business hour to same business day, answered by a person rather than a bare booking link. Chapter I takes apart the five-minute rule that would push you faster, since it never described cold outbound at all.
7.6.5 Propose two specific times and do the scheduling yourself.
The gap between meetings booked and meetings held is the largest loss you can actually recover. A campaign at Chapter III's floor, the minimum viable volume, books 2.25 meetings a month and holds about 1.5, arithmetic that chapter runs. Propose two specific times, get the booking inside three days, send three reminders rather than one, and send a short pre-read. Three reminders cut no-shows 29% against one, vendor-reported with no independent check, and every item means you absorb the friction instead of the buyer. Measure your show rate from the first month, because it multiplies meetings booked into meetings held everywhere downstream.
7.6.6 Revive a "not now" when something real changes.
A concrete new reason to return beats a check-in on a date you invented, and the strongest signal is a past champion landing at a new company. The vendor UserGems analyzed more than 5,000 opportunities among its own customers and reports a 114% higher win rate, 54% larger deals and 12% shorter cycles. Discount the magnitudes as a vendor measuring itself, but the mechanism is a warm relationship and that survives.
7.6.7 Keep a human in the loop on every warm reply.
Letting AI sort replies into buckets is safe, and a third-party test of 1,200 Reply.io replies measured about 87% accuracy. Letting it send them puts you on the hook for what it says. In Moffatt v. Air Canada (British Columbia Civil Resolution Tribunal, February 14 2024) the tribunal held the airline liable for its chatbot's promise and dismissed the separate-legal-entity argument as "a remarkable submission." Assume you are bound by whatever your agent says, so a person approves before anything leaves. The agent may draft and never send, which is least privilege, and keep a kill switch that halts every automated send within reach.
7.6.8 A reply-to-meeting rate under 15% means the list is wrong.
If fewer than 15% of your positive replies convert into a booked meeting, against the 25% planning rate, people are answering and then declining to proceed. The words worked and the audience was wrong, which is a list problem no rewritten follow-up will move, so go back to layer one and rebuild the list.
7.7 Craft wins the reader and never rescues the list.
Write the email yourself and count the AI steps in the chain, because no filter in deployment scores whether a machine wrote the text. The human final pass buys the reader, and varied sentence structure buys the inbox. Skip this and you spend the hour hunting em dashes instead of setting SPF, DKIM and DMARC. You are still rewriting openers while a reply-to-meeting rate under 15% says the list is the problem. Go back and pick a list nobody else can build.
Most builds only send. Build the half that listens back.
Almost every demo shows the same four steps, find, enrich, write and send, and you can assemble that chain in an afternoon. The architecture is what comes back. Results from the send, the reply and the calendar flow into your state store, your offer library and your ICP file, your written definition of who you target. Of the fourteen recorded builds that name every tool in the chain, thirteen have none of those return paths.
8.1 Agents cut research cost 10x and review cost by nothing.
A narrow list only works if you research every row deeply, and research used to be the cost that stopped you. Agents cut that cost roughly tenfold while cutting none of the human time it takes to check the output. Every capacity number in this chapter measures that gap.
8.1.1 Researching every row stopped being the expensive part.
In 2023, a researched write-up on one account cost $150 to $400 in analyst hours. The one cost model that prices both years puts the 2026 figure at $20 to $60 per account, most of it human review rather than compute.
8.1.2 An agent executes decisions it cannot make.
Parts IV through VII covered four choices: who you target, what you offer, whose name goes on the message, and how it reads. An agent can carry out every one while originating none, so the architecture question is a question about files. Each choice must be written somewhere the agent reads while it runs.
8.1.3 Review cost did not move, so capacity is human.
Compute got roughly ten times cheaper while human verification got no faster. That caps you at the five to ten checked outputs per person per day that Chapter IX prices.
8.2 Everyone builds the forward chain and skips the rest.
Fourteen recorded builds name every tool in the chain, and all fourteen run nearly the same forward sequence, find-enrich-write-send, without anyone arguing about it. What varies is whether anything comes back, and in almost every case nothing does.
8.2.1 Only six of the twelve layers accumulate your decisions.
The twelve, in order: targeting, data, enrichment, verification, signal, offer selection, composition, approval, send, reply, state-and-suppression, and feedback. Six get better the longer you run them, because every judgment gets written into them and stays: targeting, offer selection, composition, approval, state-and-suppression, and feedback. The claim is about layers; Chapter VI ranks the campaign choices. Four are execution you rent: data, enrichment, verification, and sending. Signal and reply sit in between, and the evidence does not settle which side they belong on.
8.2.2 Everybody already owns the forward chain.
In all fourteen named-tool builds covering the forward chain, find-enrich-write-send, Claude drives a bought data service, twelve keep versioned context files, and eight let a bought sequencer send. A further eighteen builds disclose no numbers, and in all eighteen nobody owns a suppression store; whatever suppression exists is rented from inside a bought tool.
8.2.3 Thirteen of fourteen builds ship half a system.
Thirteen run data, enrichment, composition and send, and none builds state, suppression, audit or feedback; the fourteenth builds the state half and none of the acquisition half. Finding and writing to people demos well and governance does not. Three of these builders run real client campaigns at volume, and not one mentions a suppression store even in passing.
8.2.4 Put offer selection and suppression on your diagram.
The published component catalogs for this stack leave both layers off: offer selection gets folded into "message," and suppression gets assumed into whatever tool does the sending. You own both layers outright, so a diagram missing them is describing a campaign rather than a system.
8.2.5 Color the diagram by who owns each layer.
Draw the six layers you own in one color and the four you rent in another, and the picture stops looking like a pipeline. What shows up are the backward arrows: send record into state, opt-out into the suppression store, reply classification into approval, outcome into feedback. Then feedback into the offer library and the ICP file.
8.2.6 Agents start every run knowing nothing.
A managed agent session starts with only its system prompt. A routine, a saved Claude Code job Anthropic's servers run on a schedule or trigger, starts every run from zero because the working copy gets rebuilt each time. Nate, a free-community operator with affiliate ties, openly critical of the product he was demoing, says it on camera. Each time the agent wakes up, it is completely stateless. An agent that starts blind cannot know whom it already contacted, an architecture problem that takes an architectural fix.
8.2.7 The same toolchain logs content and not outbound.
One recorded content build comes from the Boss AI channel, which sells an AI academy and takes a Blotato affiliate fee. It opens by telling the agent to log every published post with platform, date, tags and live URL. Point the same toolchain at outbound and you get nothing, across all fourteen named-tool builds.
8.3 Own the layers that compound and rent everything else.
Own the six layers that compound, buy the four that change fast and carry reputation risk, and build only the stable glue between them. Jason Lemkin and Amelia Lerutte of SaaStr, who sell conference tickets and sponsorships, run twenty to thirty agents and vibe-coded apps in production. Their rule is buy 90% and build at most 10%, and of their own flagship internal build they say, "I wish we hadn't."
8.3.1 Re-price build versus buy every quarter.
The field's cost model names three spend tiers, and building only pays at the top one. The minimal solo tier, one operator on a bought stack at $150 to $400 a month, saves about $55 by building, not worth the setup hours. The serious solo tier adds agents at $600 to $1,200 a month, where building saves about $236, or 29% of the bought equivalent. The small-team tier runs roughly $2,000 to $4,500 a month, where building saves about $2,500, or 55%. Those savings are one report's estimates, not invoices, and none holds long, because the category invalidates its own documentation roughly every three months.
8.3.2 Never build the database or run your own mail server.
Across the thirty-two recorded builds complete enough to judge, not one builds a contact database. And across the fourteen with named tools, not one sends from self-hosted mail infrastructure. Anthropic's own free Small Business plugin connects to eleven named services while building none. Operators name one sending-side reason, the 4-to-8-week fleet warmup on the clock Chapter III sets, though the one-or-two-mailbox build ramps in days.
8.3.3 Rent enrichment and expose it as one function.
A waterfall tries provider after provider until one returns a hit, and Chapter IV prices what it buys. It buys about one point of coverage over the best verify-first finder at four times the false-positive rate, and a whole category over a contact database. Either way, buy the waterfall as one API exposed to the agent as a single function returning a verified result or nothing. Ordinary code outside the agent owns retries, rate limits and verification gating.
8.3.4 Own composition, approval, the offer library and feedback.
There is no condition under which renting these four becomes correct. Composition is where your voice and evidence live, and approval is where your accountability lives. The offer library is the only place your unit economics are written down, and a rented feedback loop optimizes toward whatever the vendor decided counts as winning.
8.3.5 Own your state and suppression stores, not contact data.
The test is who the record is about. State and suppression record what you did and what you promised, and those facts do not rot. An enriched contact record is a decaying guess about a stranger, losing roughly 2.1% of its accuracy a month. That is why a list bouncing under 1% at 30 days bounces 5 to 8% at 90.
8.3.6 Choose a sequencer on what an agent can write.
Every sequencer markets deliverability, so that tells you nothing about which to pick; what differs is how much the agent is allowed to write. Each tool exposes a Model Context Protocol server, the standard interface a tool offers an agent. Some cannot manage a campaign at all while others create sequences and push contacts into them.
8.3.7 Spend your operator hours on the versioned context files.
Keep your ICP definitions, voice profiles, exclusion lists and playbooks as files with a commit history. It is the field's most common component, present in twelve of fourteen builds. The six owned layers compound because you keep writing decisions into them, and this file set is where the writing lands. It is also the only layer anyone claims improves its own output with age. Anthropic puts that as Claude on day 30 beating Claude on day one, while everything else stays flat.
8.3.8 Decide what is shared before the second operator arrives.
Suppression, the audit log, the state store, the offer library and the context files stay at exactly one copy no matter how many people run the campaign. Research, drafting and review get one copy per person. Decide this after you add someone and you end up with one suppression list per operator, and that is not suppression.
8.4 Give the send credential to code, never to the agent.
A guardrail is a check that blocks a bad action before it happens; a prompt rule is not one, because it states an intention without enforcing anything. Only three gate types appear anywhere in the field's send paths: a billing prompt, an API rate limit and an interactive approval step. All three erode as credits get cheaper, rate limits rise and schedulers improve. The one send gate that survives is one where the agent physically cannot send. A separate piece of ordinary code holds the credential and checks suppression before it uses it.
8.4.1 Blast radius decides where the autonomy boundary sits.
Blast radius is how much damage one wrong action does before anyone can stop it, and sorting the record by it gives an exact pattern. Of every autonomy boundary that shifted between 2024 and 2026, the ones that moved covered recoverable actions invisible to the buyer. The ones that held covered actions that were irreversible, visible to the buyer, legally exposed, or something a person has to answer for. Better models moved not one boundary in that second set, so that set is where your gates belong.
8.4.2 Every gate in the field is an accident of billing, rate limits or scheduling.
Six of the fourteen named-tool builds prompt the operator before a metered API call and call that prompt their safety mechanism. The billing prompt is the first of the three gate types, which is why LinkedIn connection requests, where no credits burn, have no gate at all. The second is an API rate limit stopping the agent short of send. SaaStr's chief AI officer called that on camera a failure of the APIs rather than of the agent. The third, an interactive approval step, vanishes the moment the work moves to a scheduled task, because a scheduled job cannot stop to ask.
8.4.3 The agent proposes and the code decides.
False safety escalates from "I told it never to delete my database" to "I block all delete statements" to the agent writing a script and running the script. In a 2026 academic study, 31.89% of the model-generated rules about what not to do were ungrounded, and 21.93% were actively harmful. Put the rule somewhere the model cannot reach it.
8.4.4 Put the suppression check inside the send function.
A sequencer's stop-on-reply feature protects the sequencer, and it knows nothing about your LinkedIn tool, your dialer or a manual send. List every way a message can leave your building: sequencer, cold-email tool, LinkedIn automation, manual sends, dialer, marketing automation, event lists, agent sends, and API or webhook sends. Then make each one call the same suppression store and wait for the answer before sending. The ceiling on getting this wrong is the $2.95M civil penalty of August 2024, covering unhonored opt-outs across more than thirty million emails. That stands against a CAN-SPAM statutory maximum of $53,088 per email.
8.4.5 Put an idempotency key on every send path.
An idempotency key is a unique tag on each send that lets the system spot a retry, so a crashed job resuming does not email the same person twice. The need comes from the scheduling surfaces. Desktop scheduled tasks play catch-up on missed runs when the machine wakes, and routines fired by an API call carry no built-in run identity. So a retried trigger produces a second run and a second send. Build the key from recipient, campaign and step, and check it in the code that holds the credential, not in the agent.
8.4.6 Split approval across two routines.
A running routine has no pause-and-ask mode, so routine A drafts into a review channel and stops there. A human approves, and that approval fires routine B, which sends. That split keeps a human in the loop on a surface built with nowhere to put one.
8.4.7 Split memory stores by who can write to them.
An outbound agent reads untrusted text from websites, profiles and inbound replies, and Anthropic's own documentation names the consequence. Stores attach read-write by default. So a successful prompt injection, an attacker planting instructions in text your agent reads, writes into a store that every future session reads as trusted memory. Mount suppression, exclusions and standards read-only, and make the writable store the narrowest per-contact scratchpad you can define, least privilege applied to memory rather than to tools.
8.5 Without the return paths you shipped a campaign.
Seven arrows point backward: send record to state, opt-out to the suppression store, reply classification to approval, outcome to feedback. Then feedback to the offer library and the ICP file, and state back to targeting so you can re-approach old leads. Thirteen of fourteen builds have none of them, and four out of a further eighteen have some.
8.5.1 Check the calendar, not the reply count.
Exactly one build in the field closes its loop honestly. A positive reply fires a webhook that wakes an agent with calendar access, and that agent checks whether the meeting actually got booked. The only other closed loop optimizes against reply rate, and a machine can learn to maximize replies while producing no meetings at all.
8.5.2 Write the opt-out before you draft anything.
Order the reply pipeline so that classification writes the opt-out to suppression first and drafting happens after. A crash, a timeout or a bad generation then cannot cost you an opt-out, because the one irreversible legal obligation has already landed.
8.5.3 Stamp offer variant, agent version and model version on every send row.
Six months from now you must reconstruct what was sent to whom and why. Without the offer variant on the row you cannot retire a losing offer, which means comparing booked-meeting rates across each offer's sends. Without the agent and model version you cannot tell a copy change from a quiet model dot-release. One such release made SaaStr's pitch-deck analyzer invent revenue figures where data was missing. SaaStr eyeballed the rate at about 5% of four thousand decks, with the code unchanged.
8.5.4 A human approves every change to the copy.
One recorded build pulls yesterday's results, compares them to a baseline, rewrites both the copy and the list filter, and relaunches itself unattended. It runs on two enterprise accounts at 10% of volume with no significance test anywhere in the cycle. On 10 replies out of 200 sends, a Wilson 95% interval runs from 2.7% to 9.0%, spanning everything from failing to top-decile. And that 5% is a researched-band rate, so the noise worsens at every rate beneath it.
8.5.5 Copy the daily dashboard loop that content builders already run.
Brock runs an AI-for-non-techies channel with affiliate fees from Clay, Firecrawl and Higgsfield. Every day in production he runs a scheduled task into a persistent database into a live dashboard, exactly the return path the outbound builds lack. His points at YouTube view counts, and aiming the same loop at reply, bounce and meeting data is a configuration change, not an engineering project.
8.5.6 Wire the two cheap return paths first.
Your sequencer already computes per-campaign numbers and exposes them, so making that data readable by the agent is the cheapest return path available. The second is a scheduled six-month follow-up to cold leads, pulled from your own state store, and both run on tools you have already paid for.
8.6 Use the Claude controls that already shipped.
Builders appear not to know Anthropic's permission model exists, rather than having looked at it and rejected it. One creator with a large Claude audience says most of his viewers have never heard of either of the two settings, disable-model-invocation and the permission deny rule. Those are the settings that decide whether a skill can fire itself.
8.6.1 Three send controls already ship, plus a worked example.
The three are disable-model-invocation, permissions.deny kept under version control, and the two-routine split. Anthropic's own skill catalog also holds invoice-chase, the closest thing in it to an outbound campaign. It ranks a list, matches tone per segment, drafts, stops for approval, and sends only as a separate explicit step. And its setup dialogue asks you, unprompted, when it should and should not send on its own.
8.6.2 Two lines of config, and one survives bypass mode.
disable-model-invocation: true means only a human can fire that skill. Anthropic documents it in exactly those terms for higher-risk things, such as a skill that sends a message. permissions.deny lives under version control, is inherited by every operator, and holds even in bypassPermissions mode. That is the mode people switch on when they are in a hurry, which is exactly when this matters.
8.6.3 A webhook plus a saved session gives you reply triage.
Isabella He of Anthropic's Applied AI team described the capability in May 2026. An external event arriving by webhook can resume a saved agent session or push it into a specific state. Point that at an inbox and an incoming reply wakes that prospect's session, full prior event log intact, to classify the reply, write the outcome and draft a response. Reply triage is missing from all fourteen named-tool builds. Four of the further eighteen run it by routing a webhook into a workflow tool or scheduled routine with chat-app approval, and nobody has assembled the resumable-session version yet.
8.6.4 Build CLAUDE.md first and tool integrations last.
Anthropic names building integrations first as the most common mistake teams make. The order that works runs from the instructions file to hooks, then skills, then plugins, then Model Context Protocol servers. Each layer decides what the next one needs, and integrations built before the instructions exist get rebuilt.
8.6.5 Run a separate evaluator, because self-review fails.
Anthropic says this against its own interest. An agent asked to judge its own work will likely praise it even when a human can see the quality is mediocre. And Claude out of the box has been seen identifying real problems and then approving the work anyway. Make the evaluator a separate invocation with its own rubric file, weighted toward where the model is weakest, originality and craft, to catch slop.
8.6.6 Score with yes-or-no questions, and script what you can.
Numeric self-scoring collapses to the middle of the band: a 1-to-5 rubric lands on 3 or 4, a 1-to-10 on 6 or 7. In published work by Stureborg and colleagues, a model scoring several attributes at once moved them almost in lockstep, correlating at 0.979 where 1.0 means identical. Human raters moved nearly independently, at 0.315. Ask yes-or-no questions on named dimensions instead, and convert every step you can from model inference into ordinary code.
8.6.7 Separate what you know from what patches the model.
ICP files, voice profiles, deny-lists, evaluator rubrics, the offer library and suppression classes hold what you know, and they compound. Context resets, rigid orchestration and retry scaffolding work around a model's current limits. Anthropic's own harness team built three such patterns for one model generation and deleted all three in the next release. Put a three-to-six-month review date on the second category and none on the first.
8.7 Inference is cheap, and the architecture creates nothing.
Two production go-to-market agents run at $257 a month, while the costs that actually constrain you are enrichment credits, sending infrastructure and human review time. None of those three falls when token prices fall, and once built the whole thing generates no value on its own.
8.7.1 Tokens are a small share of the cost per prospect.
On a cheap-and-wide campaign, researching and writing to one prospect costs roughly $0.002 to $0.01 in tokens against $0.10 to $0.30 all-in. So tokens are 2 to 10% of the cash. On a deep-research campaign, tokens run $1 to $4 against $3 to $9 all-in, or 33 to 44%. Both ratios come from a report author's cost model, not an invoice, and neither decides anything, because the constraining costs sit outside both.
8.7.2 The argument for buying is never inference cost.
Two go-to-market agents in production cost $257 a month in inference, about 95% of calls going to a small model at under a cent each. That is roughly one eighth of the sending-infrastructure line and about a hundred times less than the salaried headcount displaced. The same operator's fully burdened figure, $500 to $800 a month with supporting licences counted, is still not the constraining number. If a vendor's pitch rests on your model bill, they are selling against the wrong number.
8.7.3 Budget the three costs that do not fall.
Token prices fall about tenfold a year, and these three do not follow. Enrichment credits run roughly $0.70 to $3.75 per fully enriched contact on Clay, the credit-metered enrichment platform most of these builds sit on. Deliverability infrastructure runs about $0.35 to $0.50 per prospect, and operator review time, the largest line, is priced in hours rather than dollars. Budget those three and let the model bill fall wherever it wants.
8.7.4 Route calls to a small model and cap every run.
The production pattern is roughly 95% of calls to a mini-class model, with the ICP context cached, plus a hard spend cap on every run. Agent loops without bounds multiply cost about tenfold, and in one self-reported case Clay's research agent burned about $3,000 in its first month with no cap on it.
8.7.5 The architecture makes you honest, and it cannot make you informed.
Narrow, researched outbound produces a few hundred sends a month per campaign, 250 to 300 at steady state on the floor Chapter III sets. That holds no matter how many people run it, and at that volume no instrumentation gives a trustworthy answer about which message worked.
8.7.6 The architecture creates nothing on its own.
A bad list, a bad offer or a missing sender reputation passes through all twelve layers unchanged and arrives intact at the far end. Nate built the most completely specified system anyone recorded, in a video made with Clay's own team, and he says against that interest. Even if Clay gives you the best possible data, that does not mean you are going to get clients.
8.7.7 The responsible build is also the effective build.
Suppression protects the recipient, protects your deliverability, and prevents the double-send that burns your domain. The audit log is your compliance record, your debugging tool, and the only place outcomes can be measured. Cutting generative steps reduces slop and breaks up the message similarity behind the AI spam penalty. That is the roughly 2.7x flag rate the two vendor tests in Chapter VII measured.
8.8 Owning your memory and hiding the credential are one decision.
Owning memory and judgment while renting execution, and putting the sending credential behind code rather than a prompt, are two angles on a single decision. The six layers you own are exactly the ones no rented tool can check on your behalf. Ordinary code holding the credential is where that ownership stops being an intention. Skip this and you rent suppression along with sending, then read a reply count off a rented tool while trying to count meetings that actually happened. That is how thirteen of fourteen builds shipped the acquisition half with no way to know whom they had already contacted.
Count the agents doing real work, not the ones installed.
The number of agents you run is a vanity metric. What counts is your working agents: the ones whose output a person checks and acts on the same day. Human hours cap that number, not model quality, so growth comes from adding operators who each carry a small set.
9.1 Five to ten checked outputs a day is one person's ceiling.
One operator can check and act on roughly five to ten pieces of AI output per day. The figure is Nick Saraev's own daily count, stated on camera in July 2026. He runs an automation agency and sells automation education, and it is not a survey of operators. It still caps everything downstream, because the limit is the human doing the checking, and adding agents does not move it.
9.1.1 A person can handle four to six working agents as "direct reports" whose work they check.
Split a five-to-ten-output day across agents that each need review and you land at four to six per person, a derived figure, not a field measurement. Treat them as direct reports, because what you are budgeting is your own review time. Amelia Lerutte, chief AI officer at SaaStr and the one operator to describe her daily routine in detail, runs four to five core agents. And a fifteen-to-twenty-hour weekly maintenance budget carries the same four to six: two roads, one small number.
9.1.2 An agent counts only if you check it today.
An agent counts when a person verifies its output and acts on it the same day, a deliberately strict test. Personio, the HR and payroll platform, built 400 internal AI assistants, and the top ten deliver about 80% of the value. Its chief revenue officer Philipp Lohr disclosed those numbers on record in January 2026.
9.1.3 Each new agent costs two weeks you never get back.
Standing up an agent takes roughly two weeks of onboarding, down from the month or six weeks it took SaaStr chief AI officer Amelia Lerutte early on. Chapter III describes that window as a blackout in which her existing agents degrade. Her stated ceiling is one to one and a half new agents a month; the planning default rounds it down to one per operator per month. That is one operator's production experience, not a modelled number: worth having, not a law.
9.1.4 The daily attention each agent needs never ends.
Lerutte spends 10 to 60 minutes with each of her four to five core agents every day, because weekly review is too slow. The picture goes stale before the week ends. Each added agent subtracts from that operator's outreach and content time, so budget the attention as a recurring monthly cost, not a one-time build.
9.1.5 Managing agents costs the same hours as doing the work.
Jason Lemkin and Amelia Lerutte of SaaStr say on record that they each spend 15 to 20 hours a week on maintenance. What they maintain is twenty to thirty agents and vibe-coded apps in production. SaaStr sells conference tickets and sponsorships to the vendors whose agents they run, so they had every reason to report the opposite.
9.1.6 Plan inside the time envelope that actually exists.
Per operator, the week holds 15 to 20 hours of agent maintenance and 45 to 60 minutes a day of outreach. It also holds 30 to 60 minutes a day of content. The maintenance figure comes from Lemkin and Lerutte, the outreach from Nick Saraev, and the content from the writer Dan Koe, who sells a writing course. The envelope is assembled from three sources, not observed in one measured week, and it is already full. A plan that assumes more time will quietly drop one of the three, and the one it drops will be content.
9.1.7 Deeper research lowers the number of agents you can run.
The ceiling counts outputs a human can verify, and a deeply researched account brief takes far longer to check than a templated line. So the more research each message needs, the fewer agents one operator can carry: depth and breadth bid for the same scarce hour.
9.1.8 Automate a step only after you have done it by hand and stopped changing it.
The manual version of a step is the specification the agent gets built from, so skipping it means automating a process you never understood. The LinkedIn research says it outright: "No tools, no AI: learn the motion by hand so you can later judge whether an agent's drafts are good." Skip it and the two-week onboarding above goes to discovering by correction what doing the step yourself would have taught you. Chapter II's self-optimizing loop that taught itself to produce replies is this mistake at full speed. Do it by hand, stop changing it, then hand it over, and keep the judgment calls even then.
9.2 A headline agent count measures nothing.
Any headline count sweeps in idle agents, which cost nothing to accumulate. Personio's 400 assistants with ten carrying about 80% of the value is what "we run hundreds of agents" looks like. Author's note: I have seen the same shape at five times the scale. In one enterprise deployment I worked on, associates built roughly 2,000 agents and about 25 of them drove 98% of token usage. Token spend is not value and the overlap is not causation, but the concentration was the same, and that is the part that keeps recurring.
9.2.1 Idle agents pile up for free.
Nothing forces you to retire an agent that stopped mattering, so the pile grows, never shrinks, and inflates every published ratio. Treat any agent count you are quoted as an upper bound, because the fraction actually working is unknown.
9.2.2 Audit every quarter and retire the agents nobody checks.
Each quarter, list which agents produced output a human verified and acted on, then retire the rest, benchmarked against Personio's ten valuable assistants out of 400. An agent that still fires still eats review attention out of the same five-to-ten checked-outputs-a-day budget.
9.2.3 Price an agent the way you price a hire.
Personio books roughly $100,000 a year per AI SDR agent. The figure carries weight because it came from Lohr, the buyer's own chief revenue officer, not a vendor price list. That is a hiring decision wearing software clothes, so scrutinize it like a hire, not a SaaS renewal.
9.2.4 SaaStr's whole agent stack still missed its own email target.
SaaStr, the B2B software events and media company, is the most instrumented operator here, running on two to three people at roughly $15M in revenue. In February 2026 its chief AI officer, Amelia Lerutte, put the count at twenty agents and vibe-coded tools in production, almost thirty that morning. The total sweeps in dashboards counted as agents, so read it as agents and apps together. Her own agent tallied the machine's real output and told her she had sent 87 emails against a 200-email target over about six weeks. The same stack sent 83 personalized emails unattended at 12:20am and later 100 in ten minutes, so machine throughput never ran out. Those sends went to about a hundred existing sponsors, assembled from records the sponsors were already committed against, while the cold outbound agent stayed in draft mode.
9.2.5 Two large operators reached the same verdict from opposite ends.
Lohr counted Personio's assistants and found the value in ten of 400. Lerutte's own agent counted SaaStr's output and found 87 sends against a plan of 200. Neither organization knew of the other, so the agreement is worth more than either number alone.
9.2.6 Work that lives in one person's head cannot be handed over.
Lerutte says the logic routing contacts to her outbound agents lives entirely in her head. Segment rules, prompts and hard-won corrections sit in memory, not in the versioned context files Chapter VIII names among the six layers worth owning. A second operator inherits none of it, because context files compound while the agents above them get replaced.
9.3 Grow by adding operators and keep the architecture the same.
Growth is operators multiplied by the working agents each carries, and every term except operator count is fixed by human review time. The architecture does not change with size: a 700-person marketing organization runs the same shape a single person does.
9.3.1 Model growth as operators times agents times checked outputs.
Write it out: operators, times 4 to 6 working agents each, times 5 to 10 checked outputs a day. In campaign terms that budget carries one campaign comfortably and two at full load per operator. Only the first term is yours to change, because the other two are set by how much work a human can check.
9.3.2 Enterprises run the same architecture, just more of it.
Denise Persson, chief marketing officer of the data-cloud company Snowflake, spoke on record in June 2026. She said a central AI and data organization of roughly 20 people serves sales and marketing, with about two more sitting close to marketing. The marketing organization is about 700 people, and Persson says the central team cannot handle the incoming demand.
9.3.3 Verification, not compute, runs out first at every size.
Compute is cheap and abundant at every size, while the hours a person can spend checking output are neither. So large operations converge on what a solo build converges on: one shared foundation, one suppression store, one audit log, one context library. Nobody who has run this at scale ends up with a thousand independently built agents.
9.3.4 The enterprise and the solo operator drew the autonomy line in the same place.
Snowflake fully automated the account and use-case prioritization behind its expansion revenue. It deliberately kept a human in the loop on campaign activation and parts of its digital media buying, which is precisely where the one-person builds stop. When a 700-person organization and a solo operator independently pick the same line, the line is structural, not cultural.
9.3.5 Low volume comes from the campaign, not from the company.
A campaign here means one list paired with one offer. One built on heavy research and a narrow list produces a few hundred touches a month. That holds whether a single person runs it or Snowflake's 700-person marketing organization does. The small-sample limits hit them exactly as hard as they hit you, including the fact that you cannot A/B test a message.
9.3.6 A big company runs more campaigns, not bigger ones.
Size buys parallelism, not escape, because each campaign is still capped by its own operators' ceiling on checked outputs. Nor do the campaigns add up into one large sample, since each runs a different list and a different offer.
9.3.7 Run one suppression store across every operator.
Run one suppression store, one audit log and one state store for the whole operation. The moment two operators each keep their own list, someone who opted out of one is still live in the other. And the recipient experiences your company as a single sender who ignored them.
9.3.8 The person you add has to build and check.
Add people who can build and verify their own agents. Send capacity already comes from the agents and is not scarce; verification capacity is, and it only arrives attached to a person.
9.4 Buy the agent-to-operator ratio you can verify, not the one the vendor demos.
Vendor demos quote the size of the pile, because size is impressive and cheap to produce. Ask instead how many agents produce output a human checks and acts on daily, price the purchase in operator-hours, and assume you will want out early.
9.4.1 Ask how many agents a human checks daily.
Any vendor can put a large agent count on a slide, because idle agents cost nothing to keep. Make them say how many produce output a human verifies and acts on the same day, then make them describe that human's day. And ask for a dry run, a full pass with sending switched off, so you can watch who checks what before a message reaches a recipient.
9.4.2 Price the build in operator-hours rather than dollars.
The real cost is the two-week onboarding tax per agent plus 10 to 60 minutes of daily attention forever. Both come out of the same operator whose checking speed caps the campaign. Dollars are the small line item, so convert every quote into hours before comparing it to anything.
9.4.3 Nobody has audited the churn figure, so buy as if you are leaving.
The 50 to 70% churn figure everyone quotes for AI SDR products has never been audited. It circulates under five attributions: Leadriver, the consultant Michael Saruggia, unnamed practitioners relayed by Kwanzoo, Firstsales, and the signal vendor UserGems. Chapter III weighs that repetition against the incompatible windows it is quoted under and lands on a day-90 break clause you intend to exercise. Chapter I prices the one hard case, TechCrunch's March 2025 reporting on 11x, where most of the claimed revenue did not survive that clause. The revenue gap is the solid half, since the churn percentage attached to it rests on anonymous sources. Buy as if you are leaving: no annual commitment, a break clause you will use, no daily dependency that dies with the contract.
9.5 The verification ceiling is a staffing plan, not a limitation.
Counting only the agents doing real work gives you a number to plan against: four to six direct reports per operator, five to ten checked outputs a day. Operator count is the only term you control, because a better model does not lift the ceiling and growing the pile buys nothing. Owning your memory and judgment while renting the execution is what lets a second operator inherit anything at all.
If your rules are not in code, you do not have rules.
Every safeguard in the production builds examined here was either chosen on purpose or fell out of billing, rate limits and scheduling. The accidental ones went away as soon as the underlying limit got cheaper, so ask which line of code enforces each of your rules.
10.1 Check whether a human ever chose each of your send gates.
A gate here is a human in the loop: a person who approves a message before it sends. Chapter VIII catalogs the gates this field runs on and finds only three in the whole send path, none designed as a safety control. The question about your own build is not what each gate stops but who chose it, and what is left when the answer is nobody.
10.1.1 Your ethics only hold if you have time to review.
An approval gate only counts as a control if the human has time to catch the mistake before it ships. Chapter IX sets that budget at a handful of checked outputs per person per day, and above it the gate turns into a click.
10.1.2 Your send gate may just be a credit meter.
Chapter VIII records the builders who call a billing prompt their safety mechanism. That prompt watches a credit balance draining, not a recipient about to be contacted, and the tell sits where no credits burn: there the same builds prompt nobody. A gate that tracks your invoice vanishes the day the credits get cheap.
10.1.3 A rate limit is not a safeguard anyone chose.
The second gate type is an API rate limit that stopped an agent short of send. Chapter VIII records the operator who hit it reading it as a vendor failing, not a safety decision. Buying a higher limit deletes the safeguard, and nobody who never chose it is watching for the day it goes.
10.1.4 Scheduling a job quietly deletes the human step.
The third gate type is an interactive approval step, and Chapter VIII describes how it vanishes once the work moves to a scheduled task. Nobody removed it and nobody weighed the trade, because no human was present to interrupt.
10.1.5 A gate that changes nothing is not review.
One gate in this set was chosen on purpose and decayed anyway. Nick Saraev, who sells automation-agency training, placed a copy review before launch, then said on camera that "half the time I don't actually make a change". That is an admission against his own commercial interest. A review step that produces no edits is a click, so measure your gate by how often it changes the output, not by whether it sits on your diagram.
10.1.6 Put the stop where the action cannot be undone.
The only chosen gate still holding belongs to an operator who placed it by blast radius, meaning how far one mistake reaches. He did not place it by how dangerous the tool call looked. He publishes as TopFunnel, runs several client projects, and discloses that he helped build the Instantly integration he shows. He runs the model with permissions wide open and still puts a manual stop at campaign launch, the irreversible act the buyer sees. Place your own stop at the same kind of step.
10.1.7 A missing consent step tells you the liability is yours.
Voice tooling forces opt-in because the US Telephone Consumer Protection Act puts the exposure on the platform. Email and direct messages carry no such step, so the exposure there sits on you. Read a consent step as a map of who is liable, then decide for yourself what is right.
10.1.8 Every runaway send on record came from an agent misreading something.
Two mass-send incidents are documented in this material. Jason Lemkin of SaaStr asked an agent to send roughly 1,500 invitations five minutes before a keynote. It used a sending address a rule in its own memory forbade, and he reported the failure on stage against his own interest. Nate Herk, who sells a paid automation community, had an agent misread its task list and email his whole list a discount code that was never meant to ship. Neither came from bad intent, so hardening only against bad intent guards a flank nothing has come from: put the budget on misreading.
10.2 Put every rule in code, or you only have an intention.
A rule written in a prompt is a request the model can decline, and Chapter VIII measures how often it does. Suppression, send authority and every action you cannot undo belong in code the model calls, not in text it reads.
10.2.1 Never let the model write the rules meant to restrain it.
Chapter VIII carries the study behind this, a 2026 arXiv preprint numbered 2601.11908, which measured 31.89% of model-written prohibitions as ungrounded and 21.93% as actively harmful. Carry its caveat: the research report supplying those figures could not confirm the identifier independently, so check the paper before quoting the numbers to a client. A separate 2025 benchmark, AGENTIF, arXiv 2505.16944, found constraints failing even when their triggering condition was present. You cannot hand the writing of your rules to the thing they exist to restrain, so suppression has to live in code.
10.2.2 Rules written in a prompt fight each other.
Constraints in a prompt compete for the model's attention and degrade as they pile up. One documented agent had completed roughly 5,000 successful jobs, then began failing when a fifteenth guardrail was added. A suppression table checked by code carries its fifteenth row and its ten-thousandth at no cost, so that is where the rules belong.
10.2.3 Keep one suppression store and check it everywhere.
Own a single store with immutable versions and a 30-day audit trail. Give it a redaction path, so an erasure request can be honored without losing the fact that a suppression exists. Give it SHA-256 preconditions too, a fingerprint check that fails a write if the file changed since you read it, so concurrent writes cannot silently overwrite each other. Chapter VIII lists every path a message can leave your building by, and each one asks this store first. Fragmented suppression is the costliest documented failure in this field, and the opt-out section below prices it.
10.2.4 Start every campaign paused and make replays harmless.
Default every campaign to paused, and open each one with a dry run, a full pass with sending switched off, scored for how reversible its actions are. Chapter VIII builds the idempotency key, a fingerprint that makes a repeated request do nothing the second time. You need one because replays are a certainty here, not bad luck.
10.2.5 Give the reviewer evidence, not a button.
The review payload should carry a classification bucket, a confidence value and the model's written reasoning. That is exactly what one production reply-triage build ships, run by the Outlier GTM agency, which sells done-for-you outbound.
10.2.6 Build three primitives before you build anything clever.
Three primitives, small parts you build once and reuse everywhere, carry most of the weight. Chapter VIII builds all three: the one-hop cap on model-rewritten text, the calendar check that replaces counting replies, and the append-only send log.
10.2.7 Break the lethal trifecta with architecture.
The security researcher Simon Willison named this in June 2025. An agent that holds private data, reads untrusted content and can communicate externally can be tricked into shipping your data out. Any two of the three are safe together and all three are not, so break the set in your architecture, not your prompt. The cheapest break is least privilege, the smallest access that does the job: the agent that reads never holds the credential that sends. Chapter VIII shows how to split the memory stores and where the allow-list goes.
10.3 The responsible build and the effective build are the same.
Chapter VIII shows each of these controls, the suppression store, audit log and gates, paying off three ways at once: the recipient's experience, your deliverability, and your ability to measure anything at all.
10.3.1 An irrelevant send costs you the buyer, not just the send.
The cost has been measured twice, and both readings are large. Gartner surveyed 632 buyers in August and September 2024 and found 73% actively avoid suppliers who send irrelevant outreach. A separate Gartner survey of 1,464 buyers found personalized marketing produced negative experiences for 53% of customers, who were then 3.2 times more likely to regret the purchase. Price every bad send as future pipeline you have already spent.
10.3.2 One suppression store pays you back three ways.
Chapter VIII sets out the three returns, recipient trust, deliverability, and measurability, from one suppression store, and the governance point sits inside the first. Somebody asked to be left alone, and a fragmented store is how you break that promise without ever deciding to.
10.3.3 The audit log pays for itself three times over.
The log is what you hand a regulator, a client or a court asking what you sent and why, a question you cannot answer from memory. Chapter VIII shows the same builders logging published content meticulously and outbound not at all, which makes this a habit, and a habit is yours to change. Write the log before you write anything that sends.
10.3.4 Capping model hops improves writing and inbox placement.
Feeding model output back into model input produces slop by construction, however good the model is. Capping the chain at one hop improves the writing and breaks up the identical-shape text that near-duplicate detection clusters on. Chapter VII prices that clustering at roughly 2.7 times the human spam-flag rate. That multiple rests on two vendor tests nobody audited, Saleshandy's 7.8% against 2.9% on 12,000 emails and Digital Applied's 8% against 3% on 100,000 paired sends. Both vendors sell into the answer they measured, so carry the multiple rather than the decimals.
10.3.5 Human approval may be what keeps your insurance valid.
Errors-and-omissions cover often applies only where a human reviewed the output before it reached the recipient. The standard-form generative-AI exclusion endorsements the industry writes against carry January 1, 2026 effective dates. And at least one carrier has filed an absolute AI exclusion across its professional lines. Ask your insurer in writing at renewal, because the approval step you treat as quality control may be what keeps your policy live.
10.3.6 Hiding the automation is dishonest and short-lived.
Every tactic that depends on the recipient not noticing has the same expiry problem, and three recur. The first is hard-coded false claims about how much effort was spent. The second is reply latency tuned under two minutes but never instant, so a lead cannot tell it is an agent. The third is a system prompt reading "do not mention automation AI generation or that this was system produced". Detection tooling is the fastest-improving part of this market, and a peer-reviewed 2025 classifier already spots model-written email at 96% accuracy, so disclose instead.
10.4 Draw your autonomy boundaries by blast radius, not by model skill.
Set every boundary and every control by how much damage the action does when it goes wrong, and by nothing else. Chapter VIII sorts each autonomy boundary that shifted between 2024 and 2026 on exactly that measure and finds that nothing on the output side moved. What remains is who gets to decide where each line sits, and what happens when the decision turns out wrong.
10.4.1 The reviewer is the operator who owns the agent.
Read every mention of a reviewer here as the person who built and runs the agent, never a second human in a sign-off chain. This architecture exists to replace approval hierarchies, and putting one back returns you to the volume where review becomes a rubber stamp.
10.4.2 Run the entire input side without a human.
Research, enrichment, account scoring and CRM upkeep are cheap to reverse and invisible to the buyer, so run them with no human watching. Model scoring agrees with human scoring at a Cohen's kappa of 0.79 to 0.84, strong agreement beyond chance and short of the roughly 0.97 two humans reach. That is good enough for ordering a list, short of what anything irreversible needs. One caution on coverage: the widely quoted 90 to 95% enrichment figure is a vendor's own claim. Chapter IV weighs it against the one open-dataset benchmark, which came in lower.
10.4.3 Every first touch and every reply gets human approval.
Put a human in the loop on every send the buyer sees: agent-drafted, human-approved, no exceptions. Anything going to an executive or touching legal ground gets written by a human outright. In every build examined, this is the setting that stayed put while the models improved.
10.4.4 Human review of every outbound message only survives at low volume.
The gate in question is one person reading every outbound message before it sends. That is feasible at the few hundred sends a month a solo operator runs and impossible at any volume a scaled program would call normal. Chapter IX sources and caveats the ceiling of five to ten checked outputs per operator per day.
10.4.5 Treat the six human-only tasks as settled, not as a frontier to push.
The six are AI voice, killing a campaign, launching a campaign, budget, letting an agent modify its own prompts, and legal review of any new claim. The list comes from a commissioned research report dated August 2026 that sorted fifty outbound tasks by autonomy level and compared the 2024 defaults against the 2026 ones. Every one of these six sat with a human in both years. The report owns its per-task assignments as engineering judgment rather than measurement, so read the list as a well-argued default, not a finding.
10.4.6 Suppression is the one safe high-stakes automation.
Suppression is the only place where serious consequences and full autonomy coexist safely, because executing a suppression rule is a deterministic lookup, not a judgment call. Note the split: executing an opt-out is safely automatic, while spotting an opt-out phrased in free text inside a reply still needs a human.
10.4.7 Low volume is a moat, not a handicap.
At the roughly 1,000 approval-eligible events an hour that 50 agents produce at 20 tool calls each, review can only be a click. And routing even 10% of that to humans takes three or more full-time people clicking. Undifferentiated volume is what pushes you over the line, and a campaign built on a narrow list stays under it. That holds whether you run it alone or inside a marketing team the size of Snowflake's, at about seven hundred people.
10.4.8 Design against the human tendency to agree.
The tendency has a name, automation bias: a reviewer drifting toward whatever the machine already decided. A 2025 review of thirty-five studies reports reviewers routinely agreeing with incorrect AI recommendations, and Harvard Business School work reports that clearer AI explanations made people defer more, the explainability paradox. Both arrive with their underlying papers unnamed, so treat them as direction rather than numbers you can check, but act on them because the countermeasures are free. Force the human to decide before revealing the model's answer, show evidence rather than summaries, and hide confidence scores. On that last point, a 2026 measurement by the vendor Digital Applied, a marketing site flagged as directional, puts a claimed 90% confidence at roughly 75% actual accuracy.
10.5 Rank your risks by what is actually documented.
The compliance side of this field is evidenced down to named settlements and per-email penalties. The security side is still mechanism and proof-of-concept with no public incident behind it. Build against both, and order your spending by the evidence rather than the headline size.
10.5.1 Work the risk list in its documented order.
The order the evidence supports is close to the reverse of the order the market panics in, because the largest possible number is rarely the most likely event. Programs that start from the largest number spend their first year defending against the one risk that has never fired in this channel.
Mailbox providers are your first risk and the real regulator.
Getting blocked by Google, Microsoft or Yahoo is the near-certain failure, and it kills the channel outright, with no appeal and no lawyer. The bulk-sender rules Chapter I details set 0.3% as the spam-complaint ceiling and 0.1% as the target. Treat 0.3% as a wall you never approach, and run the program against the 0.1%.
Private class actions in the United States are your second risk.
These are the live US private theories with money already moving. The Telephone Consumer Protection Act, or TCPA, covers AI voice and SMS. California's Invasion of Privacy Act, the CIPA, charges $2,500 per violation under §638.51 on website visitor tracking. That statute carries two 2026 customer settlements at $3.85M and $21.5M, and more than 800 CIPA claims were filed in 2025 alone. Private plaintiffs bring these, which is why they fire far more often than regulators do.
State privacy and chatbot statutes are your third risk.
Roughly fifteen states have passed chatbot-disclosure statutes, and the state comprehensive privacy laws add data-subject rights that reach the prospect records you hold. These rarely produce a headline penalty at this size but routinely produce obligations you must satisfy on request, a different kind of exposure. The work they demand is record-keeping you should do regardless. Know where each contact came from, honor deletion and opt-out requests across every identifier, and disclose AI use in the first message.
A headline privacy fine is your last risk.
There is essentially no record of a regulator fining a small cold-email operator over its list. Enforcement lands upstream on the data brokers and enrichment vendors, where it is real and rising. Your exposure runs through the vendors you buy from rather than your own sending, which argues for knowing a provider's sourcing before its price.
10.5.2 Do the four compliance chores that cost you almost nothing.
First, record where every contact came from, because the question you will be asked is where you got the address. Second, put a real physical postal address and a working opt-out in every message, which CAN-SPAM requires and most senders get casually wrong. Third, honor an opt-out within ten business days and across every identifier you hold for that person, not just the address that asked. Fourth, disclose AI use in the first message, which satisfies the roughly fifteen US state chatbot laws with one line. All four cost almost nothing, and each is checkable by an agent before a send.
10.5.3 Suppress every identifier you hold for a person, not just the address.
About 20% of contacts change jobs each year and nobody can follow a person through those moves, so do not try. The law only asks that the identifiers you already hold are honored everywhere, so when somebody opts out suppress the name, the company, every email address on file, and the profile URL at once. Model the person separately from the job in one shared suppression store, since a per-tool list only honors the string that tool knows. The regimes disagree: under CAN-SPAM the opt-out attaches to the address, while GDPR Article 21(2), outside this guide's US scope but the strictest reading available, makes it an absolute right of the person. Chapter VIII prices the failure at the $2.95M civil penalty collected from Verkada, against a $53,088 statutory maximum per email, in an action that also covered a missing opt-out mechanism and a missing postal address. Against all of it, the fix is one day of schema work.
10.5.4 Send every review respondent to Google, or send none.
Routing four-and-five-star respondents to Google while quietly filtering out everyone below is taught in this field as ordinary automation, and it violates Google's review policies. It also sits inside the US Federal Trade Commission's rule on consumer reviews and testimonials, which reaches practices that suppress negative reviews. Treat the whole pattern as an exposure: send every respondent or none.
10.5.5 Say plainly where the evidence is thin.
Across 65 commissioned research reports and 209 practitioner transcripts, there is no documented case of a named AI sales-development agent being prompt-injected into an unauthorized commitment. And no first-person post-mortem of an injected production outbound agent exists. Everything here about injection rests on mechanism, not a recorded incident, and pretending otherwise would mean inventing evidence.
10.5.6 Build for misreading, the failure you can prove.
Every documented failure in this material is an agent misreading something. SaaStr's agenda agent pulled the first 50 records from a paginated API, silently dropped the rest when the list grew past 50, then invented an explanation when questioned. Other agents sent cold emails to their own company's existing customers, and one put the same person into two sequences at once. None of it required an attacker, so build your defenses against misreading first.
10.5.7 Keep the reading agent away from send authority.
Two published results keep the injection risk live even without a public post-mortem. ForcedLeak, a September 2025 exploit of Salesforce's Agentforce rated 9.4 out of 10 on the industry severity scale, hid instructions in a web lead form's description field. It exfiltrated CRM data to an expired allow-listed domain the researchers re-registered for $5. MCPTox, an August 2025 benchmark, found tool-description poisoning succeeded up to 72.8% of the time across 45 real Model Context Protocol servers. So keep the agent that reads inbound replies separate from the agent that can send: least privilege applied to the one split that matters most.
10.6 Count only the safeguards you can point at.
Two moves follow from all of this: the sending credential goes behind code rather than a prompt, and the mailbox providers get feared before the regulators. Counting the human gate you already have as a control, without asking who put it there or what remains once it goes, is the standing mistake. So is spending the first year defending against a headline privacy fine while the mailbox providers regulate the channel.
Measure four numbers, and write down in advance which ones you will ignore.
Four numbers survive the way email actually gets delivered: reply, positive reply, meeting held, and pipeline created. Everything else goes on a written ignore list, and the agent that reports the four needs tighter rules than the agent that sends. The reason is that a bad send costs one prospect while a confident wrong number travels further and lives longer.
11.1 An unwritten ignore list puts open rate back into the quarterly review.
A number you glance at but never record still steers your decisions while leaving nothing to audit. Name in writing the four numbers you will act on and the four you will not, because a list in your head loses to a dashboard.
11.1.1 Turn open and click tracking off in week one.
Apple Mail Privacy Protection loads images at delivery, so an open gets recorded whether or not a human ever looked. The newsletter platform beehiiv put Apple at 49.29% of tracked opens in its 2026 analysis; the email-analytics firm Litmus put it near 58% in early-2026 data. The email platform Omeda watched a natural experiment across roughly 80,000 deployments and two billion emails in 2022, where unique opens moved from 15.2% to 29.0% with no change in audience. Clicks are contaminated too, because corporate security scanners open links from datacenter IPs within seconds. Forum reports from March 2026, with no controlled measurement behind them, blame scanners for up to 90% of reported clicks in some B2B campaigns. Treat any click in the first ten to sixty seconds as a machine.
11.1.2 Four numbers earn a decision, and four go on the ignore list.
Track reply, positive reply, meeting held, and pipeline created, because a human decision sits behind each one and no scanner or pre-loaded image can fake it. Read reply rate as a diagnostic and steer on positive reply. Opens, clicks, emails sent, and sequence steps completed go on the ignore list; the last two measure how hard you worked, not whether a buyer cared.
11.1.3 Read every reply and sort it by hand.
At 250 to 300 researched sends a month you get roughly 8 to 18 replies. That is too few to test anything and exactly the right size to sort by hand. File each under a fixed objection type: wrong person, no budget, no pain, bad timing, already have a solution, and "what is this?". The sort tells you what to change next, which the statistics at this volume never will.
11.1.4 One randomized holdout a year is all you can afford.
A holdout is a group of accounts you deliberately send nothing to, so you have a comparison. A real incrementality test compares whole markets that got your campaign against markets that got none. The B2B geo-lift practitioners Amsive and Zaitz price its entry at 20 or more markets per group, four to eight weeks of runtime, and upwards of $10,000 a month in channel spend. A few hundred touches reach none of that, so run one randomized holdout a year, splitting your ideal-customer accounts at random into a group you work and a group nobody touches. Random is the whole point, because a hand-picked holdout measures your picking, and one a year means spending it on your biggest open question, not a subject line.
11.2 Lock down the reporting agent harder than the sender.
The sender reaches one prospect at a time, while the reporting agent reaches the plan, the board slide and next quarter's budget. So the guardrails go here first: the narrowest permissions that work, a required range on every rate, and a list of banned verbs.
11.2.1 Set four guardrails before the agent states a rate.
First, a floor: no stated rate until at least 100 sends and at least 10 positive events sit behind it. Below that the only allowed output is "insufficient data to conclude, n too small", where n is the number of sends. Second, a Wilson interval on every comparison, reporting each rate as a range plus a sentence on whether the two ranges overlap. Third, a ban on causal verbs such as drove, caused and improved. Fourth, printed SQL and row counts, so a human can check the sends behind the rate instead of trusting the prose.
11.2.2 Never merge two campaigns to reach a bigger sample.
When both arms of a test are too small to support a claim, the tempting fix is to merge results. You merge across operators or campaigns until the count looks respectable. Forbid it, because side-by-side campaigns use different lists and offers, so the merged number describes a group of people that never existed and manufactures significance.
11.2.3 Ignore the winner your sending tool picks.
The cold-email platform Instantly says on its own blog that its dashboard shows no statistical significance indicators and no p-values, the standard test of whether a gap could be chance. The sales-intelligence platform Apollo's help centre claims significance "should start to appear around 200 recipients (100/variant)". At 250 sends, a 5% reply rate carries a 95% Wilson interval of 2.9% to 8.5%, indistinguishable from 3% and from 8%. And 5% is the top of the researched band, the reply-rate band hand-researched email earns, not the rate at volume. Detecting a 50% lift on a 5% baseline needs roughly 1,469 sends per version, which Chapter I prices. So Apollo's 100 falls short by roughly fifteen times, and by sevenfold against the lower competing estimate of about 700, so ignore the crown your tool puts on a variant from day one.
11.2.4 Never let the agent rewrite its own prompts.
No credible published build has safely fed reply outcomes back into the prompts that write the emails. And the one documented build that closed the loop ran no significance test anywhere, which automated the failure. At this volume an agent editing its own instructions from what "worked" is overfitting, learning the noise instead of the pattern, with nobody watching. Have it propose a playbook change as a diff instead, and keep a human reading that diff before it lands in version control.
11.2.5 A confident forecast is the most dangerous output.
The SaaS conference company SaaStr runs a stack of agents and vibe-coded apps in production, counted by its own chief AI officer in February 2026 at twenty, or almost thirty that morning. One agent forecast about 1,000 extra ticket sales with confidence and data behind it. Twenty-four hours after the campaign ran it had delivered 2, off by roughly 500 times, while drawing a wave of investors applying for free passes. Neither operator could catch the forecast before it went out, and their fix, dialing the agent's goal aggressiveness down in the prompt, left the number unchecked. A softer prompt is not a control.
11.3 Make "not enough data to conclude" an expected and counted answer.
Almost everyone publishing in this field is selling something, so the public record is nearly all wins and holds essentially no rigorous low-volume experiments. An agent trained on that record will manufacture a finding rather than admit the data is thin. So put the phrase in its permitted vocabulary, count how often it comes back, and treat a month that returns it as a month that reported correctly.
11.3.1 Put the refusal phrase in the agent's vocabulary or it invents a trend.
A model with no honest way to say nothing will invent something, because generating text is the only thing it does. Write "insufficient data to conclude, n too small" into the allowed vocabulary explicitly, and expect to see it often. Treat it as a correct answer rather than a failed run.
11.3.2 "We measured nothing" is not "we measured no effect."
"We measured no effect" claims you looked with enough sends to see one and found none, a null result. "We measured nothing" says the sample was never large enough to look. At these volumes the second is almost always the true one, and collapsing the two is how a program talks itself into killing a tactic that was working.
11.3.3 Write down what you killed, and why, on a schedule.
The most credible thing across 65 research reports and 209 recorded practitioner videos is a dated admission: "we stopped doing X because Y". It runs against the speaker's commercial interest, and almost nobody writes one about their own program. Put a recurring date on the calendar to record what you killed and why.
11.4 A wrong number does more damage than a bad send.
A bad send reaches one prospect, and the blast radius, how far a single mistake travels, stops there. A wrong number reaches the plan, the board slide and the budget, and it keeps getting quoted long after the campaign that produced it is dead. So the tighter rules go on the agent writing the weekly number before the agent writing the email. Get that order backwards and you control the sending credential while the report goes unchecked. That is how you act on a forecast that promised about 1,000 extra tickets and delivered 2.
Almost everything is decided in week one.
Three things run on a clock money cannot reset, and week one starts all three: your domains, your byline and your suppression store, the record of who must never be contacted again. Everything else that week is cheap to decide now and expensive to retrofit, including where the sending credential lives. After week six the job changes shape, and what you hold is three decision dates and one hard stop on sending.
Put all eleven moves on one calendar.
Five of the eleven moves start in week one. Sort every tactic by what survives everyone copying it (1), start publishing under a real name (4), and own the memory and the judgement while renting the execution (6). Put the sending credential behind code (7), and start the clock at domain purchase rather than at first send (9). Weeks two through six add two more: pick a list nobody else can build (2) and test your offer rather than your wording (3). The last four never end: write the message yourself and count the AI steps between your judgement and the sent text (5), and count meetings held rather than replies received (8). Count the agents that do real work every day, per person (10), and fear the mailbox providers before the regulators, because that is the order the risks actually bite in (11).
12.1 Week one holds every decision a clock will not let you buy back.
Ten decisions land in week one, and most could wait until month three at the same cost. Three cannot: a domain has a fixed warming period, and an audience only compounds from the day it starts. A suppression store built after the first send has already lost the record it exists to keep.
12.1.1 Buy and warm your domains on day one.
Buy the domain on day one whichever build you are running, because domain age is the one input money cannot reset. A fleet, many domains sending at once, then warms on the clock Chapter III sets before real volume: two to four weeks on a domain you already owned, four to eight on a new one. One or two mailboxes wait only the one to three days DNS takes. Then they send real researched mail at five to ten a day ramping toward the documented cap, so warming and working are the same activity. Agents compress your research and drafting time, and they compress neither wait at all.
12.1.2 Publish under a named human byline on day one.
An audience is the only asset here you cannot buy back later, so the start date is the whole decision. Start publishing in October 2026 and you have a year-old audience in October 2027. Start in October 2027 and you are twelve months behind, and no spending closes that gap.
12.1.3 Build suppression, audit and state stores first.
Chapter VIII counts thirteen of fourteen recorded builds that shipped only the acquisition half, find, enrich, compose and send, with nothing to record what comes back. Build the missing half in week one. The suppression store holds who must never be contacted again, and the audit log holds a record of what went out that nobody can edit. The state store holds what has already happened to each contact. All three exist before you have anything to send.
12.1.4 Put the send credential in code the model cannot reach.
A gate here is a human in the loop, a person who approves a message before it sends. Week one is when you decide where that approval lives, because Chapter X found deliberately chosen gates rare in production and one decayed into a click on camera. Never count on a gate you did not put there yourself. Hand the sending credential to plain deterministic code, so the approval sits inside the send path rather than in a prompt the model can read past. Set Anthropic's disable-model-invocation: true flag on any skill that can send, documented for exactly this case ("higher risk things like a skill that sends a message"), so only a human can fire the send.
12.1.5 Add the "how did you hear about us" field on day one.
The field tells you nothing in month one, which is exactly why it never gets added. By month nine, when inbound starts arriving, it is the only thing separating work that paid off from luck, and it cannot be filled in after the fact.
12.1.6 Turn open tracking off in week one.
Apple Mail Privacy Protection records opens no human made, which Chapter XI documents, so the number steers decisions while leaving nothing to audit. In week one, write down the four measures you will act on: reply, positive reply, meeting held, and pipeline created. Write down the four you will ignore too: opens, clicks, emails sent, and sequence steps completed. Count a meeting as booked on a confirmed calendar event and as held only when it actually took place, never either one from a reply. The day-sixty and day-ninety decision dates below have nothing to read unless that counter exists before your first send, and a published show rate is somebody else's.
12.1.7 Sort every tactic on two axes before you buy anything.
Two things decide whether a tactic lasts: whether it can be mass-produced, and who owns the ground it stands on. Score everything on both axes before you spend a dollar, then check the plan against the sixteen durable items, the tactics this research found survive both axes, which takes ninety seconds. Scoring effort instead sends you wrong in the most expensive direction. The four heaviest builds in this research sit exactly where long build times meet short-lived tactics, all inside the eight that did not survive. Most of them were disowned in public by their builders ("I wish we hadn't").
12.1.8 Buy the stack, and choose the sequencer on what its connector can write.
Buy a contact data source, an email verifier, a sequencer, and mailboxes, and build nothing you could rent. Pick the sequencer on what its MCP server, the standard interface your agent calls the tool through, can actually write. Some cannot manage a campaign at all, while others create and push whole sequences. What each can read, write and authenticate against shifts month to month. So check the write boundary yourself before wiring a live send path, and treat deliverability marketing as the weaker criterion.
12.1.9 Rank the risks, and keep the list inside the country you sell in.
Write the risk order down as the evidence supports it, not as it alarms you. Inbox providers come first, because getting blocked kills the channel outright with no appeal and no lawyer. US private-plaintiff theories come second, since money is already moving through them, and state disclosure and privacy statutes third. A headline regulatory fine comes last, because there is essentially no record of one landing on an operator this size. Then one filter before you buy a tool: the build sells inside the United States, so every list is scoped to US recipients and anything outside leaves it.
12.1.10 Run one suppression store for the whole operation.
Every agent across every operator reads and writes one shared suppression store, one audit log and one state store. If each operator keeps a private do-not-contact list instead, the same person gets emailed twice by two of your agents.
12.2 Weeks two to six go to the list and the offers.
Which list you pick moves your results by roughly 100x, which Chapter IV documents, and which offer you pick moves them by two to three times. Every other variable moves them by less than a few hundred sends a month can detect. So spend the whole build window on the only two levers whose effects you will see.
12.2.1 Define one segment, then pull 150 to 200 contacts.
A usable segment is two to five rules of thumb pointing at a specific tension someone feels, which a headcount-and-industry filter cannot express. Pull 150 to 200 contacts against it, because nothing gets killed on fewer than 150 to 200 touches or fewer than four to six weeks. A smaller pull cannot produce a decision either way.
12.2.2 At day seven, read fifty rows by hand.
Spend thirty minutes reading your own list and marking who does not belong. That directly measures list quality, the one variable whose spread runs to roughly 100x, and nothing else in month one is that cheap or that informative.
12.2.3 Email three conference organizers in week two.
Conference attendee lists carry the highest cold-email reply figures anywhere in this research, 15 to 20%. They reach people whose contact details are not published and cannot be filtered out of a database. The figure is Nick Saraev's self-report on his own narrow campaigns, checked by nothing independent, and he sells scraping and automation education, so weigh it accordingly. Nobody has measured the cost per ask or the share of organizers who say yes, and Chapter VII shows the same man blaming copy for most campaign failure, the reverse of the ranking this move rests on. The ask still costs one email, and exclusivity is the only list property with high reply rates attached to it at all.
12.2.4 Build five to ten offers before rewriting any email.
Structurally different means a different promise, a different countable unit and a different risk reversal, which a reworded subject line is not. Give each offer its own list so results stay separable, and finish before spending a minute on message variants. At your volume you can tell offers apart while variants stay indistinguishable.
12.2.5 At day thirty, watch the wrong-person rate.
Sort your first twelve to twenty-four replies into fixed buckets: wrong person, no budget, no pain, bad timing, already have a solution, and what is this. The sample proves nothing statistically and still says a great deal about positioning. Watch the wrong-person rate especially, because it separates a list problem from an offer problem, and those have different fixes.
12.2.6 First email lands in week one on a small build, and week five to nine on a fleet.
On one or two mailboxes, DNS propagation and a list you have read are the only gates, so the first researched email goes out in week one. On a fleet, many domains sending at volume, domain warming and list building run in parallel, and whichever finishes second sets your start date. Both shortcuts, an unwarmed fleet domain or an unread list, cost more than the weeks they save.
12.3 Plan each week around what one person can actually check.
One operator's week has a fixed shape, and every agent you add takes a bite out of it every week rather than once. When you need more output, the move that delivers it is a second operator carrying their own agents, their own list and their own offers.
12.3.1 Send 250 to 300 a month, and hold there.
That is roughly a dozen sends on a working day across two or three mailboxes. It sits well under the 20 to 30 a day per warmed mailbox deliverability consensus permits. And it sits under the 40 to 90 a day per domain Chapter XIII derives. The number is your choice about depth rather than an infrastructure limit, and agents make exceeding it trivial. So hold the line with a cap in the sending code that does not depend on your discipline.
12.3.2 Each operator's day is already full.
Outreach takes forty-five to sixty minutes a day, content thirty to sixty minutes a day, and maintaining the agents fifteen to twenty hours a week. That last figure comes from Jason Lemkin and Amelia Lerutte at SaaStr, each of them rather than the two combined. It covers the twenty to thirty agents and vibe-coded apps they ran in production in February 2026. They report management time did not fall but shifted, a concession against the interest of an operation selling the agent story. A large marketing organization gets more copies of this day rather than a bigger one, an inference nobody has measured at that scale.
12.3.3 Five to ten verified outputs a day is the ceiling.
That is how many AI-produced items one human can read, check and act on in a day, the ceiling Chapter IX budgets. A plan that needs twenty gets twenty approvals with the reading skipped, which leaves you with nothing actually checked while you believe it is.
12.3.4 Add at most one new agent per operator per month.
An agent here is one configured worker with its own instructions and tools doing one recurring job, a researcher, a drafter, a reply sorter, and one campaign uses several. Add an agent only for a step the operator has already done by hand and stopped changing, because the manual version is the specification. Each new agent costs roughly two weeks of onboarding, during which the existing agents get worse because the attention that maintained them moved. Four to six agents doing real work every day is capacity, which Chapter IX budgets as direct reports rather than installed software. So past that point adding one means retiring one.
12.3.5 Cap the AI rewrite chain at one step.
A hop is one step where a model writes text that another model then rewrites. Past the first hop the output is slop by construction, however good the model. Allow one, and recount every time you add or change an agent, because hops pile up quietly as the estate grows. Two separate vendor benchmarks rank AI-drafted, human-finished mail above pure human and pure AI writing. So let the model fill variable slots against a locked style guide, and keep the body, the ask and the final read yourself.
12.3.6 Answer the first positive reply yourself, same day.
Propose two specific times and do the scheduling yourself, because a bare booking link spends the advantage you just earned. About a third of cold-booked meetings never happen. That is the gap between the 2.25 booked and 1.5 held that Chapter III derives, and the booking is the part you control. Book inside three days and send three reminders with something to read beforehand.
12.3.7 Model spend is the smallest line on the budget.
SaaStr's two production agents, run by Jason Lemkin and Amelia Lerutte, cost $257 a month combined in model inference. That is roughly an eighth of the sending infrastructure line and a hundred times less than the headcount they displaced. Lerutte's first reaction was "Is that in one day?", while Lemkin puts the fully burdened figure at $500 to $800 a month. That counts a $30 identity service and a share of the Salesforce bill, two to three times the quoted number and the one to plan against. Optimizing either figure optimizes a rounding error while enrichment credits, deliverability infrastructure, and operator review time go unmanaged.
12.3.8 Re-price build versus buy every quarter.
Chapter II prices any hand-built workaround for a model weakness as a wasting asset, because one model release can delete the lot. The agent-facing connectors change monthly in what they can read, write and authenticate against, so a January build-versus-buy decision describes a stack that no longer exists in April. Review it quarterly.
12.4 Set three decision dates and one kill switch before your first send.
Week-six panic and month-twelve nursing are the same failure in different clothes, and a date committed in advance is the only defence against both. At each of the three dates, write a diagnosis of your list, your offer and your deliverability in words, not as a percentage held against somebody else's benchmark. The kill switch needs no diagnosis, because it reads a number and stops the sending on its own.
12.4.1 Hard-wire a kill switch that stops sending before the mailbox providers do.
A kill switch is a rule inside the sending code that halts every send the moment a number crosses a line, with nobody deciding anything in the moment. Set yours at the 2 to 3% bounce and 0.3% complaint thresholds Chapter III applies, and let it override every other rule. Above those lines the mailbox providers are already re-rating your domain, and sending on costs you the domain rather than the campaign. Test it with a dry run, a full pass with sending switched off.
12.4.2 At day sixty, zero meetings is a verdict on the list.
The working figure is 2.25 meetings booked a month and about 1.5 held, which Chapter III derives from vendor-band inputs. So sixty days buys about four and a half booked and roughly three held. Zero qualified meetings at that point points at who you targeted, not how you worded it. Rewriting the email on this date is the standard mistake, and it costs another sixty days.
12.4.3 At day ninety, write a diagnosis instead of comparing a rate.
On the same arithmetic, ninety days at 250 to 300 sends a month buys about 6.75 meetings booked and roughly 4.5 held. That is far too few to compare against a published benchmark without fooling yourself. The honest day-90 output is a written judgement on three things: was the list right, was the offer right, and did the mail arrive?
12.4.4 At month nine, judge the topic and the audience.
Content cannot be judged at all before month nine. When you do judge it, only two things are fixable, what you write about and who you write it for. Leave the cadence alone, because killing the publishing habit gives up the one asset you cannot restart later at the same price.
12.4.5 A dead variant is not a dead campaign.
A sequence under 1% reply after 200 researched sends, judged against the 3-to-6% band, is a dead variant. So re-angle the offer or the list rather than concluding outbound does not work. Cancel third-party intent data, bought signals that a company is shopping, at month two if it has not beaten your cold baseline for two months running. Exit marketplace bidding at month six if your effective hourly rate has not passed the platform floor, and stop paying to amplify your own content at month nine. Each trigger judges one component, which is why they sit outside the three decision dates.
12.5 Seven questions here have no answer yet.
Everything above rests on evidence that was graded, dated and traced to a named source, and these seven questions have none of that. Rereading the research does not resolve them, any confident answer would be invented, and at least two of them could reverse the recommendations above.
12.5.1 Low volume may be protecting you, or may just be hiding you.
The bet is that a few hundred researched sends a month land better than a hundred thousand generic ones. GlockApps, an inbox-placement testing service, supported it in Q1 2025, when senders above a million a month fell below 28% placement while moderate-volume senders rose. Three quarters later the same test reversed. Senders at 1,000 to 10,000 a month lost 15% of their Gmail placement and 19% on Workspace, while senders above a million gained 20% on Gmail. Same vendor, same seed-test methodology, opposite directions, and two more quarters of falling low-volume placement would settle the question against this entire design.
12.5.2 Nobody has isolated what publishing is actually worth.
Every available warm-versus-cold comparison sets people who chose to engage with your content against people who did not. And the engaged were already more interested before they read a word. The comparison measures self-selection and content lift together and cannot pull them apart. So the case for publishing on day one rests on which mistake you would regret more, not on a measured effect.
12.5.3 Sender brand ranks third on testimony, not on evidence.
The strongest support comes from Sam Blond, chief executive of the message-automation company Monaco, who named sender brand as a driver of reply rate separate from the message. His words: "there's only so much we can do with sequence structure and message structure". The claim costs him sales, which makes it the best-incentivized claim in the leverage ranking. But it is still one person's assertion, because nobody has held list and message constant while varying brand.
12.5.4 No independent study shows that bought intent data predicts a purchase.
Across all 65 commissioned reports and 209 practitioner transcripts, not one independent study shows that vendor-supplied buying signals predict a B2B purchase. That is an absence of evidence rather than evidence of absence, though the absence is total. It sits in a field where a survey of 750 B2B marketers found 91% using intent data, 87% calling their own signals unreliable, and 24% reporting exceptional returns.
12.5.5 The security design rests on mechanism, not on incidents.
No first-person post-mortem exists of a production outbound agent prompt-injected into an unauthorized commitment, though the mechanism is well demonstrated. ForcedLeak, a Salesforce Agentforce exploit rated 9.4 out of 10 on the standard severity scale, pulled CRM data out through a lead-form description field. And the published MCPTox benchmark recorded tool-poisoning success up to 72.8% across 45 real MCP servers. These defences are designed against the mechanism, not against a real case.
12.5.6 Watch whether the send gate gets encoded or quietly deleted.
The most consistently confirmed finding in this research is that a human approves before an outbound message goes out. Two futures are live: the gate becomes a shipped, default-on primitive, or vendors ship autonomy and it disappears unnoticed. If it disappears, the finding everything above leans on hardest has reversed, and several chapters go with it.
12.5.7 Re-run these numbers when the next model generation ships.
Chapter II tells what happened when a single model release deleted three workarounds a vendor's own engineers had built, and treats that engineering as temporary by design. Anything here describing what a model can do has a shelf life measured in model releases, while anything describing what a human can verify in a day holds.
Implementation
The same argument turned into a system you can buy, wire together and run. Nine chapters that name every component slot, the controls that keep an agent from sending on its own, what the data and sending layers cost, where suppression and state actually live, what connects to Claude and what can fire a real message, what the whole thing costs to operate, and three priced builds you can copy. It closes with thirty-seven actions in the order to do them. One chapter here is deliberately perishable and says so; the rest are written to survive the vendors named in it.
Learn the slots before you learn the logos.
Every build in this field assembles the same eleven components, and the products filling them turn over far faster than the pattern does. Each slot below is named by its job, with two lines underneath: what you own versus rent, and whose name is on each account. Each job runs in one of three places: your monthly seat, a metered account billed per token, or ordinary code you wrote. The named products as of August 2026 sit in Chapter XXI, dated on purpose. A vendor dying or doubling its price rewrites that chapter and leaves this one standing.
13.1 Eleven slots do the whole job, and the first one is a person.
A slot is a job the system has to do, whether you buy it, rent it or write it. Naming the slots separately stops you paying two vendors for one job, or discovering in month three that nobody owns one. Most are replaceable commodities; the three you build and keep are the operator, the suppression-and-state store, and the agent layer that drives the rest.
13.1.1 The operator is a slot, not the part you are trying to remove.
A person working in Claude Code and Cowork sits in the middle of this system as a named component, not an admission of failure. The operator is the send gate, the kill switch, and the collector wherever a platform permits reading but forbids scraping. The seat licence is priced around that person too, since a human inside an Anthropic application is the condition Chapter XX's clause turns on. Design the person out and you buy every one of those jobs back in software, at the capacity Chapter IX budgets.
13.1.2 Contact data is a rented input, so buy it as a single tool.
This slot answers who exists and how to reach them. The mature vendors all publish an official MCP server, the standard interface an agent calls a vendor's tool through. The August 2026 census found nine: Apollo, ZoomInfo, Lusha, Crunchbase, Surfe, Wiza, FullEnrich, Common Room and Warmly. As data tools they read and enrich only, except Apollo, whose server also carries sequence enrolment and one-off send, which Chapter XIX deals with. Two naming traps here, Clay and the wrong Apollo, are laid out in Chapter XV.
13.1.3 Verification exists so that one false positive cannot cost you a domain.
Verification checks that an address will accept mail before you send. The expensive error is the false positive, the address marked valid that bounces anyway. Re-verification of Apollo exports marked "verified" found a small fraction genuinely valid and a majority catch-all, a domain that accepts mail to any address and confirms nothing. One bad batch on a fresh domain crosses the ceiling Chapter XII's kill switch watches and costs months of reputation. So pay for this slot separately from the data feeding it.
13.1.4 Signal sources are a slot, and most builds leave it empty.
A signal source tells you which contact to work and when. Chapter VIII's twelve-layer model carries signal as its own layer, and most recorded builds leave the slot empty. That produces a system that researches and sends but never notices anything. Fill it with observed public statements rather than a purchased score, for the reasons Chapter XVI works through. The compliant implementation is mostly a person reading, and the cheapest version costs nothing but the reading.
13.1.5 Enrichment belongs behind one waterfall call rather than a chain you maintain.
Enrichment turns a name and a company into an address, a title and context worth writing from. Buy a waterfall, a chain of providers tried in order until one answers, as a single product and call it as one deterministic tool. Practitioners put the break-even for chaining providers yourself at 500 to 1,000 contacts a month, which Chapter XV works through. And one call gives your agent one response envelope, one set of rate limits and one way to fail.
13.1.6 At low volume, mailboxes, domains and warmup are a single slot.
Reputation is measured at the domain level, so a second address on your main domain buys no isolation. Use a lookalike secondary domain that redirects to your main site instead. Per-mailbox limits run to 15 cold sends a day from Maildoso, while practitioner consensus sits at 20 to 30. Two or three mailboxes per domain puts one domain around 40 to 90 a day, with no source documenting a hard per-domain ceiling. Skip synthetic warmup at this size, because your real mail does the warming on the clock Chapter III sets. Skip dedicated infrastructure too, because shared Google or Microsoft infrastructure cannot be listed by your IP address.
13.1.7 Choose the sequencer on what its connector can actually write.
The sequencer holds your campaigns, schedules the steps and records what went out. It is the one rented slot where the agent-facing connector matters more than the feature list. Some vendors cannot manage a campaign through their connector while others create and push whole sequences, and the boundary shifts month to month. Check what the connector can write yourself before wiring a live path through it. This is the slot that can fire a real message, which Chapter XIX deals with.
13.1.8 Reply handling is the half that thirteen builds in fourteen skipped.
This slot sorts each reply into a category, routes it, and escalates the hard case to a person. The hard case is someone asking to be removed in their own words rather than by clicking unsubscribe. Chapter VIII counts thirteen of fourteen recorded builds shipping only the half that sends, and the half that listens is where every number worth measuring gets created. Give your reviewer a payload, the category, the confidence and the written reasoning, as one production triage build does.
13.1.9 Suppression and state need three layers, because no sender can suppress a person.
Every sending platform blocks on an identifier, an address or a domain, and none follows a human through a job change. So a suppressed person with a new work address becomes contactable again. The fix is the three-layer suppression design Chapter XVIII builds; the slot decision here is the store. Attio carries the right identity model, multi-value addresses on its People object, but no native do-not-contact attribute, so your agent checks a custom flag you build. HubSpot ships a native unsubscribe property it enforces itself, at the cost of a heavier data model and a slower search interface.
13.1.10 Meeting held cannot be proved from a CRM, so read the conferencing layer.
Every CRM's meeting-outcome field is manual by default, which HubSpot's own staff confirm, so a held-meeting count from your CRM counts what somebody remembered to tick. Positive proof exists one layer down: Microsoft Graph attendance reports, Zoom participant reports, or a recording that only exists if people turned up. Calendly and Cal.com flag no-shows instead, which infers attendance rather than proving it. Chapter XI makes meetings held one of the four numbers you act on, and this slot makes that number real instead of remembered.
13.1.11 The agent layer is four to six workers per operator, and it drives the rest.
The agent layer is the part you actually build. Chapter IX budgets it at four to six agents doing real work per operator. Its job is to drive the machine slots through their connectors, read and write your own stores, and stop short of the send. For an operator whose work is documents and research, Cowork is the better home as of August 2026, on paid plans only. The reasons are the ones Chapter XIV weighs against the desktop app.
13.2 Own the memory and the judgement, rent the execution, own the accounts under both.
The own-versus-rent line decides what survives a vendor change, and under it sits whose name is on each account. Rent the parts that do work, because they are commodities and competing on them is somebody else's business. Own the parts that accumulate, because memory and judgement are the only two components that get better with age. Then check that the accounts belong to the organisation the system serves, which turns out to be a licensing requirement rather than a preference.
13.2.1 Your memory is a set of versioned files, and it is also your cheapest cost control.
Memory means the files that define the system: the customer profile, the offer library, the voice guide, the output schema and the worked examples. They are identical for every contact you research, which makes them exactly what to hold in the model's cache. There a cached read costs a tenth of base input, and a five-minute cache write pays for itself after one read.
13.2.2 Own your procedures as skills, and rent vendor capability as connectors.
A skill, a folder with a SKILL.md file inside it, is where your own repeatable procedures belong. Chapter XIV covers the format, the loading model and the control that keeps a sending skill manual-only. Vendor capability arrives the other way, through MCP servers and connectors, and the go-to-market presence in Anthropic's official marketplace is thin. Chapter XIV counts it, and Lemlist publishing skills on its own site tells you where to look when the marketplace comes up empty.
13.2.3 Rent execution, and rent only what you could replace inside a week.
Execution means work that carries no memory: fetching a record, verifying an address, delivering a message, holding a campaign schedule. Each has several suppliers, and your own version would be worse and need maintaining forever. The test for a rented part is whether you could swap it inside a week. That requires that your data leaves in a format you can read and that the account carries your name.
13.2.4 Own the account every rented part lives in, because the terms require it.
The rule fits one line: whoever the system serves owns the accounts, the seats and the component licences, so firing the consultant leaves the system standing. Anthropic's terms make that a requirement rather than good practice. A subscription seat may produce deliverables but may not power somebody else's running system, which Chapter XX works through. Apply the test one slot at a time: a suppression store inside a vendor account you do not own is half owned. And rented lists get stranded the day a subscription lapses.
13.2.5 One infrastructure vendor keeps your domains, which is the slot failing this test.
Instantly's help centre states that its done-for-you accounts are configured exclusively for use inside Instantly, and that it retains domain ownership and administrator access and cannot transfer either. The August 2026 research on integrated suites calls that unusual, since most infrastructure providers register domains in the buyer's name. Domain age is the one input money cannot reset, which fails the own-your-accounts test outright, and makes this the worst slot to hold on somebody else's paperwork.
13.3 Decide where each job runs: your seat, the metered API, or your own code.
Every job runs in one of three places, each billed and governed differently. Work you sit and watch inside one of Anthropic's own apps runs on your monthly seat at no extra charge. Work that runs on a schedule or that your code starts runs on the metered API, billed per token. Sending the actual email runs in ordinary code you wrote and touches neither. Anthropic's terms of use decide the choice, not the price, so an unattended job belongs on the API even when the seat would be cheaper.
13.3.1 One clause in the terms decides this.
Consumer Terms section 3(7) bars accessing the services through automated or non-human means, with an Anthropic API key the standing exception; Chapter XX takes the clause apart. The working test for the seat-versus-API choice is whether a human sits in the loop inside an Anthropic application. The moment work becomes unattended, scheduled outside Cowork, triggered by code, serving a third party, or irreversible, it belongs on the metered API. And since Anthropic publishes no safe-harbour list, the edge cases are yours to judge.
13.3.2 A seat covers interactive work, including research you sell to a client.
A seat permits Research mode by hand and commercial use of what it produces. That is because Consumer Terms section 4 assigns output ownership to the user and conditions it on nothing. A consultant producing a client deliverable interactively on their own seat sits inside that permission, though it rests on silence plus output ownership rather than an explicit carve-out. The one seat-native exception to the no-scheduling rule is Cowork Scheduled Tasks, which Chapter XIV covers.
13.3.3 A seat cannot power anything that runs while you are absent.
Scripted access sits outside what a seat allows, along with a client's running system fed by your credentials. So does routing anybody else's requests through your account, sharing it, or reselling it. Enforcement is documented rather than theoretical: Anthropic added server-side checks in early 2026 that block third-party harnesses from using subscription credentials. It formalised the position on 4 April 2026, telling The Register that using Claude subscriptions with third-party tools "isn't permitted under our Terms of Service, and they put an outsized strain on our systems".
13.3.4 The cheap option is real, and the shortcut into it breaches the terms.
Research mode inside a subscription carries no marginal per-token cost, only a draw on the flat usage window. The August 2026 economics research prices it one to two orders of magnitude below an equivalent API multi-agent loop. A 200-source deep research run costs roughly $20.60 on Opus 5 or $8.60 on Sonnet 5 before caching. That gap tempts operators into automating the web interface, which section 3(7) prohibits, leaving by hand, Cowork's scheduler, or the API as the routes. With no endpoint, connector or SDK equivalent, structured output leaves through the copy button, so any pipeline carries a human step by design.
13.3.5 Client-facing work belongs on a commercial account.
Only the Commercial Terms carry Anthropic's intellectual property indemnity and its commitment not to train on your content. A consumer seat runs the other way: you indemnify Anthropic, and liability is capped at "the greater of six months' fees and $100". Team, Enterprise and the API sit under Commercial Terms, and the August 2026 tiers research gives a rule short enough to act on. Move to a commercial tier the moment you touch client or prospect data.
13.3.6 Dollar caps stop spend, while alerts and turn caps only describe it.
Four mechanisms actually stop money leaving; the first two are prepaid credits with auto-reload off and the Console monthly spend limit returning HTTP 429. The other two are the Agent SDK's max_budget_usd and per-request max_tokens; budget alerts and usage reports arrive after the money is gone. max_turns is no cap either, because Anthropic's own issue tracker records a subagent declared at ten turns running seventy-two to seventy-five. Check /status before you start, because an ANTHROPIC_API_KEY in the environment moves a Claude Code session onto API billing whatever your subscription says. There is no separate programmatic pot, because Anthropic announced a monthly Agent SDK credit in May 2026 and paused it before it took effect. The support article, checked 9 August 2026, says Agent SDK, claude -p and third-party usage draw from subscription limits, so scheduled jobs compete with your sessions.
13.3.7 The vendors ship the trigger wired up, so the gate is yours to build.
The August 2026 connector census found seven official servers able to fire a real, irreversible message: Apollo, lemlist, Instantly, Smartlead, Salesforge, Reply.io and Woodpecker. Most sequencer vendors do not withhold send capability, so the burden of not firing sits in the configuration you write. Apollo shows what a designed gate looks like, splitting its one-off email into a deliberate draft step and a send step. The Gmail and Outlook connectors split the other way from each other, which Chapter XIV covers. Keep all seven servers out of any loop that runs without you.
13.3.8 Two primitives make the split buildable, and the pattern you want has none.
Where a send-capable server must be connected at all, two primitives carry the weight. One is the annotation forcing a human prompt on every call, the other the deny and ask rules surviving bypass mode, both in Chapter XIV. The two-routine split, separating reading from sending, is best supported on Managed Agents, where each agent declares its own servers and vault secrets substitute at egress. The outbox pattern itself, an agent writing a draft that a separate process sends, has no first-party primitive behind it. Chapter XIV prices the Managed Agents surface and lists what you assemble the outbox from.
Anthropic documents most of the controls you need, and you have to switch them on yourself.
Most of the safeguards this guide asks for already exist in Anthropic's own reference documentation, and most have to be turned on deliberately. The five surfaces differ far more in who may operate them than in what they can do. And the permission system is the only place a send rule gets enforced rather than requested, because Anthropic's own telemetry says approval prompts decay into clicking. Each primitive below carries what the vendor documents as of August 2026, with the three places where the evidence stops short marked as such.
14.1 Choose the surface by who operates it, because the surfaces differ more than the model does.
A surface means one of Anthropic's own applications or interfaces, a different question from which model runs underneath. The five are five separate governance problems, because licence position, default permission state and administrator controls differ across them. Pick the surface for the person who will sit in front of it, then configure from there.
14.1.1 Five first-party surfaces cover the whole build, and each answers a different operator.
Chat, Cowork, Claude Code, Managed Agents and the API are the five places this work can run. They differ in who operates them, which set of Anthropic terms covers them, and what stops an action you cannot take back. Compare them on those three questions before you compare features, because the features converge while the controls do not.
14.1.2 The desktop app earns its recommendation, and it still assumes a repository.
The Code tab in Claude's desktop app removes the terminal, shows visual diffs, offers click-to-approve, and adds file and browser panes. That is why the August 2026 surfaces research calls the recommendation to business operators sound. The condition is that the app still assumes a git and repository mental model, version-controlled folders of files, and on Windows local sessions require Git for Windows first. New users start in Manual permission mode, which asks before each action, and that default is the one to keep.
14.1.3 Cowork is the better answer for an operator whose work is documents.
An operator whose day is research and briefs rather than code gets the same agentic architecture from Cowork with no terminal and no repository. Anthropic built Cowork explicitly for non-coding knowledge work, and the surfaces research names it the better surface for that reader. It runs on paid plans only, so there is no free path in for a client you are standing up.
14.1.4 Anything that runs without a human sitting there belongs on the API.
Consumer Terms section 3(7) bars access through automated or non-human means, except with an Anthropic API key or where Anthropic explicitly permits it. The rule that falls out is short: a human in the loop inside an Anthropic application can sit on a seat. And the instant work runs unattended or gets triggered by a script it belongs on the metered API under the Commercial Terms. Cost never overrides the clause, and the edges are yours to judge.
14.1.5 Move to Team or Enterprise the moment you touch prospect data.
Consumer tiers sit under Consumer Terms with a training toggle, and allowing training extends data retention to five years against thirty days if you opt out. Team, Enterprise and the API sit under Commercial Terms and are not used for training by default, which is the tiers research's stated reason for moving. Administrator-level read, write and send restrictions exist only on Team and Enterprise, so the tier you buy decides whether you can configure anything centrally at all.
14.2 Put the send rule where bypass mode cannot reach it.
Permission modes let an operator turn prompting down or off, so the question that decides your build is what still stops an action once they have. Anthropic answers it directly in its permission-mode and MCP references, and the classes of rule that survive are exactly the ones a send gate needs. Everything below is documented behaviour as of August 2026, not something inferred from watching the tools run.
14.2.1 Anthropic's own telemetry says people approve roughly 93% of permission prompts.
The figure comes from Anthropic's engineering write-up on containing Claude, where users approved roughly 93% of the permission prompts they were shown. Anthropic's own word there for per-turn human approval is "fallible".
14.2.2 The vendor's answer to a decaying prompt is the environment, so copy it.
The same write-up names three durable controls: egress allowlists, which limit where a machine may send traffic, credentials isolated from the sandbox, and enforced permission rules. Each keeps working while the operator is tired, distracted or moving fast, which is the property a click does not have.
14.2.3 One MCP annotation forces a prompt even when prompting is switched off.
Setting _meta["anthropic/requiresUserInteraction"]: true on an MCP tool forces a human permission prompt on every call. Anthropic documents that it holds in acceptEdits, auto and bypassPermissions modes, and that dontAsk mode denies the call outright rather than running it quietly. It is the strongest single primitive available to you, so put it on every tool that can contact a stranger.
14.2.4 Deny rules and ask rules survive bypass mode, and allow rules do not.
Anthropic's permission documentation states plainly that deny rules and explicit ask rules keep working when an operator switches into bypass mode. Allow rules stop mattering there, because bypass already permits everything they would have permitted. Write your send restrictions as deny rules and ask rules, because those two classes survive the setting most likely to get switched on under deadline pressure.
14.2.5 Switch bypass mode off for the whole organisation in one line of settings.
Managed settings accept permissions.disableBypassPermissionsMode: "disable", which removes bypass mode from every session in the organisation rather than trusting each operator to leave it alone. Set it before anyone wants it, because the moment it gets wanted is the moment nobody will agree to it.
14.2.6 Hooks run before the mode check, so a hook stops what a mode would permit.
A hook is a piece of your own code that Claude Code runs at a defined point in its loop. Anthropic documents hooks as running ahead of the permission-mode check, so a hook can block an action that bypass mode would wave straight through. That ordering makes hooks the right home for any rule you refuse to have overridden, such as a suppression lookup or a daily send cap.
14.2.7 The outbox pattern is sound and Anthropic ships no primitive for it.
The pattern is plain: the agent writes a draft into a queue, and a separate process holding the sending credential picks it up later. Anthropic documents none of this as a feature, so you assemble it from tool restrictions, permission policy, vault scoping and human approval. Say that out loud when you scope the work, because a reader expecting a switch will hunt for one and then skip the pattern entirely.
14.2.8 The Gmail connector cannot send and the Outlook connector can.
Most operators never send through these connectors, because the sequencer does the sending through the Gmail or Exchange mailbox connected to it. The distinction matters in one case only: when you want Claude itself to send a single message directly. There, Anthropic's Gmail connector is draft-only, its documentation saying "The send function is not enabled". The Microsoft 365 and Outlook connector does send, through outlook_send_email and the Mail.Send permission, once an administrator enables write tools.
14.3 Skills hold your own procedures, while MCP holds everybody else's product.
The two extension points look interchangeable from a distance, and the August 2026 skills research shows the go-to-market ecosystem has already chosen between them. Vendor capability arrives as MCP servers and connectors, while skills are where your repeatable procedures live.
14.3.1 A skill is a folder with a SKILL.md file, and that is the whole standard.
Anthropic publishes the format openly and skills are listed at agentskills.io, so a skill is portable text rather than a product you get locked into. A procedure written as a skill travels with you when the tooling underneath it changes, which is the ownership rule applied to process. Write your list pull, your account research pass and your suppression check as skills for the same reason you keep them in version control.
14.3.2 Progressive disclosure keeps a large skill library nearly free until it fires.
Anthropic's documented loading model reads about 100 tokens of metadata per skill at startup. The body, kept under 5,000 tokens, loads when the skill triggers, and bundled files arrive only on demand. A shelf of thirty written procedures therefore costs a few thousand tokens of standing context instead of a full prompt, so write more skills and shorter prompts.
14.3.3 Set disable-model-invocation on every skill that can send.
The flag disable-model-invocation: true blocks the model from calling a skill on its own while leaving the manual /skill-name call working normally.
14.3.4 Most go-to-market vendors are missing from the official marketplace.
Anthropic's marketplace held roughly 255 to 267 plugins when the skills research counted it in August 2026. Apollo, ZoomInfo, Hunter, Lusha and Explorium were present; Clay, HubSpot, Salesforce, Attio, Outreach, Salesloft, Gong, Instantly, Smartlead and Lemlist were not. Lemlist publishes skills on its own site instead, so the absence is a distribution decision rather than a limit on what anyone could build.
14.3.5 Vendor capability reaches you as an MCP server, so read there first.
The skills research calls this its most important finding for anyone deciding what to publish: most Claude-plus-vendor integrations in go-to-market are MCP servers and connectors rather than skills. To learn what a sequencer or a CRM can actually do from inside Claude, read its MCP server documentation and ignore the marketplace listing. The listing will mislead you about which vendors are reachable at all.
14.4 "Cloud routines" names three different products, and the difference decides your controls.
The August 2026 unattended-workloads research counted roughly fifteen distinct unattended surfaces across five families. The popular phrase covers at least three of them, differing in identity, billing and trigger model. Choosing one without knowing which you chose is how a build ends up with no mid-run control and no honest failure signal. The three are named below, the trade priced, and the work split so the routine that reads never holds the credential that sends.
14.4.1 Three features on three surfaces share one nickname.
Routines in Claude Code, Scheduled tasks in Cowork and scheduled deployments in Managed Agents are separate features, and no single cloud-routines product exists. They differ in which identity runs the job, which account carries the bill, and what is allowed to trigger a run. Name the one you mean in your own runbook, because a support answer about one of the three does not transfer to the other two.
14.4.2 Autonomy trades directly against mid-run control.
The governing rule fits one line: the more autonomous the surface, the weaker the mid-run human controls and the less it tells you on failure. Routines and Claude Code on the web run with no per-action approval by design, while the strongest gates live in the self-hosted primitives. Read that as a price list rather than a warning, and pay it deliberately for the jobs where nothing irreversible can happen.
14.4.3 A green status means the session exited cleanly and nothing more.
A routine that fetched nothing, wrote nothing and drafted nothing still finishes green, because the process ended without throwing an error. Your success check therefore has to be an artifact you can count, such as rows appended or drafts created. The routine should fail loudly whenever that count comes back zero.
14.4.4 Cowork scheduled tasks are the one scheduled thing a seat licenses.
The licensing research names Cowork Scheduled Tasks as the one documented exception to the rule that a seat cannot run scheduled work. It is available on all paid plans and runs in the cloud with the machine closed. It qualifies because the seat-holder starts and reviews it inside Anthropic's own application rather than triggering it from outside. Assume every other scheduled arrangement belongs on a metered API account until Anthropic documents otherwise.
14.4.5 Split the work into two routines, one that reads and one that sends.
The reading routine holds research access and never the sending credential; the sending routine holds the credential and never reads untrusted text. The split also buys you two schedules, so the sending routine can run only in the hours when a human is awake to watch it.
14.4.6 Managed Agents supports the split best, because the scoping is already built.
Managed Agents scope MCP servers per agent, the shape the read-versus-send routine split needs, the documentation's own example being that "only the researcher declares the GitHub MCP server, so the coordinator does not have access." Secrets are session-scoped through vaults and substituted at egress, so that "the agent never sees the secret value." The bill is $0.08 per session-hour of active runtime on top of tokens, roughly $58 a month for an agent that never stops. Weigh that against building the same two properties yourself in engineering time.
14.4.7 Three things here are unsettled, and resolving them would mean inventing evidence.
Anthropic documents no first-party outbox primitive, and it publishes no safe-harbour examples for what counts as acceptable tooling on a seat. Both gaps are recorded in the August 2026 research rather than guessed at. The third is the Agent SDK credit announced and then paused before it took effect. So SDK and scheduled work still draws on the same subscription allowance as your own session.
Buy the data layer as one tool, and fear the address marked valid.
Almost nothing an agent does at this layer is irreversible, because the pure finders in the August 2026 census of official vendor MCP connectors read and enrich only. Name the four exceptions first: Apollo, Hunter, Snov.io and Clay can each trigger a real send, Apollo's official server via one-off emails and sequence adds. Outside those four, this is the cheapest place to grant an agent real autonomy and the most expensive place to be wrong. An address marked valid that bounces costs you the domain rather than a credit. The whole layer runs under one hundred and forty dollars a month at this volume, on prices checked 9 August 2026.
15.1 Buy one waterfall as a single tool, because chaining providers yourself is a second job.
A waterfall is one service that queries a stack of data providers in order and returns the first good answer. You can buy that as a finished product with a single endpoint, or assemble the same sequence yourself from three or four separate vendors. The August 2026 contact-data research settles that choice against building for anyone running a few hundred contacts a month.
15.1.1 DIY chaining starts paying back only above five hundred contacts a month.
Two independent practitioners put the break-even in the same place, and both sell services that would benefit from the opposite answer. Ziel Lab, a revenue-operations agency that built waterfalls for nine go-to-market teams, states it plainly: "Clay is wrong if you process under 500 contacts a month. The subscription cost outweighs the savings until you hit volume." TechTower, a go-to-market engineering agency, lands on the same line, saying the bought product below it "is cheaper once you count the hours honestly". At the send floor Chapter III sets, 250 to 300 sends a month, you are nowhere near that line.
15.1.2 A hand-built chain is five separate maintenance jobs under one name.
TechTower lists what you take on: "integrate each provider, handle rate limits and failures, dedupe results, and fix the whole thing when an API changes underneath it". Florian Martens, writing independently, adds credit balances tracked across several providers, catch-all addresses handled by hand, and wrong results caught by nobody. None of that work is visible on the day you wire it up, and all of it recurs for as long as you run it.
15.1.3 One tool call leaves your agent one place to look when something fails.
The bought waterfall handles sequencing, retries, deduplication and rate limits inside its own walls. So your agent makes one call and reads one response envelope with a single field naming the provider that answered. A hand-built chain works against several rate limits, credit balances and failure modes, each failing differently, and latency compounds. Sequential runs take minutes per contact at products like BetterContact and hours across a large batch.
15.1.4 Reordering the providers moved one operator's bill by four times.
Ziel Lab reports a well-tuned waterfall at $0.08 per enriched contact against $0.52 from a single premium provider, a genuine six-and-a-half-fold saving. The same report warns the saving is operator-dependent: reshuffling the provider order changed the bill by 4.1 times on the identical list. A number that swings that far on a configuration choice is not a number you can budget against at low volume.
15.1.5 Buy to learn, and build only once your volume and your list stop moving.
TechTower's rule is "buy to learn, build to scale", and it names building while your ideal customer profile is still moving as the classic trap. DIY chaining wins only at high and stable volume, above five hundred to a thousand contacts a month. It wins sooner if you already hold direct contracts with the providers. Neither condition describes the operator this guide is written for, so revisit the question when one of them changes.
15.2 Buy on bounce rate, not on coverage.
Vendors compete on coverage, the share of your list they can find an address for, and it is the wrong number to choose on. A miss costs you one prospect, and on a pay-per-valid vendor nothing at all, because you are only charged for results. An address marked valid that then bounces costs sender reputation, and at the researched band, a few hundred contacts a month, one bad batch can end the program. Compare vendors on how often their valid addresses bounce, and treat high coverage with a high bounce rate as the worse product.
15.2.1 A miss costs you a prospect, while a false positive costs you the domain.
Ziel Lab states the consequence without hedging: "Never push an unverified email into a sequence. The deliverability damage from one batch of bounces will outlast the campaign by months." Ziel Lab sells waterfall builds, so the claim runs against its own commercial interest. Design against that asymmetry, because a miss is recoverable next month and a burned domain is not.
15.2.2 Fifty bad addresses on a fresh domain is all it takes.
Two vendor readings of the same ceiling sit close together: Prospeo holds total bounces below 2% and hard bounces, permanent rejections by the receiving server, under 1%. Typpout puts best practice below 3% to keep sender reputation with Gmail and Microsoft. At the 250-to-300-send floor Chapter III sets, six to nine bounces cross the 2-to-3% line. One fifty-address batch of unverified rows at typical database bounce rates produces ten to fifteen. Ziel Lab prices the recovery at "three months rebuilding domain warmup".
15.2.3 One in five found addresses failed the second verification pass.
A March 2026 write-up from Cold Outreach Stack, which carries a mild affiliate interest, documents a hand-built three-step chain run across ten thousand contacts. The chain reached 85% coverage, yet "about one in five found emails did not survive the final verification path". The addresses that did survive cost roughly $0.03 to $0.06 each fully loaded.
15.2.4 A re-check of nine hundred verified records found nineteen percent valid.
Community re-verification posted through the first quarter of 2026 took a 900-lead sample of Apollo exports already labelled "Verified": roughly 19% proved actually valid and 60% were catch-all. The same testing reads single-source cold sends bouncing at 32 to 38%, against 10 to 14% where a waterfall and a verification step were used together. Grade all of it directional rather than settled, because it reached the research re-reported through vendor blogs and the original source could not be confirmed.
15.2.5 Vendor accuracy pages and user-reported readings differ by more than twenty points.
Apollo's own site markets "91% email accuracy". Amplemarket sells into the same category rather than auditing it. It published user-reported accuracy of 65 to 70% in 2026 and user-reported bounce rates averaging 20 to 30% on cold sends. It also notes that applying Apollo's own "Verified Emails" filter drops the database from 275M contacts to 96M, meaning 65% of the contacts carry unverified addresses. Apollo's insights page concedes that no independent audited head-to-head benchmark exists, which describes the honest state of the whole category.
15.2.6 The one large independent benchmark reorders the rankings, and its author competes in them.
Dropcontact ran a 20,000-contact benchmark in 2026 across fifteen providers, and the results do not match the marketing. Enrow placed third at a 40.9% effective hit rate with 2.3% hard bounces. The lowest hard-bounce rates went to Dropcontact itself at 0.9% and Findymail at 1.1%. And FullEnrich, which advertises roughly 80% hit rates, found 483 emails per 1,000 leads with a 3.6% hard-bounce rate and an 11.7% wrong-domain rate. Dropcontact sells into the category it benchmarked, so carry the ordering rather than the decimals.
15.2.7 Prefer a graded risk score to a binary valid-or-invalid verdict.
Microsoft began rejecting non-compliant bulk mail outright on 5 May 2025, returning 550 5.7.515 to senders above five thousand a day into its consumer domains. That threshold sits far above the volume this guide plans for, so read it as direction. Authentication and reputation now decide delivery alongside address validity, which makes a flat valid-or-invalid verdict less informative than it once was. Favour providers returning a graded score, such as the A to F risk grade MoltSets attaches to every email. Or favour the explicit catch-all handling that Enrow, Icypeas and FullEnrich document.
15.3 Pay only for results that turn out to be real.
Billing models split this market cleanly, and the split is a better selection criterion than any accuracy claim on any vendor page. Some providers charge when they find and verify an address, while others charge for every lookup including the ones that fail. At a few hundred contacts a month, with the one independent benchmark putting real hit rates in the forties, that difference decides your effective unit cost. Every price below was checked on 9 August 2026 unless another date is given, and this field changes weekly.
15.3.1 Let the billing model decide the vendor, because pay-per-lookup charges you for misses.
Prospeo, LeadMagic, Findymail, Anymail Finder, Enrow, FullEnrich and BetterContact all charge only on a verified result. Findymail adds a bounce guarantee under 5% with refunds when one gets through. Hunter charges per lookup including misses, and Snov.io charges per search regardless of outcome against a roughly 20% verified rate in one benchmark. The list-price gap between the two groups looks small on a pricing page, and the effective gap is wider.
15.3.2 Buy the database cheaply, and treat every address it hands you as unverified.
Apollo Basic runs $49 per user per month on an annual plan and $59 monthly, which buys the search and filtering layer at a price no specialist matches. Browsing and filtering burn nothing; revealing an email costs one credit, a mobile number costs eight, and credits do not roll over between billing cycles. Given the accuracy gap measured above, route every Apollo address through your waterfall or a dedicated verifier before it reaches a sequence, on every plan and without exception.
15.3.3 The whole data layer costs under a hundred and forty dollars a month.
The research prices a minimum viable stack at $78 to $137 a month. That covers Apollo Basic at $49 as the database, one waterfall doing both enrichment and verification, and an optional cheap backstop verifier. FullEnrich Starter at $29 covers 500 credits, roughly 240 resolved work emails out of 300 attempts at about $0.12 each. BetterContact Pro at $49 does the same work at about $0.20 a resolved contact, while MoltSets at $27 lands between $0.11 and $0.15. Those are list prices against one assumed hit rate, so instrument your own first month before trusting the ranking.
15.3.4 Clay costs roughly six times more per resolved contact at this volume.
Clay Launch is $185 a month, checked 7 August 2026, and carries 2,500 Data Credits against the 720 to 1,200 that a 240-contact month consumes. The plan is over-provisioned by a factor of two or more, so the same 240 resolved contacts land at roughly $0.77 each against $0.12 from the cheapest waterfall. Clay aggregates over 150 providers as a workflow builder rather than a pure waterfall, and that capability is real. The objection is to the purchase at this volume, not to the product.
15.3.5 Mobile numbers cost five to ten times an email at most vendors, and far more at one.
Prospeo and Findymail both charge ten credits for a phone number against one for an email. That is $0.39 and $0.49 on their entry tiers, falling toward $0.10 on their largest. LeadMagic is the cheapest at five credits, roughly $0.035 to $0.15, while Enrow is the outlier at forty credits, about $0.60. Independent testing puts direct-dial accuracy near 60%, so budget mobile as a separate line and decide deliberately whether the channel earns it.
15.3.6 Two naming traps and one closed beta sit between you and a wrong purchase.
The "Apollo MCP Server" published by apollographql.com belongs to a different company and does GraphQL tooling; the one you want is mcp.apollo.io. "Clay" likewise names two unrelated products that both ship official servers, the go-to-market enrichment platform and a personal-relationships address book. MoltSets, the cheapest option in the worked comparison, remains in closed beta as of August 2026. It carries United States data only at launch, which costs a US-only list nothing. And no independent accuracy benchmark has been published for it yet.
15.3.7 Write down the two measurements that reopen this purchase.
The first is volume, because crossing a thousand contacts a month of stable targeting starts to favour a chain you own or a workflow builder you configure. The second is your own measured hard-bounce rate. Above roughly 2% hard, or 3% of total bounces, add or upgrade a dedicated verification pass whatever the source of the data. Several figures in this layer are marked undocumented in the research, Apollo's Organization-tier rate limits among them, negotiated with a customer success manager rather than published.
Buy no signal you cannot quote back to the buyer.
A signal is a public statement by a buyer that tells you who to work and when. It might be a forum question, a complaint naming a product, a hire implying budget, or a stated reason for leaving a rival. Each is a dated sentence a named person wrote, so you can quote it back, check it yourself, and act the day it appears. A purchased intent score is a different product, a probability derived from browsing data, which Chapter IV weighs separately. The quotable kind is the subject here: what it costs to collect, which platforms permit it, and how thin the proof it lifts replies.
16.1 Observed statements and bought scores are different evidence.
16.1.1 A signal you can quote beats a score you cannot audit.
Purchased intent data is a probability a vendor derived from bidstream, co-op or de-anonymisation data and sold to you. An observed public statement is something a named buyer wrote themselves, in public, with a timestamp. The August 2026 listening research holds those apart, and so should you. Chapter IV's verdict on bought intent stands: across the research base no independent study shows a vendor's buying signal predicting a B2B purchase.
16.1.2 Four trigger types are worth watching, and each one is a sentence somebody wrote.
An alternative query is someone asking in public for a better option than a named incumbent. An implementation blocker names the product and the failure; a budget signal is a hire or a filing implying newly unlocked spend. Competitor detraction is an explicit reason given for leaving a rival, and each type is quotable, dated and attributable, which a score is not. Author's note: this four-type taxonomy is mine, published in 2026 at productiveai.com; the commissioned research corroborates the collection problem and the noise, not the taxonomy itself.
16.2 LinkedIn holds the densest signal and permits the least collection.
16.2.1 Section 8.2 bars any software that touches a member session.
LinkedIn's User Agreement, in its section 8.2, bars members from developing, supporting or using "software, devices, scripts, robots, or any other means or processes (including crawlers, browser plugins and add-ons)" to scrape the services. It also bars using "bots or other automated methods to access the Services". The August 2026 research quotes that from LinkedIn's own help pages, read on 9 August 2026. It covers the browser extensions most listening products ship, and that one fact decides the design that follows.
16.2.2 The enforcement is real, and one vendor is already gone.
hiQ Labs lost on breach of contract rather than computer-crime law, taking a $500,000 consent judgment in December 2022. The judgment also ordered it to stop scraping and to destroy the data. LinkedIn then sued Nubela and Proxycurl in January 2025 over accounts used to harvest profiles at scale, settling mid-2025 with a permanent injunction. Proxycurl shut down in July 2025 at around $10M of annual revenue, and LinkedIn pulled Apollo's and Seamless's company pages in March 2025. Treat the terms as enforced rather than decorative.
16.2.3 Follow instead of connecting, because following carries no published cap.
Following is a public action LinkedIn encourages, with no published limit at the few hundred people this motion watches. Connection invitations are the constrained action, at a weekly ceiling LinkedIn declines to publish and practitioners place near 100 to 200, a figure the research marks as folklore. Build the watch list out of follows and the cap never arrives.
16.2.4 A Sales Navigator seat buys off the limit that actually stops you.
Free accounts carry a monthly commercial-use search cap that LinkedIn refuses to quantify, plus a documented ceiling near 100 pages of results per query. Sales Navigator Core at $119.99 a month, or $1,079.88 a year, lifts the commercial-use limit for prospecting. It also adds saved searches with alerts on job changes, posts and news. Nothing exports, because the data stays inside LinkedIn by design, so price the seat as the cost of the densest source rather than a tool you own.
16.2.5 The signal sits in replies, and no first-party search reaches them.
LinkedIn's content search returns posts, and no way exists to search comments and replies for a keyword, first-party or otherwise. That matters because the useful statements usually sit in a reply to a comment rather than a top-level post. A person following named accounts sees that activity in their feed and notifications, which is the only permitted route to it. Any product claiming comment-level LinkedIn monitoring is reaching it by a method it does not describe, the same class as the extensions section 8.2 names.
16.2.6 The morning routine is inside terms at every step.
Follow the watch list, work a small set of saved searches and the feed by hand, and read for the four trigger types: an alternative query, an implementation blocker, a budget signal, or competitor detraction. Reply helpfully in public, then reach out privately by hand or by InMail. The August 2026 research finds every step permitted, provided a person does the reading and the replying. The cost is time, and nobody has published a measurement of it. The research estimates 45 to 75 minutes a day for a few hundred people. It marks its own estimate as a gap, so measure your own in week one.
16.3 Run the program off LinkedIn, and bolt the watch list, your LinkedIn follow list, onto it.
16.3.1 G2 buyer intent is the best permitted source, if you have a listing to attach it to.
G2 shows a listed vendor which accounts are researching its category, its alternatives pages and its comparisons, inside my.G2. The licence comes with the listing, which makes the data first-party and free of terms friction, and Capterra, Software Advice and GetApp work the same way. If your product belongs on those directories, claim the listing and start here before you price any listening product. A company with nothing listed has no my.G2 to open, which moves EDGAR, Reddit and the watch list up one place each. Settle listability on G2's own pages rather than assume it.
16.3.2 SEC EDGAR is free, public, and the cleanest budget signal available.
Full-text search across filings back to 2001 runs at efts.sec.gov, and the REST and JSON interfaces are free. The only conditions are ten requests a second and a descriptive user agent. Hiring and expansion language in an 8-K or an S-1 is a budget signal with a date attached and nobody between you and it.
16.3.3 Reddit carries the richest complaint conversation at an enterprise price.
Reddit's free tier is non-commercial only, and commercial access runs at $0.24 per thousand calls behind a minimum that 2026 sources report inconsistently. Most place it at $12,000 a month for a fifty-million-call tier, and one places it at $12,000 a year. At the low figure Reddit is the best source here; at the high one it is out of reach for one operator. Budget the higher figure and confirm with Reddit before committing. Self-service registration is closed, approval runs two to four weeks, and search lacks date ranges, comment search, and anything past a thousand-item pagination ceiling.
16.3.4 The sorting costs about eleven dollars a month.
Hand the feed, the incoming stream of signal items, to the cheapest model rather than the best one. At 500 items a day, roughly 500 input tokens and 50 output tokens each, Haiku 4.5 at $1 and $5 per million tokens comes to about $0.38 a day. That is near $11 a month before batch or caching discounts, and at 2,000 items a day it is about $45.
16.4 Nobody has proved this works, and everybody measuring it sells something.
16.4.1 No independent study compares triggered outreach against untriggered.
The claims are real and they are all vendor-published on their own customers. Unify's product page states that signal-led outbound drives 73% more replies than cold, footnoted to its own proprietary research. Autobound puts signal-based emails at 5 to 18% reply against 1 to 3% generic. LinkedOtter claims 15 to 25% against 3 to 5%, attributed vaguely to industry benchmarks with no denominator published. Grade D: vendor case studies, self-interested, no independent replication. Run the practice because the mechanism is sound, and hold your own numbers rather than repeating theirs.
16.4.2 Nobody has measured the noise in a listening feed either.
Vendors assert that raw mention counts are mostly noise and sell relevance scoring on the strength of it, yet none publishes a measured relevant fraction. The only non-vendor figures the research located were illustrative single-post examples from a 2015 blog. Expect most of what arrives to be irrelevant, and plan the human time around that. Record your own hit rate from week one, because you will be the only person holding one.
16.4.3 Public commenting first is sound on mechanism and unproven on numbers.
No independent measurement isolates a helpful public reply before private outreach as a reply-rate variable, and everyone publishing on it sells social-selling or engagement tooling. The mechanism is plain: you are recognisable when the private message arrives, and the recognition was earned on a surface the platform permits. That justifies the practice without letting you forecast from it.
16.5 The legal floor is the one you already have, plus one change most senders missed.
16.5.1 Quoting a public post adds no obligation, and removes none either.
CAN-SPAM attaches to sending a commercial email rather than to how you found the reason for it. So accurate headers, a real physical address and a working opt-out are required whether or not you reference something the person wrote. Summarising rather than quoting changes nothing in the statute, and a paraphrase that distorts what somebody said adds deception risk under the FTC's rules. Quote accurately and stay inside the compliance chores Chapter XXII lists.
16.5.2 California's B2B exemption has sunset, which puts your contacts in scope.
The CCPA and CPRA business-to-business exemption has lapsed, so a named person at a company is personal information like anybody else. That includes where they posted the thing you are reacting to in public. Lawfully obtained public information is partly carved out of the definition, and re-using it for outreach can still trigger notice duties. This is the sharpest state-law exposure in a US-to-US motion, and it is fact-specific. It is a reason to keep the audit log Chapter XVIII builds rather than a reason to stop.
One domain carries more mail than a low-volume team can check.
The sending layer is the cheapest part of this build and the one most often bought at ten times the needed size. Two architectures exist, separated by a phase change rather than a gradient. The question that sorts them is whether one or two mailboxes can carry your whole program. Below that line the hardware sits close to idle, because the person reading every message runs out first.
17.1 Your send volume sets your domain count, and low volume needs very few.
The multi-domain playbook you will read everywhere was written for 500 to 5,000 sends a day, which the August 2026 infrastructure research confirms. A few hundred sends a month needs one domain and one or two mailboxes. The work differs at each size: the large build warms mailboxes with artificial traffic, rotates sending between them, and manages each mailbox's reputation separately. In the small one the real mail does the warming, there is nothing to rotate, and one person reads every message before it leaves.
17.1.1 Work out how many mailboxes your volume needs before you buy anything.
Divide your monthly send target by twenty working days, then divide that daily number by the documented per-mailbox limit of 15 to 30 sends a day. Three hundred a month is about 15 a day, which one mailbox covers. Two or three mailboxes sit comfortably on one domain, so a program that size needs one domain. Do the arithmetic before you shop, because sending infrastructure is sold in blocks of ten or thirty mailboxes and the vendor's minimum will otherwise choose your architecture.
17.1.2 Below the line, at one-or-two-mailbox scale, warmup and the real campaign are one activity.
A pair of mailboxes has no spare capacity for manufactured traffic, so your real researched outreach is also the reputation-building traffic. Per the August 2026 infrastructure research, register the domain today, configure MX, SPF, DKIM and DMARC, and allow one to three days for DNS. Send five to ten real messages a day in week one, ramping toward the documented cap over about four weeks. The widely cited ladder of 5 to 10, then roughly 20, then roughly 30, then full volume was written for caps well above yours. The same report tells a fifteen-a-day sender to start at 5 to 10 and ramp to 15. Practitioner consensus favoring domains older than about thirty days is close to moot, and the four-to-eight-week wait Chapter III sets leaves your calendar.
17.1.3 Above the line, running a many-mailbox fleet, tending the estate becomes the actual job.
Rotation, pacing and per-box reputation management arrive together once a fleet exists, and connection maintenance becomes a standing cost rather than a one-time setup. The author has run the twenty-to-forty-domain version and reports OAuth connections dropping and needing reconnection, so the estate becomes a herd somebody has to tend. That labor buys one thing worth having at high volume: blast-radius control, a limit on how far one bad run reaches.
17.1.4 Synthetic warmup exists because a fleet has capacity it cannot fill honestly.
A warmup pool is a network of mailboxes that email each other to manufacture engagement signals, answering a problem only a fleet has. Twenty boxes cannot be fed with human-reviewed mail, so the empty capacity gets filled with traffic no recipient wanted.
17.1.5 Crossing the line means giving up reading every message.
Moving to a fleet requires abandoning human review of each message, and that review is what makes this build work. A program that quietly adds mailboxes has not scaled; it has switched to a different book, with different economics and a different ethical position. Decide which machine you are building in week one, and hold the decision against the pressure to add capacity.
17.1.6 Adding mailboxes behind one operator is the signature of a changed program.
Scaling this build multiplies operators, each carrying their own one or two mailboxes; it never multiplies mailboxes behind a single operator. Chapter IX reaches the same conclusion from human-capacity data, and the infrastructure evidence arrives there independently through vendor documentation and deliverability mechanics.
17.2 Your reading capacity stops you first, and the infrastructure is barely engaged.
Two ceilings apply to this layer. The vendors who sell mailboxes publish the first, per mailbox; the second is how many researched messages one person can properly check in a working day.
17.2.1 Always say whether a sending limit is per mailbox or per domain.
Sending limits get quoted without their unit constantly, and the same figure is safe or reckless depending on which is meant. Fifty a day per domain sits comfortably inside documented practice; fifty a day per mailbox is roughly double the consensus cap and more than triple the lowest documented one. When you read a limit anywhere, check which unit it attaches to before you plan against it.
17.2.2 The documented per-mailbox cap runs from 15 to 30 cold sends a day.
Maildoso publishes the most conservative figure at 15 cold sends a day per mailbox. Practitioner consensus sits at 20 to 30, with Mailflow Authority planning at 25 and Mailforge recommending 30 against a stated maximum of 100. All are per-mailbox numbers confirmed in the August 2026 infrastructure research. Read them as reputation conventions rather than platform limits, since Google rates a Workspace user at 2,000 messages a day. So fifteen sends uses under 1% of the quota you already bought.
17.2.3 No source documents a hard per-domain ceiling at this scale.
Convention puts two to three mailboxes on a domain, which is Mailforge's own guidance. Unify GTM holds that two secondary domains with two to three mailboxes each cover anything under fifty a day. Multiplying the documented caps by that convention derives roughly 40 to 90 sends a day per domain, or 800 to 1,800 a month. Mark the figure derived rather than documented, because the August 2026 research found no source publishing a hard per-domain cold ceiling at this scale.
17.2.4 The person reading is the ceiling, and you use a tenth to a quarter of one domain.
The review budget is the five to ten checked outputs a day Chapter IX sets, which at twenty working days is 100 to 200 messages a month. That reconciles with the 250-to-300-send floor Chapter III sets through the sequence: 250 to 300 sends is 75 to 100 contacts at three to four touches. So the person checks four or five new researched outputs a working day while the follow-up steps reuse the research. Set the two ceilings side by side: one domain carrying two or three mailboxes supports something like 800 to 1,800 messages a month at documented caps. A cross-check lands in the same band, since two mailboxes at 25 a day across 20 working days is 1,000 a month.
17.2.5 Twenty mailboxes would take twenty to fifty people doing nothing but reading.
Reaching a fleet honestly means sending four to ten thousand messages a month, which at the Chapter IX review budget requires twenty to fifty people reviewing full time. The fleet architecture is unreachable for this build rather than merely inadvisable, because at that volume you are running the other book entirely. The vendors whose minimums start at ten or thirty mailboxes are not overselling you; they are selling to somebody else. Mailforge starts at ten slots, Maildoso at a thirty-box bucket, Zapmail at $39 a month for ten, and MailReef sells a $240-a-month server holding 150 mailboxes or more.
17.3 Four DNS records on a second domain satisfy every rule that reaches you.
The requirements everybody repeats divide into two sets: one that applies above 5,000 messages a day, and an authentication floor that applies to every sender. Only the second reaches this build, and meeting it costs one afternoon and about eleven dollars a year. What follows is that floor, the domain decision underneath it, and the honest monthly price of the layer as of August 2026.
17.3.1 Reputation attaches to the domain, so an alias gives you no isolation.
Send cold outreach from a lookalike secondary domain that 301-redirects to your main site. A 301 redirect is a permanent forward that hands visitors and search engines to the real thing. An alias on your primary domain shares that domain's reputation exactly, so a spam flag earned by outreach lands on your invoices and your support mail too. This is the one piece of the high-volume playbook that transfers down intact, and every credible source in the research agrees.
17.3.2 Authentication, complaints, TLS and reverse DNS apply at every volume.
Four requirements reach you no matter how little you send. First, publish SPF, DKIM and DMARC, three DNS records that let a receiving server check the mail really came from you. DMARC at p=none satisfies the rule, and a real rua address makes the aggregate reports arrive. Second, send over TLS, the encryption that protects mail in transit. Third, keep valid forward and reverse DNS, which your mailbox provider handles on shared infrastructure. Fourth, hold user-reported spam complaints under the 0.3% ceiling Chapter I documents from Google's sender guidance, and ideally under 0.1%. Add a visible opt-out as well, even though one-click unsubscribe, RFC 8058, is technically a bulk-sender rule.
17.3.3 The 5,000-a-day rules do not reach you, and knowing that saves money.
Google's requirements for senders above 5,000 messages a day cover personal Gmail traffic at that scale, with enforcement ramped from November 2025 and rejections at the SMTP level. Microsoft began rejecting unauthenticated mail above the same threshold on May 5, 2025, answering with "550 5.7.515 Access denied". Yahoo's bulk requirements match, though the 5,000 figure comes from the joint February 2024 announcement rather than Yahoo's current page. None of those mechanics touch a dozen sends a day, so a vendor quoting them at you is quoting somebody else's problem to sell you somebody else's architecture.
17.3.4 Own the domain yourself, at a registrar whose renewal matches its first year.
Porkbun charges about $11.08 a year for a .com, registration and renewal alike as of August 9, 2026, and can register a brand-new name. Cloudflare Registrar sells at cost, roughly $10.44 to $10.46, but accepts only transfers and renewals, somewhere to move sixty days later. Namecheap's $10.98 first year renews at about $18.48, while GoDaddy renews near $22 and reclassified all users as "Business Customers" in a February 2026 terms change. Domain age is the one input money cannot reset, so the test is narrow: registered in your name or your client's, and transferable out. Buying through an infrastructure provider passes when it registers the domain to you, which several document, and spares you the registration and DNS wiring. Instantly is the documented failure: done-for-you accounts are configured exclusively for use inside Instantly, with domain ownership and administrator access retained and untransferable.
17.3.5 You cannot measure your own complaint rate, so read your replies instead.
Google Postmaster Tools needs roughly 100 to 200 Gmail-bound messages a day, a threshold Google does not publish and practitioners only estimate. So at a dozen sends a day you will see "insufficient data" permanently. Microsoft's SNDS requires proof of IP ownership you lack on shared infrastructure; Yahoo's feedback loop earns its setup only if a real share of your list sits on Yahoo addresses. What remains is reply rate as the honest placement signal, a free Mail-Tester check before any template change, and free HetrixTools alerts on blacklist listings. Watch domain-level listings in particular, because on shared infrastructure the provider owns the IP and you cannot be listed individually.
17.3.6 Cut standalone warmup first, because the evidence trend runs against it.
Google told GMass in January 2023 to shut down its warmup system or lose Gmail API access, and GMass complied after sending 1,295,152,830 warmup emails across 236,084 accounts. Apollo removed its native warmup in 2024 for volume pacing that builds no sender reputation, then relaunched it through third-party providers during 2025. Neither Google nor Microsoft publishes a statement penalizing pools, so grade the blanket claim suggestive rather than proven. And one tester of nearly all the major warmup tools found none demonstrating meaningful improvement. At $15 to $96 an inbox a month, this is the easiest line in the budget to delete.
17.3.7 The whole sending layer costs between $8 and $75 a month.
The minimum build runs about $8 to $13 a month: one .com amortized at roughly a dollar plus one mailbox at $4 to $7. Exchange Online Plan 1 is the $4 option; Google Workspace Business Starter and Microsoft 365 Business Basic are $7 each. The manual ramp and free monitoring tier cost nothing, so the cautious build's $45 to $75 is almost entirely a paid sequencer at $33 to $47. Smartlead or Instantly buys scheduling, reply detection, and an interface an agent can call, and what the agent activates is a campaign rather than a single message. Buy Instantly only as software, never its done-for-you domains; a second mailbox on a second domain adds around $8 as suspension insurance rather than capacity. Revisit dedicated infrastructure only above roughly 30 mailboxes or 750 sends a day, and re-check every figure at checkout because prices drift weekly.
Sending tools block addresses, not people, so keep your own do-not-contact list.
Every block list in this field keys on an identifier, one email address or one domain, so a suppressed person with a new work address is contactable again. That is a structural property of the sending layer rather than a vendor failing, and no tool selection fixes it. The answer is three layers you assemble yourself: one canonical opt-out on a person record, mirrored into every sender by API. The same opt-out is written to an append-only log in a store you own. Meeting held has the same shape of problem, because no customer record can prove attendance and only the conferencing tool can.
18.1 Every sending platform blocks an address, which is not the same as blocking a person.
The August 2026 suppression research found three senders running a real workspace-wide block list with a documented API: Instantly, Smartlead and lemlist. Each still keys entries on an email address or a whole domain. Apollo screens phone numbers against official do-not-call lists. Amplemarket checks exclusion lists and a recently-contacted window before enrolling a lead, a guardrail at a different layer with a documented override flag. All of that stops the next campaign while doing nothing to keep a promise made to a person, so the record of who opted out belongs somewhere else.
18.1.1 Every sender keys suppression on an identifier, so a job change defeats it.
A block list is the roster of addresses a sending platform refuses to send to, and every platform checked keys its entries on a string of text. When a suppressed person moves companies and picks up a new work address, the old entry stops matching and that person becomes contactable again. Chapter X puts the annual job-change rate at about 20%, so this is routine, not an edge case. The August 2026 research names it as the structural reason the canonical record cannot live inside the sender, and it applies to every vendor in the category equally.
18.1.2 lemlist comes closest, blocking the whole person across every channel it knows.
Its "Do not contact" setting blocks a contact across every channel and every campaign at once, and a job change still defeats it, because the new address is a stranger to the list. A separate channel-level unsubscribe blocks one email address, phone number or LinkedIn URL for the case where a single channel is unwelcome. The API marks a contact through POST /api/v2/unsubscribes/contacts/{contactId}, and the Activities tab records whether the unsubscribe came from a user, the API or a bounce. That per-contact history and Instantly's audit log are the two best vendor audit records the research found. Unsubscriptions apply globally across campaigns and future steps get skipped mid-campaign, so the block takes hold without restarting anything.
18.1.3 Instantly and Smartlead let your software update their block lists automatically.
Both tools accept block-list changes from your own software through their APIs, so your agent keeps every list in sync with nobody copying entries by hand. Instantly's Block List Entry object sits at workspace level rather than per campaign, so one write covers every campaign you run. Its v2 API supports create, list, get, patch and delete on /api/v2/block-lists-entries, where each entry's bl_value and is_domain boolean block one address or an entire domain. Instantly also ships an Audit Log object and webhook events, better native auditability than most senders offer. Smartlead's Global Block List runs through /leads/get-domain-block-list, /leads/add-domain-block-list and /leads/delete-domain-block-list, plus a per-lead POST /leads/{lead_id}/unsubscribe. It applies across every campaign, auto-excludes future imports, and can be switched on mid-campaign.
18.1.4 Smartlead's suppression flags default to off, and its client scoping can leak.
Two hazards come from practitioner reporting rather than vendor documentation, credited in the August 2026 research to GenFlows and marked as corroboration. The block-list, unsubscribe and duplicate flags ship switched off, so a fresh workspace does not enforce the list you just finished building. The list is also scopable per client, so one client's suppressions can end up filtering another client's targeting whenever the scope is set wrong. An agency should therefore never treat that list as the truth, and the research says so in as many words.
18.1.5 Six named senders are undocumented on suppression, so check before you wire one up.
Salesforge's suppression specifics were not found in primary documentation during the August 2026 pass, only a generic blog post on global block lists. Woodpecker, Reply.io, QuickMail, Mailshake and Saleshandy were not verified against primary docs in that pass either. The category pattern is a workspace-level block list with CSV import and export behind a REST API, but an unverified list is not a control you can rely on. Confirm the endpoint and its scope yourself before you make any of them the enforcement point in a live send path.
18.1.6 A block list you cannot export is a list you will eventually lose.
Practitioner reports across cold-email communities describe suppression lists stranded when a subscription lapses and API access gets cut off. Smartlead documents a CSV download of the global block list, the best case among the three. Instantly and lemlist return their lists as JSON over the API, so an export is something you script rather than click. The canonical copy belongs where export is guaranteed: your own store, or a customer record you can dump to CSV on demand.
18.2 Canonical suppression belongs on a person record, mirrored outward and logged by you.
The August 2026 research recommends splitting the job across three homes rather than picking a single winner. Contact state and the canonical opt-out live in a lightweight customer record keyed to the person. Every suppressed identifier gets mirrored into each sender's block list by API. An append-only log in your own files or a small database carries the who, when and why, which neither of the other two layers gives you reliably.
18.2.1 Attio has the right identity model and none of the enforcement.
A Person in Attio carries a multi-value email_addresses attribute with normalized sub-properties for the address, the domain and the root domain. Each Person links to a Company through a two-way record reference, so one opt-out covers every address the person holds and survives a job change. Attio ships no native do-not-contact attribute, so you build a custom checkbox and your agent checks it before every send. The "Protected recipients" and "Blocklist" features hide email and calendar data for privacy. Neither is documented as blocking Attio's own sends, an assumption that is easy to get expensively wrong.
18.2.2 HubSpot enforces the opt-out itself and charges you in data-model weight.
HubSpot ships an "Unsubscribed from all email" property alongside per-subscription-type opt-outs and refuses to send to a contact marked either way. The platform enforces the opt-out itself rather than trusting your agent, the strongest agent-checkable opt-out of any customer record in this research. Subscription status is separate from marketing-contact status, and the free tier holds up to 1,000,000 records with non-marketing contacts free of charge. The cost is a heavier data model and a rate-limited API: Free and Starter allow 100 requests per 10 seconds. The Search API runs at roughly four to five requests a second, so for bulk pre-send checks cache the suppression set into your own store and check the cache.
18.2.3 Both customer records ship an official agent connector, and one of them is easy to mistake.
Attio's hosted MCP server, the standard interface an agent calls a tool through, covers records, lists, notes, tasks, meetings and emails. It sits at mcp.attio.com/mcp, uses OAuth, and reached general availability in February 2026 with exactly 37 tools, on every plan including Free, for every workspace member. Webhooks, schema writes, headless authentication and the suppression check itself go through the REST API, the check on POST /v2/objects/{object}/records/query. That API allows 100 read requests and 25 write requests a second, ample at a few hundred contacts a month. HubSpot's remote customer-record server at mcp.hubspot.com reached general availability in April 2026 and uses OAuth PKCE. A separate local Developer server, general since February 19, 2026, covers apps and the content system; confusing the two is a documented trap, since only the remote server touches customer data.
18.2.4 Check the price yourself, because Attio moved it in July 2026.
As of August 2026 Attio Free costs nothing and allows three seats, 50,000 records, three objects and 200 emails a month. Plus runs $35 per seat a month on annual billing, or $44 monthly, with ceilings of ten seats, five objects and 250,000 records. Pro runs $79 annual or $99 monthly for unlimited seats, twelve objects, a million records, call intelligence and sequences. Attio raised Plus and Pro in July 2026 from $29 and $69, so any comparison still printing the old pair is stale. The HubSpot Starter figures of $20 per seat monthly, or $7 per seat annually, are corroboration rather than the vendor's own page; verify both there before you commit, since prices move weekly.
18.2.5 Mirror every suppressed identifier into each sender on a schedule and on every webhook.
The single most-reported failure in this field is a customer-record opt-out that never reaches the sending platform, so the person gets contacted after asking not to be. Close it by pushing every suppressed identifier into each sender's workspace block list through its API, on a timer and on every record-updated event. A webhook is a message the customer record pushes to your code the moment something changes, and Attio fires them on record created, merged, updated and deleted. The three write targets verified in this research are Instantly's block-list entries endpoint, Smartlead's add-domain-block-list, and lemlist's unsubscribes endpoint.
18.2.6 Keep the append-only record in a store you own.
Neither the customer record nor the sender gives you a strong, exportable who-when-why history, and customer records are mutable, so their audit value stays weak without your own log. Append-only means rows get added and never edited or deleted, which is what makes the log worth anything to a regulator, a client or a court. Three shapes work; the first is a SQLite file with a people table keyed on a canonical person_id and an identifiers table linking many addresses to one person. Its append-only suppression_events table carries the person, the identifier, the reason, the source, an ISO timestamp and the actor. A Git-tracked CSV gives the same immutability through commit history, every change landing as a versioned, timestamped and attributable commit. Managed Postgres on a Neon or Supabase free tier runs roughly nothing to $25 a month when you need concurrent access.
18.2.7 Keep one row per person, and merge duplicates the moment you find them.
A suppression flag on one record cannot protect you from that person's duplicate, and the research names duplicates defeating suppression as a documented, common failure. At a few hundred contacts this is housekeeping rather than an identity problem: keep one row per human in your own store, attach every email and profile URL you learn to it, and merge rows when you discover they are the same person. No universal identifier exists, and even LinkedIn URLs vary for one person, so the merge is a judgment you make rather than a key you look up. Attio's People, HubSpot's Contacts and lemlist's do-not-contact all protect you only when the record really is one per person.
18.3 Meeting held is a conferencing fact, and your customer record cannot verify it.
Chapter XI counts meeting held among the four numbers worth acting on, which makes how you record it an architectural question. Every customer record examined here treats the meeting outcome as a field a human edits by hand. Positive proof a meeting happened exists in three places, all in the conferencing layer.
18.3.1 Every customer record's meeting outcome is a field somebody has to remember to change.
HubSpot's meeting outcome offers Scheduled, Completed, Rescheduled, No-show and Canceled, and it does not flip to Completed when the meeting time passes. HubSpot staff and community confirm it as manual, and it sits among the most up-voted ideas on HubSpot's meetings backlog as of August 2026. Attio's meeting field and Call Intelligence behave the same way, staying manual unless a conferencing integration writes to them. A field that depends on somebody remembering undercounts, on the exact number Chapter XII uses to deliver your day-sixty verdict.
18.3.2 Attendance reports are the only positive proof an agent can read.
Microsoft Graph exposes a meeting attendance report whose records carry join and leave times for Teams meetings, which an agent can read as proof the meeting happened. Zoom's participant report API does the same job, alongside its meeting_ended and participant joined and left webhooks. Two caveats apply: duplicate join and leave rows, and data that can lag the meeting_ended event. A recording from Fireflies, Gong or Grain is proof by existence, since the artifact only exists if the meeting did. Google Meet's conference records and participants likely serve the same purpose; that one was not verified against primary documentation in the August 2026 pass.
18.3.3 Accepting the meeting does not mean they showed up.
Google Calendar's API returns an attendee responseStatus of accepted, declined, tentative or needsAction, which is an RSVP and nothing more. Microsoft Graph exposes the same responseStatus, and its attendance report is a separate object. Treating an accepted invitation as a held meeting inflates your most reliable count with people who agreed and then did not show. Chapter XII counts a booking from a confirmed calendar event and rests the verdict on meetings held.
18.3.4 Calendly and Cal.com automate the no-show, so held stays an inference.
Calendly fires invitee.created on a booking and invitee.canceled on a cancellation, and exposes invitee_no_show.created and deleted alongside a Create Invitee No Show API. That automates the inverse of what you want: held gets inferred from the absence of a marked no-show rather than proven by attendance. Cal.com offers booking created and cancelled webhooks with a mark-no-show action, and Chili Piper carries completed and no-show status. Neither was verified against primary documentation in the August 2026 pass. Inference is workable at a few hundred sends a month, and it degrades quietly every time nobody marks the no-show.
18.3.5 Pull attendance from the meeting tool, and record meetings actually held.
Your conferencing tool knows who showed up, so let it update your records by machine: subscribe to Zoom's meeting_ended event or the Teams attendance report, write held into the contact record, and store the attendance report's identifier as evidence. Mark no-shows the same way from Calendly's invitee_no_show.created. Never rely on a human to flip HubSpot's meeting outcome, which the research states plainly. The evidence identifier turns your meetings-held count into something a client or an auditor can check.
18.3.6 Stage the build so the version you run today costs nothing.
Stage zero is one file you control, holding the truth and the audit trail, pushing to one sender's block list. A spreadsheet is a fine first version, free, readable, exportable and agent-writable through an API; timestamp every row so it answers when each person opted out. What eventually pushes you off a sheet is needing to prove a record was never altered, not volume at a few hundred messages a month. Concurrency is not the trigger either: the Sheets API append resolves the target row on the server, so two agents appending at once produce two rows, and you only race by reading the last row yourself and writing after it. Stage one adds Attio Free or HubSpot Free as a front end once you run several senders or want an interface; the append-only log stays where it is as the evidence. Stage two is a paid Attio tier past three seats or for sequences and call intelligence, re-checked since Attio moved the figure in July 2026; a native opt-out points at HubSpot, the Search API throttle at caching locally, and several clients away from Smartlead's client-scoped list.
The connecting layer will send a real email unless your configuration stops it.
Model Context Protocol, or MCP, is the standard interface your agent calls a vendor's tool through, and by August 2026 almost every layer of an outbound stack ships one. A census of primary vendor documentation, run on 9 August 2026, asked one question of each server: can it cause an outbound message nobody can recall? The answer was yes for most of the sequencing vendors it checked. Two vendors designed a gate on purpose while everyone else left the trigger exposed. So the rule that nothing fires without a human has to live in configuration you control or in code you wrote yourself.
19.1 Seven official servers can fire a message you cannot take back.
The census, the August 2026 survey of official vendor MCP servers, states it without hedging: "most sequencer vendors do NOT withhold send capability". Apollo, lemlist, Instantly, Smartlead, Salesforge through Agent Frank, Reply.io and Woodpecker each ship an official server that reaches a live send or a sequence enrollment. Enrollment means placing a contact into an automated sequence that then sends on its own. Every one of those actions is visible to the buyer and impossible to withdraw once it fires. An agent you connect inherits that reach immediately, because the capability arrives with the connection rather than with any later decision.
19.1.1 Configure the stops before you connect anything, because a connected agent can send.
Assume any connected agent will send unless your configuration stops it. The census's conclusion carries this whole part: "the burden of not auto-firing sits largely on the operator's MCP-client configuration". The client is whatever runs your agent; the configuration is the permission rules, tool allow list and approval settings written there. Vendors ship the trigger wired up and expose it to whatever connects, so a rule that appears nowhere in that configuration is an intention rather than a control.
19.1.2 Keep every send-capable server out of anything that runs unattended.
Connect the seven named servers only where a human is present and the approval lives in the client configuration or in deterministic code. Deterministic means ordinary code that behaves the same way on every run. If you cannot name the line that enforces that today, the honest reading is that nothing enforces it, because the vendor is not supplying the enforcement.
19.1.3 The stack splits into three layers, and only the third one is irreversible.
Contact data and enrichment servers read and enrich, across every vendor the census checked, and nothing they do reaches a buyer. CRM servers write to your system of record, and some ask for confirmation first, so the worst they do is data damage you can repair. Sending and sequencing servers produce a live send or a sequence enrollment, which the buyer sees and nobody can withdraw. Sort every connection you hold into those three layers before you decide which ones need a gate.
19.1.4 The best-governed layers are the ones carrying no irreversible risk.
Official servers are plentiful at both safe ends of the stack as of August 2026. Apollo, ZoomInfo, Lusha, Crunchbase, Surfe, Wiza, FullEnrich, Common Room and Warmly cover data, and HubSpot, Attio, Salesforce, Pipedrive, Close, Zoho and Dynamics cover CRM. The census's one-line verdict: the maturity is concentrated exactly where the risk is not, so a mature ecosystem tells you nothing about whether your send path is safe.
19.1.5 A confirmation prompt on a recoverable write is guarding the wrong door.
Some CRM servers ask before they write, the only default confirmation anywhere in this stack. A bad CRM write costs a cleanup job, while a bad send costs a buyer and a piece of your domain reputation. So the industry put its prompts where the consequences are smallest. Move your attention to the layer that has none.
19.2 Two vendors design a send gate, and a third puts the send behind a license.
Apollo split the send into two steps deliberately, and Google left the send tool out of Gmail entirely. Microsoft shipped a real send behind a license and an admin registration. Those are the only gate designs the census found as of August 2026, and only two of them put a human between the agent and the buyer.
19.2.1 Apollo split draft from send on purpose, which is the first designed gate in this research.
Apollo's 2026 release notes describe the choice in its own words: "One-off email: AI tools can now compose and send individual emails through MCP. The workflow is split into two deliberate steps, draft then send, so nothing goes out by accident". That is a vendor building the gate on purpose rather than one that fell out of billing, a rate limit or a schedule. It is one product in one census, so read it as proof the design is possible rather than evidence the market is moving.
19.2.2 Google achieved safety by leaving the send tool out of Gmail entirely.
The official Gmail MCP server exposes gmail.create_draft and ships no send tool at all, which the census records as safety by omission. An agent connected to it can research, write and stage a message, after which a human opens the draft and presses send. Nothing in that arrangement depends on your configuration being correct, which makes it the strongest gate in the census and also the least flexible one.
19.2.3 Microsoft exposes a real send, and a license check is not a gate on any message.
The official Outlook MCP server exposes real sendMail and sendDraft tools behind a Copilot license and admin registration in Entra, Microsoft's identity directory. Both checks decide who may connect at all, and neither looks at an individual message. Once the connection exists the agent can send, so the licensing gate defends the boundary of the tenant while leaving every send inside it ungoverned.
19.2.4 The Gmail and Outlook asymmetry decides which mailbox your agent may touch.
The same split appears one layer up, inside Claude's own connectors, where the August 2026 subscription-tiers research recorded both sides. Of Gmail it reports that "The send function is not enabled", so every message goes out by hand. The Microsoft 365 and Outlook connector can send, through outlook_send_email and the underlying Mail.Send permission, once an administrator turns on write tools.
19.2.5 Administrative control over read, write and send exists only on Team and Enterprise.
The same research, the August 2026 subscription-tiers report, records admin-level restrictions on reading, writing and sending as available on Team and Enterprise plans and absent on Pro and Max. So an operator on a personal plan has nowhere to enforce a send restriction above the individual session. That is a second reason to put client-facing work on a commercial plan, alongside the data-handling one. Team, Enterprise and API sit under Commercial Terms and are not used for training by default.
19.2.6 One documented flag forces a human prompt even when permissions are wide open.
Anthropic documents _meta["anthropic/requiresUserInteraction"]: true on an MCP tool, which forces a permission prompt on every call even in acceptEdits, auto and bypassPermissions modes. In dontAsk mode it denies the call outright. That is the strongest single primitive available for this problem, and it belongs on every tool that can send. The census puts the burden on the operator, so plan to set it yourself, on a server or a wrapper you control.
19.2.7 Write your send rule as a deny or an ask, because those survive bypass mode.
Anthropic's permission documentation states that deny rules and explicit ask rules keep working in bypass mode, while allow rules are the one class that stops mattering there. An organization can switch bypass off across the board with permissions.disableBypassPermissionsMode: "disable" in managed settings. Hooks, your own code running at defined points in the agent loop, execute before the mode check and can block a call outright. All three are documented rather than inferred; build on them.
19.2.8 The outbox pattern is sound, and no vendor ships it for you.
An outbox is a staging table: the agent writes a draft, and a separate process holding the only sending credential delivers whatever a human released. Anthropic documents none of that as a primitive, so you assemble it from tool restrictions, permission policy, credential scoping and one human approval step. Claude Managed Agents supports the assembly best: MCP servers there are scoped per agent, Anthropic's own example being that "only the researcher declares the GitHub MCP server, so the coordinator does not have access". Secrets are scoped per session through vaults and substituted at egress, so "the agent never sees the secret value". Runtime bills at $0.08 per session-hour on top of tokens, which the permissions research puts at roughly $58 a month for an always-on agent.
19.2.9 Anthropic's own telemetry says the prompt you are relying on gets approved anyway.
Anthropic's engineering write-up on containing Claude reports that users approved roughly 93% of permission prompts, and concludes that per-turn human approval is "fallible". The controls it calls durable are environmental: egress allowlists, credential isolation from the sandbox, and enforced permission rules. Set requiresUserInteraction on your send tools regardless, because a prompt approved 93% of the time still stops the other seven. Treat it as your second line, with the credential split doing the work that holds.
19.3 The names in this layer will produce a wrong build if you trust them.
Three naming problems in the census can send a competent builder to the wrong server, and one sends the build outside a platform's terms entirely. Each is cheap to check once and expensive to discover later, so check the vendor's own domain against the server you are about to connect, every time.
19.3.1 "Clay" names two unrelated products, and both ship official servers.
One Clay is the go-to-market enrichment platform at clay.com; the other is a personal-relationships CRM with no connection to it. Both publish official MCP servers, so an agent told to connect Clay can arrive at either one. The failure mode is a build that appears to work while quietly doing the wrong job.
19.3.2 The "Apollo MCP Server" you find first probably belongs to a different company.
The enrichment research flags that "Apollo MCP Server" from apollographql.com is GraphQL tooling from an unrelated company. The sales-data product you want is Apollo.io, whose server lives at mcp.apollo.io and is the send-capable server named earlier in this chapter. Check the domain before you connect.
19.3.3 LinkedIn has no official server, and the community ones breach its terms.
The census found no official LinkedIn MCP server as of August 2026. The community servers work by wrapping a browser-session cookie that lets them drive the site as though they were your logged-in browser. That violates LinkedIn's User Agreement section 8.2 on scraping and automation, whose enforcement Chapter XVI documents. Anything advertising LinkedIn automation through MCP sits outside the platform's terms whatever it claims about safety. So the honest build treats LinkedIn as a surface a human works by hand.
19.3.4 Vendor capability arrives as MCP servers, while your own procedures live in skills.
The skills research checked Anthropic's official marketplace and counted roughly 255 to 267 plugins. Apollo, ZoomInfo, Hunter, Lusha and Explorium appear there. Clay, HubSpot, Salesforce, Attio, Outreach, Salesloft, Gong, Instantly, Smartlead and Lemlist do not, and Lemlist publishes skills on its own site instead. The report concludes that most Claude-plus-vendor integrations in go-to-market arrive as MCP servers and connectors rather than as skills. Build to that shape: vendor capability comes through MCP, while your repeatable procedures live in skills, where disable-model-invocation: true blocks automatic invocation and leaves a manual call working.
Spending a dollar to research a contact record that costs cents is a design error.
Two things decide what this build costs: a licence clause that puts unattended work on the metered API, and the shape of your research loop. The loop is where a dollar per contact and three and a half cents per contact both come from. The arithmetic is derived from published August 2026 rates, not measured on a running pipeline, so instrument your own first month before trusting any line of it. Only four mechanisms actually stop spend, and the cap operators reach for first has documented enforcement bugs.
20.1 The licence decides where work runs.
Read the terms before the price list. Anthropic sells the same models as a flat monthly seat and as a metered account billed per token. The seat is far cheaper for deep research, but the terms still send some of that work to the API. Chapter XIV covers the surfaces; this chapter covers what each arrangement costs and permits.
20.1.1 The terms bar automated access except through an API key.
Section 3(7) of Anthropic's Consumer Terms, effective 8 October 2025, bars reaching the Services "through automated or non-human means, whether through a bot, script, or otherwise." The one general exception is access via an Anthropic API Key "or where we otherwise explicitly permit it." The August 2026 licensing research quotes both passages from the terms document. In plain words, a script driving a Claude application is barred; the same script calling the API is what the exception describes.
20.1.2 The test is who starts the work and where it runs, not whether you watch every minute.
Section 3(7) bars driving Anthropic's applications with outside scripts, and that is the whole line. A long multi-agent session you start and steer inside Claude Code or Cowork is the product working as designed, however many agents it spawns and however long it runs, and Cowork's scheduled tasks are the documented unattended exception. What moves work to the metered API: an external script driving the apps, schedules outside the documented surfaces, serving another company's requests through your seat, or anything irreversible running with nobody able to stop it. Anthropic added server-side checks in early 2026 blocking third-party harnesses from subscription credentials, then formalised the position on 4 April 2026, so the clause is enforced rather than merely written.
20.1.3 A seat licenses interactive research and commercial use of its output.
You own what Claude writes for you, and you can sell it. Consumer Terms section 4 assigns Anthropic's rights in Outputs to the user, with nothing conditioning that on personal use. So a consultant may run Research mode by hand and put the result into a client deliverable. What stays constrained is access: a seat does not license scripted use, powering a client's running system, routing anyone else's requests through your credentials, account sharing or resale. Anthropic publishes no safe-harbour examples, so the edges of that list are yours to judge.
20.1.4 Cowork scheduled tasks are the one documented seat-native exception.
The licensing research names them as the single documented case where a seat runs work with nobody watching. They run in the cloud on all paid plans, and qualify because the seat-holder starts and reviews them inside Anthropic's own application. Treat every other scheduled arrangement as API work until Anthropic documents otherwise, the default Chapter XIV applies to the unattended surfaces.
20.1.5 Use the seat for research you do yourself.
Research mode inside a chat subscription carries no marginal per-token cost, per Anthropic's help-centre language; a deep query only draws down the flat usage window faster. The August 2026 economics research priced the same 200-source run on the API at roughly $20.60 on Opus 5 or $8.60 on Sonnet 5 before caching. It put heavier fan-out at $50 to $100, and called Research mode one to two orders of magnitude cheaper per deep query. But it has no official export and no programmatic access, no API endpoint, no MCP surface, no Agent SDK equivalent. So a human moves its output through the copy button, and anything scheduled or chained runs on the API.
20.1.6 Buy a seat for judgement and a hard-capped key for the pipeline.
The economics research priced one operator running 300 researched emails a month plus daily agent work. A subscription alone costs $200 flat on Max 20x, where interactive work lives, because Pro affords roughly one deep research query per five-hour session. The API alone runs $400 to $700 a month with modest optimisation, over $1,000 if the work is Opus-heavy or caching is off. The recommendation is both: Max 20x plus a Console key hard-capped at $50 to $100, roughly $250 to $300 a month with a real ceiling.
20.1.7 The cheaper arrangement is the one that protects you least.
Two protections come with the Commercial Terms and not with a seat: section K.1 has Anthropic defend the customer against claims that its paid use or Outputs infringe intellectual property. Consumer Terms section 11 runs the other way, you indemnify Anthropic, with liability capped at "the greater of six months' fees and $100". Read K.3 first, because it excludes "an alleged violation of trademark based on use of an Output in trade or commerce", exactly what a marketing system does with output. On training, Commercial Terms section B bars Anthropic from training on customer content, while Consumer Terms section 4 permits it "unless you opt out", with carve-outs for feedback and safety-flagged content. Both documents were read at Anthropic's terms pages on 9 August 2026; put anything client-facing on a commercial account for the indemnity alone, a second independent reason each client owns their own Claude account.
20.2 Six moves take a research pass from a dollar to cents.
The economics research modelled a per-email research pass at roughly $0.96 on Sonnet 5. That is an order of magnitude more than the enriched contact record it acts on, which Chapter XV prices at a few cents. Almost none of that dollar is thinking, and five changes to the loop's shape remove most of it, each free at build time and expensive to retrofit.
20.2.1 Almost none of the dollar-a-contact research pass is thinking.
The modelled pass burns roughly 400,000 input tokens against 8,000 of output on Sonnet 5, at rates dated 9 August 2026. A fifty-to-one input ratio describes a system re-reading the same pages, not one reasoning hard. The design test set a ceiling of $1 a contact and aimed far below it, because research should not cost an order of magnitude more than its data.
20.2.2 Turns and raw content drive the bill, and neither carries value.
Twenty fetched pages at about 2,500 tokens each is 50,000 tokens of raw content. And a ten-turn agent loop pays for all of it ten times, because every turn re-sends the whole conversation. Anthropic's own multipliers agree: agents use about 4x the tokens of a chat interaction, multi-agent systems about 15x, and the multiplier measures re-sent context, not intelligence.
20.2.3 Never let the expensive model read a raw page.
Fetch pages with plain code, or with web_fetch capped by max_content_tokens, so nothing unbounded reaches the reasoning step. A single research PDF can run 125,000 tokens, and one uncapped fetch that size can cost more than everything else you spend on the contact. The cap is one parameter and removes the loop's largest tail risk.
20.2.4 Compress every page with a cheap model before anything reasons over it.
Run an extraction pass with Haiku 4.5, at $1 per million input and $5 per million output. It turns roughly 2,500 tokens of page into about 150 tokens of structured fact, so twenty pages reach the expensive model as 3,000 tokens instead of 50,000.
20.2.5 One-shot the judgement, because account research rarely needs a second step.
The compressed facts, the output of the cheap extraction pass, go into a single call, which has no conversation to re-send. Loop only where a step genuinely depends on the one before it, which for account research is uncommon.
20.2.6 Cache the context that is identical for every contact you research.
Your segment definition, offer library, voice profile, output schema and few-shot examples do not change between contacts. Cached reads cost 0.1x base input, a 90% discount, and a five-minute cache write costs 1.25x, so the write pays for itself after one read.
20.2.7 Research the account once and let two or three contacts share it.
Company-level research is identical for everybody at that company, so two or three contacts per account divide the expensive half by two or three. A small per-contact pass adds only what is specific to the person. Nothing about the output gets worse, which is why this move goes in first.
20.2.8 Batch the whole pass, because research is never urgent.
Research happens days before a send, so nothing needs a synchronous response. The Batch API takes 50% off both input and output, and the discount stacks with caching.
20.3 Once you compress the tokens, web search becomes the bill.
The figures below are derived from published rates with stated assumptions, not measured, and that distinction matters more than any decimal. Assume eight pages at 2,500 tokens each, four web searches, a 15,000-token cached context, 800 tokens of account brief, and two contacts per account.
20.3.1 Eight cents per account, three and a half cents per contact.
On those assumptions the Haiku extraction pass costs $0.026, four web searches cost $0.040, and the Sonnet account brief against a cached read costs $0.014. That totals $0.080 an account, or about $0.060 through the Batch API. A per-contact personalisation pass adds $0.005, so two contacts per account works out at roughly $0.035 each, one contact at roughly $0.065.
20.3.2 Fifteen to thirty times cheaper puts research beside the price of data.
Against the naive loop's $0.96, the restructured loop does the same work for one fifteenth to one thirtieth of the money. That puts research in the same order of magnitude as the contact record it acts on.
20.3.3 Web search becomes the dominant cost the moment the tokens are compressed.
At $0.040 of an $0.080 account, search is half the bill once extraction has done its job, so search count replaces tokens as the lever. Budget four searches an account, and write the budget into the tool definition rather than a prompt.
20.3.4 Mark all of this as derived, and instrument your own first month.
None of these figures were measured on a running pipeline; they are arithmetic over published August 2026 rates, weaker evidence than a bill you have paid. Log actual cost per account from your first run, because one month of real numbers replaces every estimate in this chapter.
20.3.5 Sonnet 5's introductory price rises on 1 September 2026.
The rates above are dated 9 August 2026. Sonnet 5's introductory $2 per million input and $10 per million output move to $3 and $15 on 1 September 2026, shifting every Sonnet line by about half again. Date your own cost spreadsheet, and re-run it whenever a price moves.
20.3.6 Write a run's expected cost into the runbook beforehand.
At $0.08 an account, a job touching 100 accounts should cost about $8. If it costs $80, something is fetching uncapped pages or looping where one call would do. The expected figure beside the job makes the anomaly visible within the hour, not at the end of a billing month.
20.4 Four mechanisms stop spend, and everything else only reports it.
Anthropic gives you several ways to watch a bill and exactly four ways to stop one. A control that reports has already let the thing happen: the cost twin of Chapter III's kill switch. Two more facts follow. One environment setting silently moves you from flat billing to metered, and the cap operators reach for first does not reliably hold.
20.4.1 Four mechanisms actually stop spend, and budget alerts are not among them.
The August 2026 economics research found four. Prepaid credits with auto-reload off make your ceiling whatever you loaded, and the Console monthly spend limit returns HTTP 429 once reached. The Agent SDK's max_budget_usd halts the run on error_max_budget_usd, and per-request max_tokens caps a single call. Budget alerts, the Usage API and the Cost API all tell you afterwards. Set the four before your first production run, and file the alerts under reporting, not braking.
20.4.2 Turn caps have documented enforcement bugs, so dollar caps are the brake.
Anthropic's own tracker carries GitHub issues 41143 and 1177, where a subagent declared at ten turns ran between seventy-two and seventy-five. A turn cap is a suggestion, and treating one as a spending control runs an unbounded bill behind a number that looks bounded. Cap the dollars, which is enforced; keep turn caps as a smell test on a loop that should have terminated.
20.4.3 An API key in the environment quietly moves you onto metered billing.
If ANTHROPIC_API_KEY is present in the environment, Claude Code bills that session to the API regardless of the subscription you already pay for. Run /status before you start work, so a metered bill never arrives as a surprise.
20.4.4 Managed Agents bill runtime as well as tokens.
Anthropic charges $0.08 per session-hour of active runtime on top of token spend, roughly $58 a month for an agent that never stops by the permissions research. That is small against an operator's time and large against the cents-per-account arithmetic above. The alternative is building per-agent tool scoping and session-scoped secret handling yourself, which Chapter XIV describes and which costs engineering time instead of an invoice line.
20.4.5 The Agent SDK credit was announced and then withdrawn, so budget without it.
Anthropic announced a monthly programmatic credit in May 2026: $20 on Pro, $100 on Max 5x and $200 on Max 20x. It covered Agent SDK, claude -p and GitHub Actions use and was due 15 June, then Anthropic paused it before that date. The support article "Use the Claude Agent SDK with your Claude plan" still carried the pause on 9 August 2026: "We're pausing the changes to Claude Agent SDK usage described below. For now, nothing has changed: Claude Agent SDK, claude -p, and third-party app usage still draw from your subscription's usage limits." Underlying reports disagreed on this point, so it was checked at the source. Plan on one allowance shared between your own work and anything you schedule.
20.4.6 Three costs that never fall deserve the attention you are giving tokens.
Chapter XII prices two production agents at $257 a month in model inference. Fully burdened, with an identity service and a share of the CRM bill, the figure is $500 to $800. Either number leaves model spend a rounding error beside the headcount it displaced. The lines that never fall with a price cut are enrichment credits, deliverability infrastructure and operator review time. Optimise the loop once, then stop, because your next hour is worth more spent on the list.
Buy one stack whole, because three vendors each assume they own the sequence.
Three complete builds are priced as rungs on one ladder, every price read from a vendor's own page on 9 August 2026 unless another date sits beside it. Each needs re-verifying at checkout, because this category moves weekly. Whole stacks appear rather than a component menu because an integrated suite assumes you will not also run its competitors. A component list invites three products that each expect to own the sequencing layer, the part that decides what happens next and when. Start on the lowest rung that does the job and climb only when a measurement in the last chapter forces you.
21.1 Every price here expires, and every suite tier is sized for large-volume operations.
The figures decay on a timescale of weeks, so each carries its date and source. The tiers behind them were designed for programs running 15 to 1,700 times your volume. So the cheapest paid plan at almost every vendor is already over-provisioned for a few hundred researched sends a month. Read what follows as purchase instructions with expiry dates.
21.1.1 Re-verify every price at checkout, because two of them moved inside this research.
The prices below were checked on 9 August 2026, and two are already moving. Sonnet 5's introductory rates, $2 input and $10 output, rise to $3 and $15 on 1 September 2026. That lifts every Sonnet line by about half again while Opus and Haiku stay put. Attio changed its pricing in July 2026. Treat what follows as the shape of the bill, and read the vendor's page before you enter a card number.
21.1.2 Buy the domain and the mailbox in the next hour, because both start clocks.
Two purchases start clocks that money cannot reset, and together they cost less than a restaurant meal. A lookalike .com runs roughly $11 a year, about 92 cents a month; point it at your real site with a 301 redirect, a free permanent forward. One mailbox on that domain costs $4 a month on Microsoft Exchange Online Plan 1, or $7 on Google Workspace Business Starter. Write SPF, DKIM, DMARC at p=none with a real reporting address, and MX in the same session; those records take one to three days to propagate, the only wait. Setup comes to under $20, the entire sending infrastructure at 300 researched emails a month.
21.1.3 Send from a domain you can afford to lose.
Your aged domain with years of real traffic is the one place every source in this research agrees cold mail must not come from. Google and Microsoft score sender reputation at the domain level, not the mailbox level, so a complaint earned by cold mail lands on your invoices, client threads and newsletter. Three underlying reports state the rule independently and none qualifies it. Age is worth real time, two to four weeks of warmup against four to eight on a fresh domain, on the clock Chapter III sets. That argues for buying the sending domain early, not for sending from the one you cannot replace.
21.1.4 Check DNS, blocklists and one test message, because reputation tools need volume.
Checking a domain you already own protects your real mail, but most reputation tooling shows a sender this size nothing. Google Postmaster Tools reports little below roughly 100 messages a day. Microsoft SNDS keys on the sending IP, which you do not own on Google Workspace or Microsoft 365. A low-volume sender can still check that SPF, DKIM and DMARC resolve, that neither the domain nor its host sits on a public blocklist, and how one test message scores against a seed checker. A vendor selling one domain health score is handing you a derived number, not a measurement.
21.1.5 A suite assumes you will not run its competitors, so take a stack whole.
This is an editorial choice rather than a research finding. An integrated suite assumes it owns the sequence of steps, which is reasonable for the vendor and costly for you. Pull one module out, pair it with a rival's, and two products each wait to be in charge. Month three then finds them disagreeing about what already happened to a contact. Three internally consistent builds follow; take one rather than shopping across all three.
21.1.6 Ignore anything priced per email sent, because volume never constrains you.
The August 2026 research on integrated suites states it plainly. Every suite's email-volume tier runs from 5,000 to 500,000 a month, 15 to 1,700 times the 300 you send. What you actually pay for is the unit cost of a mailbox, the quality of the data, and the number of seats. A pricing page leading with emails per month describes a ceiling you will never approach. So compare on data quality and on what the agent-facing connector can write, a connector being the standard interface your agent calls a vendor's tool through.
21.1.7 Every mailbox-fleet product's minimum is larger than your entire program.
Mailforge and Infraforge sell in ten-mailbox minimums and Maildoso in thirty, against the one or two mailboxes a researched program uses. Dedicated infrastructure and dedicated IPs solve a problem shared Google or Microsoft infrastructure does not have, since the provider owns the address and you cannot be listed individually. Standalone warmup pools cost $15 to $96 an inbox a month, while one practitioner test of nearly all the major tools found none showing meaningful improvement. No provider document proves a penalty either, so read pools as decreasingly effective rather than punished.
21.1.8 Four named products price a program you are not running.
Clay Launch costs $185 a month, checked 7 August 2026, with 2,500 data credits. A 240-contact month consumes 720 to 1,200 of them, putting a resolved contact near $0.77 against $0.12 from the cheapest waterfall, a chain of providers tried in order. Apollo's Organization tier starts at three seats you do not have, and Salesforge's Agent Frank costs $499 a month. Any Unlimited tier charges to remove a ceiling you will never reach. The objection is to the purchase at this volume, not the product: Clay's aggregation across more than 150 providers is real capability you cannot use yet.
21.2 Three complete builds, and you start on the cheapest one that can send.
Each rung is a stack you could buy this afternoon and run. If you want one answer, take the two-app build in the first node and stop reading the section. Rung one is the smallest thing that sends real mail to real people. Rung two is the standard build at a few hundred researched sends a month. And rung three is what the same architecture becomes when a second operator joins, multiplying almost every line and deliberately none of the stores underneath.
21.2.1 The starting build is two SaaS applications, and it needs no code at all.
The shortest path from nothing to sending real mail is one data vendor, one verifier, a domain and a mailbox. Apollo Basic at $49 a month covers four layers: database, sequencing, its own do-not-contact screening, and an official MCP splitting send into a draft step and a send step. FullEnrich Starter at $29 is the line you cannot drop, because Apollo's addresses bounce at 15 to 30% in critic tests and user reports. That sits against the 2% ceiling that costs you the domain. Your domain and one or two mailboxes add $8 to $15, so the total is roughly $90 a month across two SaaS subscriptions with no code written. No suppression mirror is needed since contact record and sender are one product; write the two stop thresholds, bounce and complaint, before your first send, and the scheduled check can wait.
21.2.2 What the two-app build costs you is ownership, and one habit buys it back.
Holding the contact record and the sending in one product puts your do-not-contact list inside a vendor. That trade is reasonable at a few hundred messages a month if made knowingly. The August 2026 suppression research did not re-verify Apollo's do-not-contact behaviour, so confirm it before leaning on it. Buy the ownership back with one habit: on every opt-out, the agent appends a timestamped row carrying person, identifier, reason and source to a spreadsheet you keep. That is one tool call, costs nothing, and is the only record that survives you leaving the vendor.
21.2.3 Rung one's sending layer costs eight to thirteen dollars, and verification completes the rung near forty.
The first rung is one lookalike domain at about a dollar a month, one mailbox at $4 to $7, and four DNS records. It adds a free customer record in Attio or HubSpot and an append-only log kept as a SQLite file or a git-tracked CSV. The August 2026 sending research prices the layer at $8 to $13 a month. The two priced lines sum to $5 to $8, the top being the two-mailbox version. Attio ships its official connector on every tier including free, while HubSpot Free trades that for a native unsubscribe property it enforces itself. Nothing here schedules anything, so a human opens the mailbox and presses send, which at a dozen messages a working day is the review step.
21.2.4 Rung one buys verification and nothing else in the data layer.
A hand-built list of 150 to 200 contacts needs one purchase: a waterfall, a chain of providers tried in order until one answers. FullEnrich Starter, bought as a product and called as a single tool, costs $29 a month for 500 credits and charges only for verified results. The data research puts that at roughly 240 resolved work emails from 300 attempts, near $0.12 each. Pay for this before you own a searchable database at all, because a wrong address costs the domain rather than the credit. Rung one lands between $37 and $42 a month with verification on.
21.2.5 Rung two adds a searchable database and lands near eighty-six dollars a month.
Start from rung one, add Apollo Basic at $49 a month on annual billing plus the same $29 waterfall. The total lands near $86 as of 9 August 2026, with the state layer staying free. Apollo earns its slot as a cheap searchable database and for its connector. It is the only official server in the August 2026 census that fires a genuine one-off send, while the others can only activate a campaign. The data research prices the whole data layer at $78 to $137 depending on the waterfall, so read $86 as the low end of a real range.
21.2.6 Treat every Apollo address as unverified, whatever its label says.
Apollo markets 91% email accuracy, but independent and community measurement puts user-reported accuracy at 65 to 70% and cold-send bounce rates at 20 to 30%. That sits against the roughly 2% hard-bounce ceiling providers re-rate a domain at, a hard bounce being an address that rejects the message outright. Sending an unverified Apollo export from a fresh domain is the fastest documented way to lose the domain. So both rungs carry a verification line even though the database advertises its own. Buy Apollo as a database and a send interface, and run every address through the waterfall first.
21.2.7 A sequencer at thirty-nine to forty-seven dollars buys the agent a send path.
Add Smartlead at $39 or Instantly at $47 a month when you want scheduling, reply detection and an agent-callable send path. That takes the bill to roughly $125, nearly the whole gap between the plain build and the cautious one, the minimal and the fully instrumented versions of this stack. Buy Instantly's sequencer if it suits you, never its done-for-you mailbox package: its help centre states the company keeps domain ownership and administrator access and cannot transfer either. Choose between the two on what the connector can actually write, not on deliverability marketing.
21.2.8 Rung three multiplies operators, and each operator brings one or two domains.
Growth adds people carrying their own agents, their own list and their own one or two domains, with a mailbox or two on each. The rung-one and rung-two lines therefore repeat once per person, and the second domain buys isolation, so one bad domain cannot stop the program. Adding mailboxes behind a single operator switches programs, because one person can check 5 to 10 researched outputs a day, Chapter IX's ceiling. A fleet then carries volume nobody reads, so budget rung three as the rung-two figure times operator count, plus the two lines below.
21.2.9 Every operator writes into one suppression store, or the same buyer hears from you twice.
One suppression store, one audit log and one state store serve the whole operation however many operators you add. A suppression store is the record of who must never be contacted again. Per-person do-not-contact lists are partial copies of a list that only works when there is one. Two of your agents then contact the same buyer in the same week. Getting this right on day one costs nothing, and the record you need later is exactly the one nobody wrote. The only line that grows is seats, which the threshold chapter prices.
21.2.10 Rung three usually needs no new software, and Managed Agents is the optional exception.
More operators mostly means more of what you already run: more domains, more mailboxes, one shared suppression store. The one purchase worth pricing is Claude Managed Agents at $0.08 per session-hour, roughly $58 a month always-on, which splits the work into a reading routine and a sending routine, each agent declaring its own connectors with secrets substituted at the point of egress. It is a real step up in complexity, an always-on runtime to configure and watch, so buy it only when hand-carrying that split across operators keeps failing, and not before. You are paying for scoping rather than a finished control, because the outbox pattern, staging sends for review before release, has no first-party primitive behind it.
21.2.11 The model seat is the largest line on the whole bill.
The workload is 300 researched emails a month plus daily agent work, which Chapter XX prices at $200 flat on Max 20x. Pro affords roughly one deep research query per five-hour session. The metered API alone runs $400 to $700 a month with modest optimisation and passes $1,000 on Opus-heavy or uncached work. The recommended shape is Max 20x plus a hard-capped Console key at $50 to $100, roughly $250 to $300 with a real ceiling. Every tool line in this chapter together comes to less than that, which reverses what most operators expect.
21.2.12 Research costs cents per contact once you stop re-sending pages to the model.
The naive loop was modelled at roughly $0.96 per email on Sonnet 5, almost all of it the same fetched pages re-sent every turn. Compress each page with a cheap model, judge once, cache what never changes, and research the account rather than the person. The figure then falls to about $0.08 per account and $0.035 per contact at two contacts per account, on the arithmetic Chapter XX walks through. The figures are derived, not measured, so instrument your first month. At 300 contacts a month they keep the research line below the sequencer's.
21.2.13 The whole solo build is eleven components, and the first one is you.
The entire stack for one operator is eleven components, and the first one is you, reading and pressing send. Add a domain you own, four DNS records, one or two mailboxes, a searchable contact database, and a waterfall called as a single tool. Add a watch list of a few hundred people you follow, a customer record holding the canonical opt-out, and an append-only log in a file you keep. Finish with a job that mirrors every opt-out into each sender's block list, and two written thresholds with a named person attached. At rung two only two of those carry a monthly invoice; the vendorless five, which almost every published stack list drops, are the domain and its DNS, the watch list, the mirror job, the log and the thresholds. A two-person startup buys the same list on day one; headcount changes only the domain count, the customer-record seats, and how many outputs get checked in a day.
21.3 Climb on a measurement, and re-price the whole thing every quarter.
Four measurements move you between rungs, and each is a number you can already read off your own system. Write all four down before you have results to argue with. Two habits keep the plan honest.
21.3.1 A thousand stable contacts a month reopens the build-your-own question.
Independent practitioners put the break-even for chaining enrichment providers yourself at 500 to 1,000 contacts a month; below it, maintenance and debugging cost more than the fees saved. The second condition is that your ideal customer profile has stopped moving, because rebuilding a chain against a list definition that changes quarterly means paying the build cost repeatedly. Cross both conditions before reopening the question; crossing one is not enough.
21.3.2 Thirty mailboxes or 750 sends a day is where dedicated infrastructure starts paying.
Below that line the per-mailbox economics favour first-party mailboxes from Google or Microsoft, which carry the best baseline reputation; above it the fleet products start to make arithmetic sense. Reaching either figure means a human no longer reads every message, since one person checks 5 to 10 outputs a day, Chapter IX's ceiling. That pace will never fill thirty mailboxes. Read this threshold as a different program, not a bigger version of this one.
21.3.3 A measured hard bounce above two percent buys a verification pass whatever the source.
Two percent is the ceiling mailbox providers re-rate a domain against, so a rate above it is an emergency. The fix is a dedicated verification pass in front of the sequencer, applied to every address whatever the vendor called it. Buy it the day you measure the number, because the domain degrades while you investigate. Chapter XII's kill switch, the rule in the sending code that halts every send when a number crosses a line, shows you the figure before the providers do.
21.3.4 A fourth seat in Attio decides your customer record for you.
More than three seats in Attio moves you to a paid tier or to HubSpot Free, a genuine fork because the two products make opposite trades. Attio has the better identity model and no native do-not-contact attribute, so your agent checks a custom flag you built. HubSpot enforces its own unsubscribe property and charges you in data-model weight and a slower search interface. Attio moved its pricing in July 2026, so re-run the comparison rather than remembering it.
21.3.5 Re-price the whole stack every quarter, because the connectors change monthly.
What each vendor's connector can read, write and authenticate against shifts month to month. So a build-versus-buy decision from one quarter describes a stack that no longer exists in the next. Set a standing quarterly review and check three things: whether the connector still writes what your build assumes, and whether a cheaper waterfall has moved the per-resolved-contact figure. The third check is whether any line renewed above its first-year rate. The review takes an hour and is the only defence against a stack that decays unnoticed.
21.3.6 Send from week one on one or two mailboxes; only a fleet waits for warmup.
A small build, one domain carrying one or two mailboxes, sends real researched mail from week one, ramping 5 to 10 a day toward 15 with no standalone warmup pool, which is what the August 2026 infrastructure research recommends. A fleet, many domains sending at volume, spends four to eight weeks on artificial warmup before any live send, which is where the oft-cited 110-to-130-day clock, warmup plus ramp to full volume, comes from. Buying the domain and setting its DNS today is right under both, because a domain's age is the one input money cannot buy back.
Do these thirty-seven things, in this order.
This chapter is the thing you do, and the order carries as much weight as the contents. It mixes three kinds of item on purpose, the first being things you buy or write once, which Chapter XXI prices. The second is habits with a cadence and a number, so the kill switch is two written thresholds and a named person, not a subscription. The third is two publishing items that wait on nothing, run on a nine-month clock, start on day one and get judged last. Three items run on clocks money cannot reset, the sending domains, the publishing habit and the suppression store, so all three sit near the top of Phase 1. Each item names the part that argues it, which is where the evidence sits when a line surprises you.
Phase 1. Do all of this before anything sends.
These commitments must exist before your first message leaves. Some are clocks you cannot hurry; the rest are controls far cheaper to build now than to retrofit. None shows a result in month one, which is exactly why they get skipped. Chapter XII argues the whole phase: what decides the program is decided in week one.
01 Set the country filter to the US, and know which risk actually bites.
Do both before you buy a single tool, because this guide covers US senders writing to US recipients. Put the country filter in the list-building step, then rank what can hurt you, worst first. Mailbox providers come first, because being blocked kills the channel with no appeal, then US private class actions, where money is already moving. State disclosure and privacy statutes follow; a headline regulatory fine ranks last, since essentially no record exists of one landing on an operator this size. Chapter X sets that order, and Chapter XII puts it in week one.
02 Buy the sending domains today, and warm them only if you are standing up a fleet of them.
Buy the domains today under either build, because their age is the one clock money cannot reset. On one or two mailboxes there is nothing to warm: wait the one to three days DNS takes to propagate. Then start real researched mail at five to ten a day and ramp, because that sending is the warmup. Standalone warming is for fleets only, meaning several domains and more than one operator. It runs two to four weeks on a domain you own and four to eight on one you just bought, before real volume. Chapter III starts the clock at purchase, which is why fleet programs declared dead in month three had only been sending for eight weeks.
03 Set up one or two mailboxes and authenticate them the same week.
Do this immediately after the domain: no mail should flow before the records are right, and you cannot warm a domain with no mailbox. SPF, DKIM and DMARC prove your mail really came from your domain. Google, Microsoft and Yahoo all check them under the February 2024 Google and Yahoo rules Chapter I dates. Setup takes a morning, which Chapter VII calls a better hour than polishing copy. Passing all three gets you considered, not delivered; what recipients do with your mail decides the rest.
04 Publish your first piece under a real human byline on the day you buy the domains.
Start publishing content on day one, knowing it takes time to build a flywheel and an owned audience that pays dividends over years. Do not let working on outbound stop you from building momentum for inbound, because nothing downstream waits on publishing, and that is exactly how it starts never. Run it beside the build and judge it at month nine on its own clock. Post three to five times a week; the cost runs from an hour each morning down to fifteen minutes a week per publisher, so budget the larger figure. Chapter VI is honest that nobody has isolated how much publishing lifts, so the case rests on the clock rather than the effect size.
05 Buy the four tools you should not build.
Buy a contact data source, an email verifier, a sequencer and mailboxes, then build only the glue between them. Choose the sequencer on what its MCP server, the connector your agent calls it through, can actually write. Some cannot manage a campaign at all, while others create whole sequences and push contacts into them. Every sequencer markets deliverability, whether your mail reaches inboxes at all, so that separates none of them. The write boundary shifts month to month, so check it yourself before wiring a live send path. Chapter VIII draws the own-versus-rent line: own memory and judgement, rent execution.
06 Build the suppression store, the audit log and the state store before you build anything that sends.
Chapter VIII counts thirteen of the fourteen builds that name every tool shipping only the half that finds people and writes to them. They shipped with nothing to record what came back. The suppression store holds who must never be contacted again, so one opt-out has to suppress every identifier you hold: name, company, every address, the profile URL. Model the person separately from the job, because one law attaches the opt-out to the address and the other to the human being. Keep one store shared across every operator, checked on every path a message can leave by. One tool you buy may already hold a do-not-contact list you can adopt as canonical, the trade Chapter XXI prices in the two-app build. So decide that before writing your own. The audit log records what went out and nobody can edit it, and the state store holds what has already happened to each contact. Build all three in week one, because a suppression store built after your first send has already lost the record it exists to keep. Chapter IX prices fragmented suppression as the costliest documented failure in this field, and Chapter X sets the two legal regimes against each other.
07 The model never holds the sending credential, so it cannot send on its own.
Nothing leaves without a person, and this mechanism makes that rule true rather than hopeful. Hand the credential to plain deterministic code, so the suppression check and the approval sit inside the send path rather than in a prompt. The idempotency key also lives there, stopping a retried job emailing the same person twice. Set Anthropic's disable-model-invocation: true flag on any skill that can send, a skill being a reusable instruction file the model can call by itself. Chapter X found just two gates anybody chose on purpose: a copy review that decayed into a click, and a manual stop at launch that held. Every other gate fell out of billing, rate limits or scheduling.
08 Read and approve every first touch and every reply before it sends.
Every message the buyer sees is drafted by an agent and approved by a person, no exceptions. Chapter X reports it as the one setting that stayed put while the models improved. And Chapter XII calls it the most consistently confirmed finding in the research base. Default every campaign to paused, and open each with a dry run, a full pass with sending switched off. Then put your stop at the irreversible act the buyer sees, the campaign launch, not whichever tool call looked most alarming.
09 Wire a kill switch that pauses your system below the providers' thresholds.
A kill switch is a rule inside the sending code that halts every send when a number crosses a line. It is two thresholds in a file, a scheduled check, and one named halt owner, and no vendor sells it. Set yours at 2 to 3% bounces or 0.3% complaints, treat 0.3% as a wall, and run against the 0.1% target. Write both thresholds before your first send, let them override every other rule, and test the switch with a dry run. Above those lines the providers are already re-rating your domain; Chapter X sets the complaint ceiling, and Chapter XII puts the switch in week one.
10 Disclose that you use AI in your first message, and do the three other compliance chores.
Roughly fifteen US states have chatbot-disclosure statutes, and one plain line in the first message satisfies all of them, so write it in before your first send. The other three chores: record where every contact came from, and put a real postal address and a working opt-out in every message as CAN-SPAM requires. Then honor an opt-out within ten business days across every identifier you hold. Chapter X prices all four at almost nothing, so there is no defensible version of skipping them.
11 Add the "how did you hear about us" field before your first visitor arrives.
Put it on the form as one free-text box, not a dropdown, and land the answer on a contact record you can query. A field that exists only in a form-submission email cannot be counted. It tells you nothing in month one, which is why it never gets added. By month nine your content, cold sends and referrals run at once, and the self-reported answer alone separates work that paid off from luck. Chapter VI shows it cannot be filled in after the fact, and Chapter XII puts it on the day-one list.
12 Write down the four numbers you will act on and the four you will ignore.
Act on reply, positive reply, meeting held, and pipeline created, because a human decision sits behind each, and record meetings booked beside held, because the booked-to-held gap is itself a diagnostic. A positive reply is one that wants to continue, not one that merely arrived, and two later decisions turn on which you counted. Count a meeting held only when it took place, since about a third of cold-booked meetings never happen. Put opens, clicks, emails sent, and sequence steps completed on a written ignore list. Turn open tracking off in week one, because Apple Mail Privacy Protection records an open whether or not anyone looked. Chapter XI shows why a number you glance at but never record is worse than one you never collected.
Phase 2. Build the list, the offers and the sending setup, then send.
The job in these weeks is to reach a first send without wasting it. Which list you pick changes results by roughly a hundredfold and which offer by two to three times. Everything else moves them by less than your capped volume, 250 to 300 sends a month, can detect. A send is one email leaving your system, and a touch is one attempt to reach one person on any channel. Chapter XII gives this build window weeks two to six, with the verdict on your first segment landing well after it. Each contact absorbs several sends before its sequence ends, so 150 to 200 contacts fills more than a month of your capped volume.
13 Write one segment definition in two to five rules of thumb.
A usable segment points at a specific tension somebody actually feels, which a database filter on headcount and industry cannot express. Chapter IV's worked example turned "all SaaS companies" into "Series B SaaS companies, using Salesforce, 50 to 200 employees". Reply went from 2% to 11% on that one researched campaign. Write your rules down before you pull a single contact, since the list moves meetings per prospect about a hundredfold. Chapter IV builds the segment from outcomes and proves it by sending, not by filtering harder.
14 Pull 150 to 200 contacts against the segment you just wrote.
Nothing gets killed on fewer than 150 to 200 touches or four to six weeks, and Chapter IV sets both floors. A smaller pull cannot produce a decision either way. Hold the message constant while the segment runs, because two variables moving at once produce a result nobody can read. Expect your verdict in month three: the first email lands in week one on a small build, one domain with one or two mailboxes, and week five to nine on a fleet. The four-to-six-week minimum counts from that first email either way.
15 Read fifty rows of your own list by hand before you send to any of them.
Do it on the first list you build and again whenever the segment definition changes. Spend thirty minutes reading and count how many companies you would refuse to sell to. That measures the one variable whose spread runs to roughly a hundredfold, and nothing else in month one is that cheap or that informative. Chapter III places this as the first honest reading you get, inside a window where the statistics can tell you nothing.
16 Run every address through the verifier and drop whatever it flags.
Verifiers disagree far more than their marketing admits, so treat any verdict as a risk score. One live test put the same thousand addresses through several tools and measured real bounce rates from 0.10% to 3.70%, a 37-fold spread. That spread sits against vendor accuracy claims clustered at 97 to 99%, so suppress everything that does not come back cleanly valid. Bounce damages the domain rather than the campaign, so this runs before the first send. Chapter IV carries the test, the spread and the tools.
17 Ask three conference organizers for their attendee list in week two.
Ask for the attendee list itself, or permission to buy it, the two routes operators here actually used. Conference attendee lists carry the highest cold-email reply figures in this research, 15 to 20%, and reach people whose contact details are nowhere on the internet. Weigh that figure carefully: it is Nick Saraev's self-report, and he sells scraping and automation education. Nothing independent checks it, and nobody has measured what share of organizers say yes. The hit rate is genuinely unknown, while the ask costs one email.
18 Build five to ten structurally different offers before you rewrite a single email.
Structurally different means a different promise, a different countable deliverable and a different risk reversal, who carries the loss when the thing does not happen. Run the test as a sacrificial rig on throwaway domains whose volume never touches your real sending identity. Retire the rig the day a winner appears, give each offer its own list, and finish before you touch wording. Chapter V prices both sides: an offer change moves results two to three times, visible at your volume. A wording change needs about 1,469 sends per version, a computed sample size you will never have.
19 Send three or four emails, and let the other channels carry the rest.
The measured evidence points one way: Belkins, across 16.5 million cold emails through 2024, puts the highest single-email reply rate at 8.4%. The same data shows a third email cutting replies by as much as 20%, and complaints tripling from 0.5% to 1.6% by the fourth. Woodpecker's 2024 study puts more than six touches on one contact at negative return with deliverability damage. Send three or four emails, three to four days apart, each carrying something new, and work hardest on the first, which draws roughly 58% of replies. The eight-to-fourteen-touch cadences you will read about are multichannel counts across email, LinkedIn and phone from 30MPC and Jason Bay, not a licence to send fourteen emails. One outbound researcher calls rising average touches a counting artifact, not a target; finish with a hard stop and a 30-day cool-down.
20 Plan your first email for week one.
The first researched email goes out as soon as DNS propagates and your list has been read by hand, because on one or two mailboxes the real sending is the ramp. Do not cut either gate short, since an unread list costs more than the days it saves. If you are running a warmed fleet instead, several domains under multiple operators, your first email waits on the warmup clock, week five to nine, and Chapter III holds stakeholders to the clock that matches the build.
Phase 3. Settle into a week one person can actually hold.
Once the first message goes out, the job becomes holding a week you can repeat. Chapter III fixes the two units below. Every capacity figure counts one operator, the named human who checks the work, and every volume figure counts one campaign, one list run against one offer. When you need more output, add a second operator carrying their own agents, list and offers.
21 Hold one campaign to 250 to 300 sends a month.
That is roughly a dozen sends on a working day across two or three mailboxes, well under the 20 to 30 a day per warmed mailbox that deliverability practice treats as safe. The cap is your own choice about depth, and because agents make exceeding it trivial, put it in the sending code as a guardrail. Chapter XII is candid that this low-volume design may be the wrong bet, because the best placement data reversed in three quarters. It now shows senders at 1,000 to 10,000 a month losing ground while the largest gain. Read it every quarter, and move weight to LinkedIn, video and inbound-led work if low-volume senders fall for two more quarters. Chapter III derives the return, about two meetings booked a month and one or two held.
22 Check only the five to ten outputs one person can verify in a day.
That is how many AI-produced items a human can read, check and act on; a plan that needs twenty gets twenty approvals with the reading skipped. Chapter IX sources the number from Nick Saraev counting his own days on camera in July 2026, one practitioner's practice rather than a survey. It still caps everything downstream, because the limit is the person checking and adding agents does not move it.
23 Add at most one new agent per operator per month.
Add an agent only for a step you have done by hand and stopped changing, because the manual version is the specification. Each new agent costs roughly two weeks of onboarding, during which the agents you already run get worse, because the attention that maintained them moved. Four to six agents doing real work daily is one person's capacity, which Chapter IX budgets as direct reports rather than installed software. Past that, adding one means retiring one.
24 Count only the agents whose work you checked today.
An agent counts when a person verifies its output and acts on it the same day. Personio, the HR software company, built 400 internal assistants and reports that ten carry about 80% of the value, so the other 390 do real work for nobody. Idle agents cost nothing to keep and still eat review attention, so audit the list every quarter and retire whatever nobody checks. Chapter IX shows why a headline agent count measures nothing.
25 Read your watch list before you draft anything, every working morning.
Follow a few hundred people rather than connecting with them, and work your saved searches and the feed by hand. Read for the four trigger types in Chapter XVI. Reply helpfully in public before you write privately. Budget 45 to 75 minutes, an estimate nobody has verified, because no published measurement exists. This is the one action here whose compliant implementation is a person rather than a product, because no permitted software reaches the replies where the signal sits.
26 Answer the first positive reply yourself, inside the business day.
Propose two specific times and do the scheduling yourself, because a bare booking link hands the work back at the moment you finally have their attention. About a third of cold-booked meetings never happen, the whole distance between booked and held. So book inside three days and send three reminders with something to read beforehand. Chapter III derives both counts, and Chapter VI sends a raised hand to a person rather than a link.
27 Sort every reply by reason rather than counting them.
At the researched band, a few hundred sends a month, you will get roughly 8 to 18 replies a month. That is far too few to test anything and exactly the right size to read by hand. File each under wrong person, no budget, no pain, bad timing, already have a solution, or "what is this?". Chapter XI calls this the highest-value manual work at a few hundred sends, because it tells you what to change next and the statistics never will.
28 Cap the AI rewrite chain at one step between human checks.
A hop is one step where a model writes text another model rewrites; past the first, output reads as templated because no human judgement entered. Hops pile up quietly as the estate grows, so recount them every time you add or change an agent. Two separate vendor benchmarks rank AI-drafted, human-finished mail above pure human and pure AI writing. So let the model fill the variable slots while you keep the body, the ask and the final read. Chapter VII prices what sameness costs at the spam filter, which catches sameness rather than machine authorship.
Phase 4. Judge the program on dates you set in advance.
Week-six panic and month-twelve nursing are the same failure in different clothes, and a date committed to while calm is the only defence against either. Write all of these down before you have results to argue with, which is where Chapter XII puts them. None is a gate, because a gate in this guide means a human approving a message before it sends, nothing else.
29 Write both kill rules down, because they judge two different things.
The segment rule kills a segment below roughly 4% positive reply after 150 clean touches, clean meaning delivered mail to verified addresses at companies that passed your disqualifiers. The version rule kills a sequence still under 1% reply after 200 sends, and condemns that one pairing of list and offer rather than outbound itself. So change one of the two and run again. Chapter IV holds the pair apart, positive replies against a segment versus all replies against a version, and settles the tiebreak: the segment verdict wins.
30 At day thirty, read the wrong-person rate.
Sort your first twelve to twenty-four replies into the fixed reply-reason buckets, wrong person first, and look at one number: how many landed on the wrong person. That separates a list problem from an offer problem, which have opposite fixes. The sample proves nothing statistically and still tells you a great deal about how you are being understood. Chapter III places this as the day-30 reading worth having.
31 At day sixty, count meetings held and read zero as a verdict on the list.
Sixty days at this volume, 250 to 300 sends a month, buys about four or five bookings and roughly three conversations that happened. Zero is too large a gap to argue at the edges, so repair the list, not the wording. Rewriting the email on this date is the standard mistake and costs another sixty days. Chapter III derives both counts from vendor bands rather than anyone's dashboard, so treat them as the shape of the floor.
32 At day ninety, write a diagnosis instead of comparing a rate.
Ninety days buys roughly seven meetings booked and four or five held, and no rate computed on four or five events is a verdict. The honest output is a written judgement on three questions: was the list right, was the offer right, and did the mail arrive? Chapter III replaces the 90-day clock with the 120-day one, and that is the argument for whoever is judging the program.
33 At month nine, judge the topic and the audience, and leave the publishing schedule alone.
You cannot judge content before month nine, because sourced pipeline from publishing takes 9 to 12 weeks and the effect on acquisition cost takes 12 to 18 months. When the review comes, change only what you write about and who you write it for. Killing the habit is almost never right, because it gives up the one asset you cannot restart at the same price. Chapter VI carries the three clocks and the warning that every source behind them is self-reported.
Phase 5. Keep the system from fooling you.
These four exist because the system will mislead you quietly and in your own favour. One agent reads the stored results and writes the summary you act on, and it needs tighter rules than the one drafting the email. A bad send reaches one prospect and stops, while a confident wrong number reaches the plan, the board slide and next quarter's budget.
34 Set four guardrails on the reporting agent before it states a rate.
First, forbid stating a rate until at least 100 sends and 10 positive events sit behind it. Second, require a range on every comparison, plus a sentence on whether the ranges overlap. Third, ban causal verbs such as drove, caused and improved. Fourth, make it print the query it ran and the row counts returned, so a human can check the sends behind the number. Chapter XI sources all four, and explains why reporting is the riskier of the two jobs.
35 Let "not enough data to conclude" be an answer your reporting agent can give.
A model with no honest way to say nothing will reach for a trend, because generating text is the only thing it does. Write the phrase into its permitted vocabulary, count how often it comes back, and treat a month that returns it as a month that reported correctly. Chapter XI separates "we measured nothing" from "we measured no effect", and at these volumes the first is almost always the true one.
36 Never let an agent rewrite its own prompts, and ignore the winner your sequencer crowns.
One documented build fed its reply rate into its own optimiser and taught itself to farm replies nobody was paying for. Chapter XI reads that as automating the failure rather than fixing it, so have the agent propose a change as a diff a person reads and merges. Your sequencer will also crown a winning version from day one. But at 250 sends a measured 5% reply rate covers everything from 2.9% to 8.5%, so the crown means nothing. Chapter XI does that arithmetic, and Chapter X keeps prompt self-editing on the short list of tasks that never left a human.
37 Re-price build against buy every quarter, and re-date anything that describes a model.
Chapter II prices any hand-built workaround for a model weakness as a wasting asset: one release deleted three workarounds a vendor's own engineers had built. Agent-facing connectors change monthly in what they can read, write and authenticate against, so a January decision describes a stack gone by April. Sort your notes into two piles today and put a review date on one. What a model can do expires with the next model generation, while what a human can verify in a day keeps holding. Put the build-versus-buy question itself on the quarterly calendar.
At day 120 a working system looks unremarkable, and that is the point. Your domains are warm, one segment and one offer run at 250 to 300 sends a month, and you read every message before it left. Behind that sit a written day-90 diagnosis, a few meetings that actually happened, and four months of published work under your own name.
How this guide was made, and how far to trust it.
A guide that grades everyone else's evidence should show its own work. This appendix explains the process that produced these pages, then the grading rules the numbers live under.
A.1 The guide was built the same way it tells you to build.
The system this guide describes is agents doing the heavy reading while humans make the judgment calls, with the checks written into code. That is also how the guide itself was produced, so the method doubles as a working example.
A.1.1 Sixty-five research reports were commissioned under a skeptic contract, in two waves.
Every research prompt forced the same discipline: name who makes each claim and what they sell, date every figure, and say how many cases each number came out of. The contract earned its keep, because nearly every headline reply rate practitioners quote dissolved once a report chased it back to a source.
A.1.2 1,821 recordings were cut to 209, then read into about 1,190 graded claims.
Candidate videos were triaged by title, then re-scored against a rubric for operator-level substance drawn from the research corpus itself, leaving 209. Extraction ran in two waves, 76 recordings and then 133. Every second-wave finding had to declare whether it confirmed, contradicted or extended the first wave, so the second pass added instead of repeating. With the report extracts, that reading produced about 1,190 claims, each carrying a grade, a date and a source, before a page of this guide was written.
A.1.3 Every named claim was then re-checked against its source, one at a time.
Once practitioners were named, every sentence naming one became checkable, so 604 attributions were verified individually. Of those, 516 held, 62 were wrong and corrected, and 16 could not be supported and were cut back to descriptions. And 10 were flagged as unverifiable and kept only with that label.
A.1.4 Production hit the same failure modes this guide warns about.
Three times during synthesis an AI layer widened a real figure into a plausible range or invented a specific nobody had measured. Each read so naturally it survived several reviews before a check caught it. Once, a merge step silently dropped most of the document while reporting success. Those failures are why every count, closing and character in the final file is verified by code rather than anyone's assurance.
A.1.5 Built with Claude Code, as a working demonstration.
The research fan-out, the extraction, the drafting and the checks all ran as Claude Code agents over plain files. Project state lived in versioned documents rather than anyone's memory. The judgment calls, the corrections and the final read of every line stayed human. That split, machine volume under human judgment with verification in code, is the architecture the twenty-two chapters before this one recommend.
A.2 Numbers were graded before they were used.
Every figure carries a grade, a date and a named source, and figures that could not carry all three were cut. That rule is why 51 strategy reports, 14 implementation reports and 209 recorded practitioner videos yielded far fewer usable facts than their page count suggests.
A.2.1 Ask who measured the claim, and out of how many cases.
Two questions settle a claim's grade: is the source independent of the thing being measured, and did they say how many cases the number came out of? A-grade means independent with that count published, and C-grade means self-reported with nothing to check it against. D-grade means paid for by a vendor selling the thing measured, which describes most benchmarks in this field.
A.2.2 Author's notes are marked, and you cannot check them.
A few figures come from my own client work rather than the research base, and they appear as author's notes with the client never named. That is because some things are only visible from inside a deployment and nobody publishes them. They break this guide's own rule that every number trace to a named source in one step, so they are marked every time instead of blended in. Read them knowing three things: I sell the services these observations point toward, the same conflict disclosed for every practitioner cited here, and the clients cannot be named. You also have no way to verify any of it, which is why no author's note is ever the only support for a claim.
A.2.3 Trust the source who admits a loss.
Sometimes a source says something that costs them money, such as a vendor admitting a limit or a buyer publishing that they shut a program down. There the incentive to exaggerate runs backwards. A dated "we stopped doing X and here is why" outranks every reply rate anyone in this field has published.
A.2.4 Trace a statistic to a named source in one step, or delete it.
If a statistic cannot be traced to a named study or dataset in one step, it is deleted. Several of the most-repeated numbers in outbound survive here only as debunks, because the trail ends in a blog post citing another blog post.
A.2.5 Our own research failures are printed here.
One commissioned report had to be re-run and another was delivered twice. We redid the arithmetic in the statistics report ourselves, confirmed 12 of its 16 sample-size cells, and set the other four aside rather than quoting them.
A.2.6 A missing proof is itself a result.
Vendors sell buying signals that are supposed to predict when a company will purchase, and nowhere in this research base is there independent proof that they do.
A.2.7 Count independent sources, not voices.
Six named practitioners saying the same thing is not six pieces of evidence, so look at who they are. Jordan Crawford of Blueprint has advised and invested in Clay, the go-to-market data platform, since January 2021. Eric Nowoslawski of Growth Engine X is Clay's largest user by enrichment volume, and Crawford publicly promotes Nowoslawski's method as the better answer wherever his own breaks.
A.3 The ideas in this guide belong to the people who did the field work.
This guide stands on the shoulders of the practitioners named below. They ran the campaigns, published their numbers, taught their methods and, in the most valuable cases, said in public what stopped working and why. My contribution was to collect their work in one place, check it, grade it and arrange it into a system. Where the guide disagrees with someone, it does so in the open, with the evidence beside it. The interpretations, the rankings and any errors that survived the checking are mine alone.
A.3.1 The names, in alphabetical order.
Each name links to where their work lives, using only addresses that could be verified. The short descriptions say what each person does, because that is who they are, and it is also how this guide weighs what anyone says.
Will Allred, co-founder of Lavender and publisher of the 231,818-email benchmark study.
Kareem Amin, co-founder and CEO of Clay.
Varun Anand, co-founder and COO of Clay.
Jason Bay, founder of Outbound Squad.
Christian Bonnier, co-founder of ListKit and originator of the permission-first video playbook.
Josh Braun, teacher of the honest first touch.
Jaspar Carmichael-Jack, founder and CEO of Artisan.
Kellen Casebeer, founder of The Deal Lab.
Jordan Crawford, founder of Blueprint and originator of the pain-first list doctrine.
Craig Elias, creator of trigger-event selling and co-author of the 2010 book the signal category rebuilt.
Armand Farrokh and Nick Cegelski, hosts of 30 Minutes to President's Club and carriers of the Gong evidence base.
Alex Hormozi, whose offer vocabulary the offer chapters borrow.
Morgan J. Ingram, the deepest teacher of video prospecting in this research base.
Christian Kletzl, co-founder and CEO of UserGems.
Gene Lee, co-founder of Ramp, whose candor about shutting his own outbound automation team taught more than most successes.
Jason Lemkin, founder of SaaStr, who published his own disappointing AI test.
Amelia Lerutte, chief AI officer at SaaStr and the most detailed operator voice on running agents in production. She published as Amelia Ibarra until 2026.
Michel Lieben and Alex Vacca, founders of ColdIQ and publishers of disclosed-number vendor tests.
Samantha McKenna, founder of #samsales Consulting.
Kyle Norton, CRO of Owner.com and host of The Revenue Leadership Podcast.
Eric Nowoslawski, founder of Growth Engine X.
Jeff Pedowitz, author of the 2026 attribution manifesto.
Kyle Poyar, writer of Growth Unhinged and the nearest thing this field has to a neutral referee.
Adam Robinson, founder of RB2B and the most valuable disclosure source in this research base.
Nick Saraev, automation educator and the largest single voice in this research base.
Michael Saruggia, GTM engineering advisor.
Nadav Shanun, co-founder and CEO of Geodo, who published the honest infrastructure math.
Tibor Shanto, co-creator of trigger-event selling.
Florin Tatulea, writer of Prospecting from the Trenches.
Marina Temkin, the TechCrunch reporter whose 11x investigation anchors the autonomy argument.
Kyle Vamvouris, founder of Vouris and the most sober voice in the video literature.
Richard van der Blom, author of the annual LinkedIn Algorithm Insights report.
Chris Walker, demand-generation strategist of Refine Labs and Passetto.
Beyond these names, the wider research base drew on scores of other practitioners. Their talks, posts and course material shaped this guide even where no single claim of theirs appears in it. The register behind this list records what each person sells alongside what they taught, and that pairing is respect rather than suspicion.
A.3.2 The measurement came from organisations, and they are owed their own list.
The names above taught methods; the numbers in this guide came mostly from organisations that published a dataset with a denominator attached. Belkins (16.5 million cold emails through 2024) supplies the sequence-length evidence and the complaint curve, and Woodpecker's 2024 study puts more than six touches on one contact at negative return. Gong's call and email corpora sit behind most credible cadence claims, and The Bridge Group supplies a decade of SDR structure benchmarks nobody selling software produced. Lavender ran the 231,818-email study, 6sense surveyed four thousand buyers, and RAIN Group and Backlinko published touch-count and response-rate studies with their samples stated. RevenueHero and Ziellab split no-show rates by source, while ColdIQ and Cleanlist ran the vendor accuracy benchmarks that let this guide contradict vendor marketing. Prospeo, Anymail Finder, LeadMagic and FullEnrich published verification methods precisely enough to compare, and the guide notes wherever a figure comes from a vendor selling something adjacent to what it measured.
A.3.3 The build half was written from documents rather than from people.
Aug 2026 edition v1. Published at productiveai.com. Every number in this guide carries a grade, a date and a named source. Back to top.