7 min readDan Mercer

    What to Demand From an AI Outbound Tool Before You Buy

    Before you buy an AI outbound tool, demand six things: traceable real metrics, hard deliverability controls, honest data sourcing, ICP discipline, clear guardrails, and references you can actually call. Here is the checklist and the failure modes it catches.

    Recap

    • Demand metrics you can trace to a specific message and reply. Vanity ARR claims with no audit trail are the category's signature problem.
    • Make deliverability controls (SPF, DKIM, DMARC, and an automatic stop on bounces and complaints) a hard requirement, not a nice-to-have.
    • Ask exactly where the contact data comes from, and ask to see how the tool defines and enforces your ICP before it sends a single email.
    • Get references in your size and motion, and ask them what broke. A vendor that cannot produce one is telling you something.

    Before you buy an AI outbound tool, demand six things: traceable real metrics, hard deliverability controls, honest data sourcing, real ICP discipline, visible guardrails, and references you can call. The AI SDR category has a trust problem: splashy revenue claims that fall apart under inspection and tools that quietly torch sender reputation. This checklist is built to catch exactly those failures during the demo, not three months after you sign.

    How do I tell real metrics from vanity claims?

    Click into a number until it stops being a number. A real metric traces from a headline figure down to the specific message, the send, the reply, and the meeting that produced it. A vanity claim stops at the dashboard. If a vendor shows you pipeline influenced or ARR attributed but cannot let you open the actual thread behind one of those dollars, you are looking at a story, not a measurement.

    The category is full of stories. Plenty of these tools inflate their own results, and more than one startup has overstated how much of its work is actually automated. You do not need to litigate any single company. You just need a tool whose claims survive a click.

    What deliverability controls are non-negotiable?

    Authentication and an automatic kill switch. Since the 2024 Google and Yahoo bulk sender requirements, mail without SPF, DKIM, and DMARC set up correctly gets throttled or dropped, and a high complaint rate gets your domain flagged fast. The tool must authenticate every send and must stop on its own when bounces or complaints cross a threshold. If a human has to notice and pull the plug, the damage is already done.

    One more thing to watch. Open rates are a weak success metric now, because Apple Mail Privacy Protection pre-loads tracking pixels and inflates the number for a large share of recipients. A vendor that still leads with open rate is either behind or hoping you are.

    Where does the contact data actually come from?

    Ask the question plainly and listen for a plain answer. Good vendors will tell you which providers and which public sources feed their contacts, how fresh the records are, and how they handle suppression and opt-outs. Vague answers about a proprietary database with no detail usually mean scraped or resold data of unknown age, which ages into bounces and complaints, which brings you right back to the deliverability problem.

    You are also buying compliance risk along with the data. If the vendor cannot explain sourcing, they cannot help you defend it later.

    Does it enforce ICP discipline or just spray and pray?

    Make the tool show you its targeting before it sends anything. Discipline looks like a defined buyer thesis, filters you can inspect, and ideally a trigger (a new hire, a funding round, a tech change) that explains why this person, this week. Spray and pray looks like a giant list and a volume slider. Volume without fit is the single fastest way to burn a domain, and it is what gives the whole category its reputation.

    SignalSpray and prayICP discipline
    List sizeAs large as possibleAs tight as fit allows
    Why this person nowNo answerA specific trigger or signal
    Targeting visibilityHidden behind a sliderFilters you can inspect and edit
    Reply qualityMostly unsubscribesConversations worth a human
    Domain riskHigh and risingManaged by design

    Can I see the guardrails and the off switch?

    Autonomy is fine. Autonomy you cannot see is not. Demand sending caps, suppression lists, a clear off switch, and a log of every decision the tool made and why. You should be able to set the constraints, walk away, and still reconstruct what happened while you were gone. The goal is to be out of the per-message loop without being out of the picture.

    This is the line between a tool that runs your outbound and a tool that runs off with it. The autonomous operators worth buying, the category LaunchSurface sits in, treat your guardrails as the contract, not as suggestions. Ask to see where you set the limits and where you can read the audit trail. If the answer is a shrug, keep looking.

    How do I check references the right way?

    Talk to a customer who looks like you and ask what broke. Same company size, same motion, same channel. Then ask two questions: what went wrong, and what they would change. Real users always have a rough edge to name. A reference who can only gush, or a vendor who cannot produce one in your segment at all, is a signal in itself.

    Here is a compact version of the whole checklist to take into your next demo.

    DemandThe question to askThe failure it catches
    Traceable metricsCan I click from this number to the exact reply behind it?Inflated revenue claims
    DeliverabilityShow me SPF, DKIM, DMARC, and the auto-stop on complaints.Burned domains
    Data sourcingExactly where do these contacts come from, and how fresh are they?Scraped, stale, risky lists
    ICP disciplineWhy this person, this week?Spray and pray
    GuardrailsWhere do I set caps and read the decision log?Runaway automation
    ReferencesWho like me uses this, and what broke for them?Fabricated proof

    None of this requires you to be a deliverability expert or a data broker. It requires you to keep clicking until the story either holds up or falls apart.

    Frequently asked questions

    What is the single biggest red flag in an AI outbound demo?
    A metric you cannot trace back to a specific message, send, and reply. If the dashboard shows pipeline or revenue but you cannot click into the exact emails that produced it, treat the number as marketing, not measurement.
    Why are reported open rates unreliable now?
    Apple Mail Privacy Protection pre-loads tracking pixels, which inflates open rates for a large share of recipients. Any tool that leans on opens as its headline success metric is measuring a number it cannot trust. Ask about replies and meetings instead.
    What deliverability controls should be non-negotiable?
    SPF, DKIM, and DMARC set up correctly, plus an automatic stop on bounce and complaint thresholds. Since the 2024 Google and Yahoo bulk sender rules, sending without authentication and a low complaint rate is a fast way to burn your domain.
    How do I check a vendor reference without wasting an hour?
    Ask for a customer in your size and motion, then ask that customer two things: what broke, and what they would change. Glowing references that cannot name a single rough edge usually were not real users.
    Is fully autonomous outbound safe to buy yet?
    It can be, if the autonomy comes with guardrails you set and can see: caps, suppression lists, an off switch, and a log of every decision. Autonomy without visibility is the part to avoid, not autonomy itself.

    Dan Mercer writes about outbound and go-to-market at LaunchSurface.

    AI agentsBuyer guideOutbound