Recap
- Demand metrics you can trace to a specific message and reply. Vanity ARR claims with no audit trail are the category's signature problem.
- Make deliverability controls (SPF, DKIM, DMARC, and an automatic stop on bounces and complaints) a hard requirement, not a nice-to-have.
- Ask exactly where the contact data comes from, and ask to see how the tool defines and enforces your ICP before it sends a single email.
- Get references in your size and motion, and ask them what broke. A vendor that cannot produce one is telling you something.
Before you buy an AI outbound tool, demand six things: traceable real metrics, hard deliverability controls, honest data sourcing, real ICP discipline, visible guardrails, and references you can call. The AI SDR category has a trust problem: splashy revenue claims that fall apart under inspection and tools that quietly torch sender reputation. This checklist is built to catch exactly those failures during the demo, not three months after you sign.
How do I tell real metrics from vanity claims?
Click into a number until it stops being a number. A real metric traces from a headline figure down to the specific message, the send, the reply, and the meeting that produced it. A vanity claim stops at the dashboard. If a vendor shows you pipeline influenced or ARR attributed but cannot let you open the actual thread behind one of those dollars, you are looking at a story, not a measurement.
The category is full of stories. Plenty of these tools inflate their own results, and more than one startup has overstated how much of its work is actually automated. You do not need to litigate any single company. You just need a tool whose claims survive a click.
What deliverability controls are non-negotiable?
Authentication and an automatic kill switch. Since the 2024 Google and Yahoo bulk sender requirements, mail without SPF, DKIM, and DMARC set up correctly gets throttled or dropped, and a high complaint rate gets your domain flagged fast. The tool must authenticate every send and must stop on its own when bounces or complaints cross a threshold. If a human has to notice and pull the plug, the damage is already done.
One more thing to watch. Open rates are a weak success metric now, because Apple Mail Privacy Protection pre-loads tracking pixels and inflates the number for a large share of recipients. A vendor that still leads with open rate is either behind or hoping you are.
Where does the contact data actually come from?
Ask the question plainly and listen for a plain answer. Good vendors will tell you which providers and which public sources feed their contacts, how fresh the records are, and how they handle suppression and opt-outs. Vague answers about a proprietary database with no detail usually mean scraped or resold data of unknown age, which ages into bounces and complaints, which brings you right back to the deliverability problem.
You are also buying compliance risk along with the data. If the vendor cannot explain sourcing, they cannot help you defend it later.
Does it enforce ICP discipline or just spray and pray?
Make the tool show you its targeting before it sends anything. Discipline looks like a defined buyer thesis, filters you can inspect, and ideally a trigger (a new hire, a funding round, a tech change) that explains why this person, this week. Spray and pray looks like a giant list and a volume slider. Volume without fit is the single fastest way to burn a domain, and it is what gives the whole category its reputation.
| Signal | Spray and pray | ICP discipline |
|---|---|---|
| List size | As large as possible | As tight as fit allows |
| Why this person now | No answer | A specific trigger or signal |
| Targeting visibility | Hidden behind a slider | Filters you can inspect and edit |
| Reply quality | Mostly unsubscribes | Conversations worth a human |
| Domain risk | High and rising | Managed by design |
Can I see the guardrails and the off switch?
Autonomy is fine. Autonomy you cannot see is not. Demand sending caps, suppression lists, a clear off switch, and a log of every decision the tool made and why. You should be able to set the constraints, walk away, and still reconstruct what happened while you were gone. The goal is to be out of the per-message loop without being out of the picture.
This is the line between a tool that runs your outbound and a tool that runs off with it. The autonomous operators worth buying, the category LaunchSurface sits in, treat your guardrails as the contract, not as suggestions. Ask to see where you set the limits and where you can read the audit trail. If the answer is a shrug, keep looking.
How do I check references the right way?
Talk to a customer who looks like you and ask what broke. Same company size, same motion, same channel. Then ask two questions: what went wrong, and what they would change. Real users always have a rough edge to name. A reference who can only gush, or a vendor who cannot produce one in your segment at all, is a signal in itself.
Here is a compact version of the whole checklist to take into your next demo.
| Demand | The question to ask | The failure it catches |
|---|---|---|
| Traceable metrics | Can I click from this number to the exact reply behind it? | Inflated revenue claims |
| Deliverability | Show me SPF, DKIM, DMARC, and the auto-stop on complaints. | Burned domains |
| Data sourcing | Exactly where do these contacts come from, and how fresh are they? | Scraped, stale, risky lists |
| ICP discipline | Why this person, this week? | Spray and pray |
| Guardrails | Where do I set caps and read the decision log? | Runaway automation |
| References | Who like me uses this, and what broke for them? | Fabricated proof |
None of this requires you to be a deliverability expert or a data broker. It requires you to keep clicking until the story either holds up or falls apart.
