Every AI SDR demo is impressive, and that is the problem. The demo leads were chosen, the training material was prepared, and the output you are shown is the good version.
Here is how to evaluate one on something more informative than a sales call.
The five capabilities that decide the outcome
1. What the agent is trained on
This is the single biggest predictor and the least discussed on pricing pages.
Out of the box, a model knows how to write sales email in general. It does not know what you sell, what you charge, which objection kills your deals, or the sentence that makes your best customers nod.
Ask exactly what the agent can be trained on:
- Can it read your website and actually retain it?
- Can you upload documents: service descriptions, pricing logic, case notes, your best-performing emails?
- Can you give it your objection handling in your own words?
- Does that training persist across campaigns, or is it prompt text you paste each time?
An agent with no specific training produces grammatically perfect mail that could have come from any vendor in your category. That is not automation, it is a faster route to the trash.
2. What context the lead records hold
Personalisation is capped by data, not by the model. If your lead record contains a name, a company and an industry, no agent can write anything but a well-phrased guess.
The interesting personalisation comes from things like: the business has no booking system, their reviews complain about response times, they just opened a second location, they are hiring for a role that implies a bottleneck.
So the question is not "how good is the AI." It is "what does this platform know about each lead before the AI starts writing?" A platform that discovers businesses and enriches them has more to work with than one that imports a CSV of names.
3. Draft-and-approve versus autonomous
For first-touch messages, autonomous sending is fine: nobody is on the other side yet.
For replies, insist on drafting. When a prospect answers with a real objection, that is the most valuable moment in the whole sequence and the worst possible place for a model to improvise about your pricing or commit you to something.
One click of approval per reply costs a few minutes a day and eliminates an entire category of failure. Any vendor who treats that as a limitation rather than a design choice is optimising for their pitch rather than your outcome.
4. Whether deliverability comes with it
An AI SDR is a machine for sending more, faster. Without verification, warm-up and sending limits underneath, that is a machine for damaging your domain faster.
Check for: verification at send time, continuous warm-up, per-mailbox health, sensible daily caps, one-click unsubscribe. If the AI product assumes you have solved this elsewhere, budget for solving it elsewhere. Our deliverability checklist sets out what "solved" means.
5. One record across channels
If the agent writes email and something else writes LinkedIn, and they keep separate lists, the same prospect will eventually receive both in the same week with different framing.
Ask directly: one lead record, or two lists that sync?
The four claims that do not predict anything
"Trained on billions of sales emails." Trained on the average of all sales email is a description of the problem, not the solution. What matters is what it knows about your business.
"Fully autonomous." For first touches, fine. For conversations, this is where the stories come from.
"Hyper-personalisation." Personalisation is bounded by lead context. Ask what context exists before you ask how personal the writing is.
"Replaces an SDR for a fraction of the cost." It replaces the hours, which is most of the cost and none of the judgement. Useful framing for a budget conversation, misleading as a description of the job.
The one-week test
This settles it faster than any comparison table.
Day 1. Give each candidate the same five real leads from your actual market. Not their samples. Train each one on the same material, and note how long that takes: a platform where training is fiddly will not get trained properly in real life either.
Day 2. Read the twenty-five messages. Score each one on a single question: could this have been sent to any other business in this industry? If more than half could, the personalisation is cosmetic.
Day 3. Reply to one of your own sequences as a prospect would, with a real objection. Watch what happens. Does the sequence stop? Is the reply drafted or sent? Is the draft any good?
Day 4. Check the deliverability layer. Is warm-up running? Was the address re-verified before sending? Can you see health per mailbox?
Day 5. Look at the second and third follow-ups. Do they reference the first, or restate it? Restating is the tell that you have a template engine with better vocabulary.
Whatever survives that week will survive production.
What to measure once it is live
Reply rate and positive reply rate. Not sends, not opens.
Open rate stopped being a measurement when Apple Mail Privacy Protection began pre-fetching images for a large share of recipients. Send volume is trivially easy to increase and tells you nothing about whether the writing works.
Hold volume constant while you test messaging. If replies per hundred sends is not moving, sending more is the wrong response, and it is the response most teams reach for.
Where Leads Ranger fits
Our AI sales agents train on your own website and documents into a Business Brain that persists across campaigns, write per-lead rather than per-template, and draft replies for one-click approval rather than answering unsupervised.
The reason they sit inside a discovery platform rather than being sold separately is capability two above. The agent writes from what discovery and enrichment already found about that business, which is a different starting point from a CSV of names. Verification, warm-up and per-mailbox health run underneath it in the same product, because in practice campaigns rarely fail at the copy.
Everything is free while we are in pilot. If you want to run the one-week test, five of your real leads is all it takes, and the results will tell you more than this article can.
