How Lead Generation Fits Into an Agent-Native Prospecting Workflow: A 7-Step Checklist for SDR Teams

2026-09-30 · Julian Hartwell

Who This Checklist Is For — and the Problem It Solves

If you're an SDR lead, RevOps manager, or agency owner trying to figure out where lead generation actually sits inside an agent-native prospecting workflow, this is for you.

Quick context on me: I manage quality and brand compliance for a mid-market B2B SaaS company. My job is to review every outbound sequence before it ships — roughly 40 to 60 per month. In 2024, I rejected about a third of the first drafts I saw, and the reason was almost never the copy. It was almost always a data gap or a missing verification step upstream.

I've also been on the receiving end of what happens when lead gen gets treated as an afterthought. We sent a batch of enriched contacts that hadn't been verified, bounced 12%, and spent the next six weeks rebuilding sender reputation. That one cost us roughly a quarter of pipeline flow.

Here's the checklist I now run before any agent-native workflow goes live. Seven steps. Each has a checkpoint. If you can't answer the checkpoint cleanly, that step isn't actually done.

Step 1: Spell Out Your ICP Before Agents Do Anything

This sounds obvious. It isn't. Most teams I've seen define their ICP in prose — "we sell to mid-market SaaS companies with 200–2,000 employees" — and then hand it to an agent to interpret. That's where drift starts.

What worked for us was translating the ICP into structured filters the agent can actually evaluate: firmographic (headcount bands, funding stage, tech stack), behavioral (recent job postings, product launch signals), and disqualifiers (existing customers, active opps, named competitors).

We learned this the hard way. "Mid-market" meant one thing to me and another thing entirely to the SDR who'd built the list. I said "mid-market," he heard "anywhere from 150 to 5,000 seats." Discovered it when the first 200 leads came back and a third of them were startups with 12 employees.

Checkpoint: Can you list your disqualifiers with the same precision as your target criteria? If your agent doesn't know what to exclude, it'll include it.

One more thing — I don't think there's a universal right answer for how narrow the ICP should be. It depends on your sales cycle and how much disqualification your AEs can absorb. We went too broad in 2023 and paid for it in noise. Your mileage may vary.

Step 2: Clean the Source Before You Enrich

Enrichment can't fix a broken list. It just makes broken data look more authoritative.

Before you run anything through a waterfall enrichment provider — okkigo, your CRM's built-in tool, or whichever stack you've got — do the boring stuff:

  • Deduplicate on domain + last name, not on email
  • Normalize company names (there are at least four ways to spell "Alphabet Inc.")
  • Strip obvious junk — "info@," "noreply@," role accounts
  • Flag records with no company-domain match

We didn't have a formal pre-enrichment cleaning process until 2023. Cost us when an entire batch of 4,000 contacts got enriched on top of duplicate source data — we paid enrichment credits for records that were already in CRM in some form.

Checkpoint: Does your enrichment log show you its source for each field? If not, you can't tell a verified email from a guessed one.

Step 3: Run Waterfall Enrichment as Layers, Not a Single Pass

This is the step most teams skip. They pick one enrichment vendor and treat its output as ground truth.

Waterfall works differently: you chain providers, and each one fills what the last one missed. The order matters. Generally, the more expensive or specialized provider goes first — because it covers the hardest cases — and the cheaper, broad-coverage providers fill the tail.

We run waterfall through okkigo's enrichment pipeline with three providers chained in that order. Coverage went from 61% to 88% on the same input list. But — and this is important — coverage isn't the goal. Accuracy on the fields you actually use is the goal.

I can only speak to our setup, but we've found adding a fourth provider produced diminishing returns. Stack size has a ceiling measured in relevance, not in credits spent.

Checkpoint: Do you know your coverage rate per field, not just overall match rate? A 90% match rate skewed toward direct-dial mobiles is worse than a 70% match rate that leans toward verified work emails.

Step 4: Verify Emails — Twice, and by Segment

Email verification isn't a single event. It's a process that runs at two points: when the list is built, and again when the sequence is about to send. If you're running continuous outreach, add a rolling re-verification schedule on top.

The reason is simple: emails go stale. People change jobs, companies get acquired, domains expire. A list verified in January can have 8–12% newly invalid addresses by Q2. I don't have a great explanation for why the decay rate varies so much between industries — my suspicion is it's tied to job mobility, but I've never been able to verify that.

Two things to check on your verification setup:

  1. Verification granularity: Are you getting risk-scored results (valid / risky / invalid / unknown), or just a binary pass/fail? The risky bucket is where the interesting decisions happen.
  2. Segment-level thresholds: A CEO-level list needs a stricter bar than a mid-manager list. Same tool, different thresholds.

When we moved to okki go email verification with risk scoring, our hard bounces dropped from about 4.8% to under 1.5%. The compound value of stopping that bounce rate in our outbound pipeline was worth more than the annual cost of the tool — but again, that's our volume and domain setup, not a universal number.

Checkpoint: What's your current bounce rate by segment? If you don't know, you're not verifying — you're just hoping.

Step 5: Wire Intent Data In Before the Agent Scores, Not After

This is the one most agent-native stacks get backwards. Teams run the agent, get a ranked list, then try to layer intent on top as a filter.

Too late. If your agent's scoring criteria don't include intent signals, the ranking it produces is essentially a firmographic list. Useful, but not agent-native.

What we do: intent data — job changes, tech installs, content engagement, buying-committee activity — feeds into the ranking prompt at step one. The agent sees intent as an input, not a post-hoc filter.

The tradeoff is latency. Intent signals can take 24 to 72 hours to propagate. Our solution was a rolling refresh: the agent re-scores on a schedule rather than on demand. That works for us because our sales cycle is 3–6 weeks. If you're selling something with a 5-day decision window, this whole step might not fit your setup.

Checkpoint: Can you trace a specific intent signal from "detected" through "included in ranking" through "produced outreach"? If yes, it's wired in. If no, it's decoration.

Step 6: Keep the Human in the Loop — on the Right Things

"Human-in-the-loop" often gets interpreted as "review every message." That's not scalable, and honestly, that's not what HITL is supposed to mean in an agent-native workflow.

The human review should sit at three specific points:

  1. Prompt and criteria changes — anything that changes what the agent considers qualified
  2. New sequences — the first send of a new play runs reviewed-before-send
  3. Anomaly flags — when the agent's confidence drops below a threshold, or when a signal contradicts a firmographic

Everything else — copy variations, minor timing adjustments, follow-up sequencing — runs without human review. We track rejection rate at the review points as a quality signal. In our case, about 18% of new sequences get sent back for revision. That number is slowly going down, but it shouldn't hit zero. If it's zero, the reviewer isn't actually reviewing.

Checkpoint: Can you name the three review points in your workflow? If the answer is "the human reviews everything," you're not running agent-native — you're running direct mail with an AI typewriter.

Step 7: Measure Pipeline Attribution, Not Reply Rate

Reply rate is a vanity metric in an agent-native workflow. Agents can inflate reply rates by targeting easier segments, sending more volume, or writing more "curiosity-style" openers. None of those correlate reliably with pipeline.

What to measure instead:

  • Qualified meetings booked — meetings that passed your qualification criteria, not just meetings on the calendar
  • Pipeline influenced per 1,000 contacts reached — normalized for volume so sequences are comparable
  • Bounce rate by week — as a leading indicator of list health
  • Cost per qualified meeting — the only "price" number that ultimately matters

On that last point: this is where the "cheaper is better" argument falls apart. We ran an experiment in early 2024 where we sourced leads from a provider at roughly a third the cost of our usual waterfall stack. Reply rates were comparable in week one. By week four, bounce rates had climbed 5x, sender reputation had degraded, and we had to pause the entire sending domain for three weeks.

That cheap-per-thousand sourcing turned into a three-week outreach blackout. The math rarely works the way procurement thinks it does.

Checkpoint: If you remove reply rate from your dashboard, do you still know whether the workflow is performing? If not, you're measuring the wrong thing.

Notes, Caveats, and the Mistakes I'd Flag First

Before you take this checklist and run with it, three things.

On compliance: CAN-SPAM in the US and GDPR legitimate-interest rules in the EU apply to outreach regardless of whether an agent wrote the email or a human did. The "human-in-the-loop" defense gets weaker when the loop is thin. Get your legal or privacy team involved at step one, not step seven. We had to retrofit consent tracking into our pipeline in 2023 because we treated it as a legal afterthought, and that was not a fun quarter.

On domain warm-up: If you're starting a new sending domain, nothing on this checklist compensates for skipping warm-up. Three weeks minimum, ramp volume gradually. I've watched teams hit 90% of this checklist and then get blacklisted in 48 hours because they sent 5,000 emails on day one from a fresh domain.

On the "agent-native" label: I'm somewhat skeptical of how freely the term gets used. A pipeline where an agent drafts emails is not the same as a pipeline where an agent owns qualification, ranking, and sequencing. Both are useful, but they need different QA processes. Don't confuse "agent-assisted" with "agent-native" — the failure modes are different, and review checkpoints belong in different places.

And one honest caveat: everything above is calibrated to our context — a mid-market B2B SaaS company with a 3–6 week sales cycle, a two-person RevOps team, and enough domain reputation to afford a hard-bounce learning curve. If you're a small agency running outreach for 20 clients on shared infrastructure, the calculus changes. If you're an enterprise with a dedicated compliance team, you'll probably add steps I haven't listed.

This worked for us. Your mileage may vary — but the seven checkpoints are the questions I'd ask in any setup, whatever your size.