Seatext library / BotRefund evidence
How to Compare Lead Quality Across Ad Campaigns Using a Baseline
Start by calculating a quality baseline for your account — landing-page sessions per click, contactable leads, verified leads, qualified opportunities, and revenue by campaign. Then normalize each campaign's metrics against that baseline to see...
✓ Built for advertisers who need clear, refund-ready traffic evidence.
To compare lead quality across campaigns, first establish a baseline using your own CRM and analytics data. Measure landing-page sessions per click, contactable leads, verified leads, qualified opportunities, and revenue broken down by campaign, placement, audience, creative, device, geography, and time. Then divide each campaign's metrics by the baseline to get a ratio — campaigns above 1.0 outperform the norm, campaigns below 1.0 underperform. This normalization removes volume bias and lets you compare a $500 test campaign against a $50,000 evergreen campaign on equal footing.
Why a Baseline Matters for Campaign Comparison
Raw lead counts and cost-per-lead figures mislead when campaigns differ in spend, audience, or placement mix. A campaign generating 200 leads at $10 CPL looks better than one generating 50 leads at $25 CPL — until you learn the first campaign yields 2 qualified opportunities and the second yields 15. The baseline converts raw numbers into a common language: performance relative to your account's normal.
Without a baseline, you optimize for volume or cost efficiency while the actual business outcome — qualified pipeline — drifts. The baseline also protects you from overreacting to small samples. A sudden quality dip in one ad set might be noise; a consistent gap across multiple clusters signals a real problem.
Building Your Quality Baseline: What to Measure
Pull data from three sources: ad platform (Meta Ads Manager, Google Ads), website analytics (GA4, server logs), and CRM (Salesforce, HubSpot, Close). Join them on click ID (fbclid, gclid) and timestamp. For each campaign, calculate:
- Click-to-session rate: Landing-page views divided by link clicks. A low rate suggests tracking gaps, slow loads, or non-human clicks.
- Session-to-lead rate: Form completions divided by sessions. Isolates landing-page conversion efficiency.
- Contactable-lead rate: Leads with working phone/email divided by total leads. Filters typos, fake details, and bot submissions.
- Verified-lead rate: Leads where a human confirms interest (call connected, reply received, booking made) divided by contactable leads.
- Qualified-opportunity rate: Verified leads that meet your ICP criteria divided by verified leads.
- Revenue per click: Closed-won revenue attributed to the campaign divided by clicks. The ultimate north star.
Compute these rates for the trailing 90 days (or your sales cycle length) across the whole account. That aggregate is your baseline. Then compute the same rates per campaign, per placement, per audience, per creative, per device, per geo, per landing page, and per week. Each slice becomes a comparison cluster.
The Four-Layer Audit Framework
BotRefund's CRM lead quality audit structures investigation in four layers, each adding evidence before you change targeting or request refunds.
Layer 1: Platform Delivery
Compare reach, link clicks, landing-page views, placements, and spend. A cheap placement isn't a win unless it produces contacts you can reach and qualify. Avoid eliminating an entire audience from a small sample; use enough volume to see a consistent quality pattern.
Layer 2: Landing-Page Evidence
Measure page loads, redirects, consent behavior, form start, form completion, time to completion, and meaningful engagement (scroll depth, field corrections, dwell time). A click-to-session gap can have ordinary explanations — in-app browsers, tracking consent, slow loads, analytics misconfiguration. Investigate those before concluding the gap is bot traffic.
Layer 3: Lead Verification
Record whether an email is deliverable, a phone connects, duplicate details recur, and the prospect confirms interest. Add qualification questions that reveal fit, not just extra fields that make the form longer. For high-value offers, a confirmation step or booking flow can be more valuable than the cheapest raw lead.
Layer 4: Sales Outcome Feedback
Give sales a small, mandatory set of dispositions: verified, contacted, qualified, disqualified, duplicate, invalid details, no response. Feed these dispositions back into the ad platform via offline conversion APIs (Meta CAPI, Google Enhanced Conversions). This teaches the algorithm which leads actually matter.
Normalizing Metrics Across Campaigns
With baseline rates and per-cluster rates in hand, calculate a quality index for each cluster:
Quality Index = (Cluster Rate) / (Baseline Rate)
An index of 1.0 means the cluster performs at the account average. Above 1.0 outperforms; below 1.0 underperforms. Apply this to every rate in the funnel — click-to-session, session-to-lead, contactable-lead, verified-lead, qualified-opportunity, revenue-per-click.
Example (hypothetical): Your baseline verified-lead rate is 12%. Campaign A shows 18% (index 1.5). Campaign B shows 6% (index 0.5). Campaign A delivers 50% more verified leads per contactable lead than average; Campaign B delivers half. Even if Campaign B has lower CPL, its true cost per verified lead is higher.
Plot indices in a heatmap: rows = campaigns, columns = funnel stages. Green cells = outperformance, red = underperformance. This visual makes cross-campaign comparison instant.
Decision Criteria: When to Act on Quality Differences
| Criterion | Threshold | Action |
|---|---|---|
| Statistical significance | Minimum 100 clicks and 30 leads per cluster | Below threshold: flag for monitoring, don't optimize yet |
| Consistency | Index below 0.7 or above 1.3 for 3+ consecutive weeks | Persistent gap: investigate root cause (placement, creative, audience, bot traffic) |
| Funnel depth | Gap appears at verified-lead or qualified-opportunity stage | Deeper gaps matter more — they reflect sales reality, not just form fills |
| Revenue impact | Cluster drives >10% of spend but <5% of revenue | High spend, low return: pause or restructure |
| Bot signals | Fast form completion, identical field structures, placement-level spikes, no page engagement | Run client-side behavioral audit (BotRefund) before changing targeting |
These criteria prevent knee-jerk reactions. A single bad week on a new creative isn't a trend. A placement that consistently delivers unverifiable leads across months is a structural problem.
Common Pitfalls and Limitations
- Treating every bad lead as fraud. A weak campaign attracts real people who aren't ready to buy. Bot traffic leaves repeatable technical patterns — unusually fast form completion, identical field structures, sudden placement-level spikes, conversion events with no meaningful page engagement. Investigate signals before accusing fraud.
- Using industry benchmarks instead of your own baseline. Imperva reported automated traffic represented more than half of web traffic in 2025; that does not mean half of a Meta advertiser's clicks are fraudulent. Treat broad statistics as context, then measure your own sessions and leads.
- Changing campaign settings before preserving attribution. Keep campaign, ad set, creative, placement, click ID, timestamp, URL parameters, CRM record, and any verification result before you change targeting or make a refund request.
- Ignoring click-to-session gaps. A gap can stem from app browsers, consent banners, slow loads, or analytics config. Rule out ordinary causes before assuming invalid traffic.
- Over-segmenting. Slicing by campaign + placement + audience + device + geo + hour creates clusters too small to decide on. Roll up until each cluster clears the minimum-volume threshold.
Practical Scenarios: Applying the Framework
Scenario 1: Audience Expansion Looks Cheap But Converts Poorly
Meta's Advantage+ audience expansion delivers $8 CPL vs. $18 CPL for core audience. Baseline normalized index shows expansion verified-lead rate at 0.4x baseline. True cost per verified lead: expansion $20, core $15. Decision: keep expansion but exclude placements driving the gap (often Audience Network), or add a verification step for expansion leads.
Scenario 2: New Creative Spikes Leads Then Flatlines
A new video creative generates 3x leads in week one. By week three, lead volume normalizes but verified-lead index sits at 0.6. The creative attracted curiosity clicks and bot traffic that triggered conversion events. Decision: pause creative, audit sessions for behavioral anomalies, retrain pixel with verified conversions only.
Scenario 3: Mobile vs. Desktop Quality Divergence
Mobile delivers 60% of leads at 0.8x baseline verified rate. Desktop delivers 40% at 1.4x. Revenue-per-click index: mobile 0.7, desktop 1.6. Decision: bid adjust -20% on mobile, +30% on desktop; add mobile-specific qualification question to filter low-intent taps.
Key Terms and Definitions
- Baseline: Aggregate funnel rates (click-to-session, session-to-lead, contactable-lead, verified-lead, qualified-opportunity, revenue-per-click) calculated across the whole account over a representative period.
- Quality Index: Cluster rate divided by baseline rate. 1.0 = average. >1.0 = outperformance. <1.0 = underperformance.
- Click ID (fbclid, gclid, msclkid): Unique parameter appended to landing-page URLs by ad platforms. Enables joining ad-click data to website sessions and CRM records.
- Pixel Poisoning: Bots triggering conversion events, causing the ad platform's ML to optimize for non-human behavior.
- Offline Conversion API (CAPI): Server-to-server endpoint sending CRM dispositions (qualified, disqualified) back to the ad platform to improve optimization.
- Client-Side Behavioral Audit: JavaScript-based detection of non-human interaction patterns (mouse movement, scroll, timing, form velocity) that server logs miss.
Key Facts from Source Pack
| Fact | Source |
|---|---|
| Start with a quality baseline: landing-page sessions per click, contactable leads, verified leads, qualified opportunities, and revenue by campaign | S6 |
| Quality normally changes by placement, audience, creative, device, geography, landing page, and time | S6 |
| Preserve click identifier, campaign context, timestamp, URL parameters, CRM record, and verification result before changing campaign settings | S6 |
| Four-layer audit: Platform delivery, Landing-page evidence, Lead verification, Sales outcome feedback | S6 |
| Bot traffic leaves repeatable patterns: fast form completion, identical field structures, sudden placement-level spikes, conversion events with no meaningful page engagement | S1 |
| Meta Audience Network historically shows high CTRs and near-instant bounce rates | S4 |
| BotRefund detects non-human traffic with 99% confidence and builds compliance-grade evidence for refund claims | S7 |
| 83% approval rate across client refund claims filed with ad platforms | S7 |
| Industry audits place automated traffic between 9% and 20% of paid clicks | S7 |
| Client-side audits analyze visitor browser behavior; server-side audits only see IP, headers, user-agent | S3 |
Limitations: When This Advice Does Not Apply
- Brand-new accounts with < 500 clicks and < 50 leads — no stable baseline exists yet. Use platform benchmarks cautiously, then build your own.
- Single-campaign accounts — nothing to compare against. Focus on absolute funnel health instead.
- Lead-gen without CRM integration — if sales dispositions aren't recorded, you can't compute verified-lead or qualified-opportunity rates. Fix data plumbing first.
- E-commerce with instant purchase — the funnel compresses; revenue-per-click becomes the primary metric, verified-lead rate is irrelevant.
- Accounts where click IDs are stripped (some redirect tools, certain AMP setups) — attribution breaks, normalization fails. Restore click-ID passthrough before auditing.
FAQ
How long a lookback window should I use for the baseline?
Match your sales cycle. If leads typically close in 45 days, use 90 days of data to capture full funnel outcomes. For longer cycles, use 180 days but weight recent months higher.
What if a campaign has high volume but low quality index?
That's the most dangerous quadrant — it burns budget at scale. Pause or restructure immediately. Audit for bot traffic (check Audience Network placement, behavioral signals) before blaming creative or audience.
Can I use this framework for Google Ads and Meta Ads together?
Yes. Compute separate baselines per platform (different audiences, different fraud vectors), then normalize within each platform. Cross-platform comparison only works at the revenue-per-click level.
How do I handle campaigns with different objectives (leads vs. sales)?
Don't mix objectives in one baseline. Build a lead-gen baseline for lead campaigns, a purchase baseline for sales campaigns. Compare only within objective type.
What's the minimum data needed before I can trust a quality index?
At least 100 clicks and 30 leads per cluster. Below that, the index is noise. Flag the cluster for monitoring and revisit when volume accumulates.
Should I exclude bot traffic from the baseline calculation?
Ideally yes — run a client-side behavioral audit (BotRefund) first, flag bot sessions, exclude them from baseline rates. If you can't, note that your baseline includes some invalid traffic and interpret low indices cautiously.
How often should I recalculate the baseline?
Quarterly for stable accounts. Monthly if you've made major changes (new offer, new pixel, new CRM, seasonality shift). Always recalculate after a confirmed bot-traffic cleanup.
Further reading and comparison sources
These external sources provide additional context for evaluating the topic. Their inclusion is not an endorsement.
Learn more
Visit the website for more information.