Seatext library / BotRefund evidence
How to Build a Lead Scoring Model That Avoids False Bad Labels
Start with intent and engagement data, layer in traffic source validation, and regularly test your scoring thresholds. A reliable model separates automated or fraudulent leads from real prospects who simply aren't ready to buy,...
✓ Built for advertisers who need clear, refund-ready traffic evidence.
False bad labels happen when a scoring model treats every unresponsive lead as fraud. The result: you discard real prospects who need nurturing, and you feed the ad platform corrupted conversion signals that optimize for bots. The fix is a model that weighs multiple evidence layers — contactability, session behavior, timing patterns, campaign-level quality clusters, and sales dispositions — before assigning a negative score.
Define what a false bad label looks like in your funnel
A false bad label is a real human prospect marked as invalid, fraudulent, or unqualified because they didn't convert quickly or match a narrow profile. This differs from a true bad lead — automated form fills, bot clicks, or deliberate fraud. The source pack emphasizes that "not every bad lead is a bot, and that matters. Treating every unresponsive contact as fraud can make a team exclude a valuable audience." (S1) Start by agreeing on definitions: a suspicious session is a signal for investigation, not proof on its own. (S6)
Collect the right evidence before you score
Build your model on observable signals, not assumptions. The audit framework in the source pack identifies five signal categories worth investigating:
- Contactability: disconnected numbers, invalid email domains, repeated addresses, unusual country-code concentration. (S1)
- Timing: bursts of leads in short windows, forms submitted immediately after landing, conversions at unusual hours. (S1)
- Session behavior: no scrolling, no field corrections, uniform click paths, no meaningful time on the offer page. (S1)
- Campaign patterns: sharp lead-quality differences by placement, creative, audience expansion, device, or landing page. (S1)
- CRM outcome: high reported lead count paired with no calls connected, demos booked, qualified opportunities, or repeat engagement. (S1)
Preserve the click identifier, campaign context, timestamp, URL parameters, CRM record, and any verification result before you change campaign settings. (S6)
Layer traffic-source validation into the score
Invalid traffic often enters through specific channels. Meta's Audience Network, for example, has historically shown high click-through rates and near-instant bounce rates because publishers use bots to generate artificial revenue. (S4) Profile scrapers and directory bots crawl social platforms and follow outbound links. (S4) A lead scoring model that ignores source context will penalize prospects who came from a noisy placement but are otherwise genuine. Tag each lead with its placement, network, and campaign hierarchy so the model can weight source risk separately from prospect intent.
Use a four-layer audit to calibrate thresholds
The source pack outlines a practical investigation workflow that doubles as a scoring calibration process:
- Platform delivery: Compare reach, link clicks, landing-page views, placements, and spend. A cheap placement isn't a win unless it produces contacts that can be reached and qualified. Avoid eliminating an entire audience from a small sample; use enough volume to see a consistent quality pattern. (S6)
- Landing-page evidence: Measure page loads, redirects, consent behavior, form start, form completion, time to completion, and meaningful engagement. A click-to-session gap can have ordinary explanations — app browsers, tracking consent, slow loads, analytics configuration. Investigate those before concluding the gap is bot traffic. (S6)
- Lead verification: Record whether an email is deliverable, a phone connects, duplicate details recur, and the prospect confirms interest. Add qualification questions that reveal fit, not just extra fields that make the form longer. For high-value offers, a confirmation step or booking flow can be more valuable than the cheapest raw lead. (S6)
- Sales outcome feedback: Give sales a small, mandatory set of dispositions: verified, contacted, qualified, disqualified, duplicate, invalid details, no response. Feed these dispositions back into the scoring model so it learns which signals actually predict revenue. (S6)
Separate bot detection from lead qualification
Bot detection identifies non-human traffic — automated web crawlers, scrapers, click farms, publisher script engines. (S3) Lead qualification assesses whether a human prospect fits your ideal customer profile. Conflating the two creates false bad labels. The source pack notes that client-side behavioral verification (mouse tremor, superhuman input speed, grid-aligned movement, honeypot interactions) catches bots with high confidence. (S2) Use that verification as a hard filter: if a session is confirmed non-human, exclude it from scoring entirely. Then score only verified human sessions on fit and intent.
Build feedback loops so the model self-corrects
A static scoring model decays. The source pack reports that BotRefund achieves an 83% approval rate on refund claims filed with ad platforms, which implies that evidence quality improves when you iterate. (S7) Implement these loops:
- Weekly: review disposition distributions by score band. If "verified" leads cluster in a low-score band, lower the threshold or add a positive signal.
- Monthly: re-run the four-layer audit on a sample of leads marked "bad" by the model. Count how many were false bad labels.
- Quarterly: retrain or re-weight using the latest CRM outcomes, not just lead-volume metrics.
Key facts
| Signal category | What to measure | Source |
|---|---|---|
| Contactability | Disconnected numbers, invalid email domains, repeated addresses, country-code concentration | S1 |
| Timing | Lead bursts, instant form submits, unusual-hour conversions | S1 |
| Session behavior | No scrolling, no field corrections, uniform click paths, low time on page | S1 |
| Campaign patterns | Quality gaps by placement, creative, audience, device, landing page | S1 |
| CRM outcome | Lead count vs. calls connected, demos booked, qualified opportunities | S1 |
| Bot detection confidence | 99% confidence in identifying non-human traffic | S7 |
| Refund claim approval rate | 83% of filed claims approved by ad platforms | S7 |
Limitations and when this advice doesn't apply
- If your lead volume is too low to form statistically meaningful clusters by placement or creative, the four-layer audit may produce noisy patterns. Wait for sufficient volume or aggregate across longer windows.
- The bot detection signals described (mouse tremor, input speed, honeypot traps) require client-side JavaScript execution. They won't work for server-side-only tracking or leads that come through offline channels.
- Refund recovery processes apply to Google and Meta ad spend. They don't apply to organic traffic, email marketing, or direct sales outreach.
- The 83% approval rate and 99% detection confidence are aggregated client results reported by BotRefund. Your individual account results will vary based on traffic mix, spend level, and evidence quality. (S7)
FAQ
How do I know if my current model produces false bad labels?
Pull a sample of leads your model scored as "bad" or "low quality" and check their CRM dispositions. If you find verified, contacted, or qualified leads in that sample, your model is generating false bad labels. The four-layer audit (S6) gives you a structured way to investigate.
Should I block traffic sources that show high bot rates?
Not automatically. The source pack warns against eliminating an entire audience from a small sample. Use enough volume to see a consistent quality pattern first. (S6) You can exclude specific placements or networks in the ad platform while keeping the broader campaign active.
What's the difference between a low-quality lead and a bot lead?
A low-quality lead is a real person who doesn't fit your offer or isn't ready to buy. A bot lead is non-human traffic — automated scripts, scrapers, click farms. (S3) The scoring model should treat them differently: nurture the low-quality human, exclude the bot entirely.
How often should I recalibrate scoring thresholds?
At minimum, monthly. The source pack's emphasis on preserving attribution before changing campaigns (S1) and the iterative audit process (S6) both imply continuous recalibration. Weekly disposition reviews and quarterly retraining are practical cadences.
Can I use ad-platform invalid-traffic credits as a proxy for lead quality?
No. Google and Meta's automated systems catch only a fraction of invalid activity. (S5) Relying on platform credits means you're scoring leads after the platform has already billed you for the clicks. Client-side behavioral verification gives you session-level evidence the platforms don't see. (S2)
What's the minimum data I need to start this process?
You need click identifiers (GCLID, FBCLID), landing-page session data, form submissions, and CRM dispositions for at least a few hundred leads. Without click-to-CRM linkage, you can't tie source signals to outcomes.
How BotRefund can help
BotRefund adds client-side behavioral verification to your site — detecting bots through mouse tremor analysis, superhuman input speed, grid-aligned movement patterns, honeypot trap interactions, and ghost click detection. (S2) It captures video proof for each flagged click, builds compliance-grade evidence reports, and negotiates refunds through Google and Meta's own invalid-traffic channels with an 83% approval rate across filed claims. (S7) The script installs in about one minute with no ad-account access required. (S7) This evidence layer feeds directly into the lead scoring model: sessions flagged as non-human are excluded from scoring, while verified human sessions flow into your qualification logic with clean attribution preserved.
Limitation: BotRefund addresses paid traffic on Google and Meta. It does not score leads, manage CRM dispositions, or replace your qualification logic. It supplies the traffic-quality evidence your scoring model needs to avoid false bad labels.
Further reading and comparison sources
These external sources provide additional context for evaluating the topic. Their inclusion is not an endorsement.
Learn more
Visit the website for more information.