Seatext library / BotRefund evidence
How Machine Learning Improves Human and Bot Behavior Detection
Machine learning models analyze 106 browser, network, hardware, and behavior signals together instead of scoring single indicators. This pattern-based approach catches sophisticated bots that use residential proxies and browser automation, which traditional IP blacklists...
✓ Built for advertisers who need clear, refund-ready traffic evidence.
Machine learning improves bot detection by evaluating how hundreds of signals fit together rather than checking each one in isolation. BotRefund's prediction AI examines 106 browser, network, hardware, and behavior signals as a combined pattern to classify traffic as human or bot with 99% accuracy. Single signals like IP reputation or user-agent strings can be spoofed; the full pattern cannot be easily faked.
Why single-signal scoring fails against modern bots
Traditional click-fraud tools rely on IP blacklists, rate limits, and user-agent checks. These methods miss bots that rotate residential proxies and run real browser engines. A bot on a residential IP with a valid Chrome user-agent looks identical to a human in server logs. Server-side audits only see IP addresses, request headers, and user-agent data, which catches basic scrapers but struggles with advanced botnets.
Machine learning changes the game by moving detection to the client side. The browser itself becomes the sensor. When a visitor loads a page, the ML model collects hardware fingerprints, network timing, mouse dynamics, and JavaScript engine behavior. These signals are difficult to forge simultaneously because they come from different system layers.
How the prediction AI evaluates 106 signals as a unified pattern
BotRefund's prediction AI sees how 106 browser, network, hardware, and behavior signals fit together before deciding whether a visit is human or automated. The model does not assign a risk score to each signal independently. Instead, it learns the joint distribution of legitimate human sessions across devices, networks, and geographies. A visit that matches the marginal distribution of each signal but violates their conditional dependencies gets flagged.
For example, a visitor may have a correct timezone, language, and IP geolocation individually. But if the WebRTC network leak reveals a different location, the DNS tunnel shows a mismatched route, and the TCP TTL doesn't match the claimed OS, the combination is statistically impossible for a real user. The ML model catches this inconsistency without any single signal crossing a hard threshold.
Network, VPN, and geolocation evasion detection
Bots often hide behind VPNs, proxies, or spoofed geolocation settings. The ML model checks for coherence across network-layer signals:
- WebRTC Network Leak: Checks whether browser network paths reveal conflicting locations.
- DNS Tunnel Leak: Checks whether DNS and web traffic follow the same route.
- DNS Challenge Blocked: Checks whether DNS and web traffic follow the same route.
- Timezone Evasion: Checks whether location and language settings agree.
- Latency Mismatch: Checks whether connection and browser request details stay consistent.
- Suspicious Ports: Checks whether the visitor's network identity is coherent.
- UTC Timezone Bias: Checks whether location and language settings agree.
- Languages Mismatch: Checks whether location and language settings agree.
- Netprobe Telemetry Missing: Checks whether the visitor's network identity is coherent.
- IP Address Inconsistency: Checks whether the visitor's network identity is coherent.
- OS / TCP TTL Mismatch: Checks whether the visitor's network identity is coherent.
- HTTP User-Agent Mismatch: Checks whether connection and browser request details stay consistent.
- Accept-Language Mismatch: Checks whether location and language settings agree.
- HTTP Protocol Mismatch: Checks whether connection and browser request details stay consistent.
- DNS Routing Mismatch: Checks whether DNS and web traffic follow the same route.
Each check alone produces false positives. Corporate networks, privacy tools, and mobile carriers create legitimate mismatches. The ML model learns which combinations occur in real traffic versus bot traffic, reducing false blocks.
Evasion, debugger, and anti-stealth trap detection
Sophisticated bots use automation frameworks like Puppeteer, Playwright, or Selenium, often wrapped in stealth plugins that patch browser APIs. The ML model looks for traces these tools leave:
- CDP Debugger Leak: Checks for traces left by browser automation or masking tools.
- Native Patching: Checks whether the browser profile behaves like a real device.
- Engine Mismatch: Checks whether the browser profile behaves like a real device.
- Rebrowser Leaks: Checks for traces left by browser automation or masking tools.
- JS Engine Mismatch: Checks whether the browser profile behaves like a real device.
- Automation Properties: Checks for traces left by browser automation or masking tools.
These signals detect when the JavaScript engine, DOM APIs, or Chrome DevTools Protocol have been modified. Stealth plugins can hide individual properties, but they rarely replicate the full behavioral profile of an unmodified browser across all 106 signals.
Behavioral biometrics: mouse, speed, path, engagement, and session
Human interaction has micro-patterns that automation struggles to replicate. The ML model analyzes:
- Pointer behavior: Robotic linear mouse movements flag unnaturally straight pointer paths. Absence of humanlike mouse tremor looks for tiny imperfections and jitter typical of human movement.
- Speed behavior: Superhuman input speed (<1ms) identifies interactions faster than a person could realistically perform.
- Path behavior: Grid-aligned movement patterns detect movement that snaps to precise lines or blocks instead of natural curves.
- Engagement behavior: Absence of clicks or scrolling highlights sessions that stay too static to match a real browsing journey.
- Session behavior: Unnatural session durations catch visit lengths that are too short, too long, or too uniform to be human.
Click farms using real smartphones bypass IP filters but still produce detectable behavioral signatures: uniform timing, missing scroll events, and repetitive click coordinates. The ML model learns these patterns from millions of labeled sessions.
Client-side detection versus server-side audits
Server-side audits examine logs after the fact. They see IP addresses, headers, and user agents. Client-side detection runs in the visitor's browser during the session. This enables real-time filtering and captures signals impossible to see server-side: canvas fingerprints, WebGL renderer details, audio context behavior, and precise input timing.
The distinction matters for refund evidence. Ad platforms like Google and Meta require Google Click IDs (GCLIDs) or Facebook Click IDs (FBCLIDs) linked to behavioral proof of invalidity. Client-side ML captures the click ID at the moment of interaction and attaches the full 106-signal evidence package. Server-side tools cannot reliably connect a click ID to a specific browser session.
From detection to refund evidence: ML-powered evidence generation
Detection alone doesn't recover money. The ML system auto-captures click IDs with behavioral evidence and generates compliance-ready refund reports. BotRefund helps large advertisers and agencies prove invalid clicks, prepare the evidence, and negotiate directly with Google and Meta to recover wasted ad spend. The 83% refund success rate for high-volume advertisers comes from evidence that meets platform dispute requirements.
The workflow: ML classifies the session as invalid in real time, the click ID is stored with the 106-signal fingerprint, a dispute report is generated automatically, and the advertiser submits it through the platform's billing dispute process. Refunds can be recovered for Google Ads spend dating back to 2017.
Limitations and when human review is needed
ML detection has boundaries. Legitimate users on unusual network configurations (corporate VPNs, privacy browsers, satellite internet) can trigger signal mismatches. The model minimizes false positives by learning the joint distribution, but edge cases exist. Not every bad lead is a bot; treating every unresponsive contact as fraud can exclude valuable audiences.
Signals worth investigating before labeling fraud: contactability issues (disconnected numbers, invalid emails), timing anomalies (burst leads, instant form submits), session behavior (no scrolling, uniform paths), campaign patterns (sharp quality differences by placement or device), and CRM outcomes (high lead count but no qualified opportunities). A structured audit comparing ad-platform data, website sessions, and CRM outcomes should precede refund requests.
Key facts
| Metric | Value | Source |
|---|---|---|
| Signals analyzed by prediction AI | 106 browser, network, hardware, and behavior signals | S1 |
| Classification accuracy claim | 99% accuracy | S1 |
| Ad spend drained by bots | Up to 20% of Google Ads and Meta spend | S2 |
| Refund success rate | 83% for high-volume advertisers | S2 |
| Refund lookback window | Google Ads spend dating back to 2017 | S2 |
| Detection method | Client-side behavioral analysis with ML pattern recognition | S1, S3, S7 |
| Evidence captured | GCLIDs and FBCLIDs linked to 106-signal behavioral proof | S2, S7 |
Frequently asked questions
How does ML detection differ from traditional IP blacklists?
IP blacklists block known bad addresses. Bots rotate residential proxies daily, making blacklists obsolete. ML detection analyzes behavior patterns that are expensive to forge at scale, regardless of IP reputation.
Can ML detection run without slowing down page load?
Yes. The client-side script loads asynchronously and collects signals during the session. Classification happens in real time without blocking page rendering.
What happens when a legitimate user gets flagged?
The system minimizes false positives by requiring multiple signal inconsistencies. Edge cases (corporate VPNs, privacy tools) are reviewed before any blocking or refund claim is made.
Does ML detection work on mobile apps and AMP pages?
Client-side detection requires JavaScript execution. Mobile web and AMP pages support it. Native mobile apps need SDK integration for equivalent signal collection.
How much historical data is needed to train the model?
The model comes pre-trained on millions of labeled sessions. It adapts to your traffic patterns within days of installation.
Can I use ML detection alongside existing fraud tools?
Yes. ML detection complements server-side filters. It catches bots that bypass IP and user-agent checks, providing evidence for refunds that other tools don't generate.
What ad platforms support refund claims with ML evidence?
Google Ads and Meta (Facebook/Instagram) have formal invalid-click dispute processes that accept behavioral evidence linked to click IDs.
Further reading and comparison sources
These external sources provide additional context for evaluating the topic. Their inclusion is not an endorsement.
Learn more
Visit the website for more information.