TL;DR: GA4 automatically excludes known bots, but that does not establish that all remaining traffic is human. This guide explains how automation can distort ecommerce reports, how to investigate suspicious sessions, and why analytics filtering and blocking requests are different controls.
On this page
- What GA4 filters, and what it misses
- Five metrics bots quietly distort
- How to find bot traffic in Google Analytics: a six-step audit
- Build a clean reporting baseline
- Why filtering is not the same as fixing
- AI crawlers and shopping agents: the 2026 wrinkle
- Frequently asked questions
What GA4 filters, and what it misses
Google says GA4 identifies known bots using its research and the IAB International Spiders and Bots List. This exclusion is automatic: you cannot disable it or view the amount excluded. See Google’s documentation.
A known-bot exclusion is not a complete security control. Investigate unusual traffic rather than assuming that every automated request is removed or that every suspicious session is a bot.
- Browser automation: examine repetitive behavior and request patterns, not just the browser name.
- Proxy traffic: an IP address or geography alone does not establish whether a visitor is human.
- Suspicious paid traffic: compare campaign events with actual orders and the ad platform’s invalid-traffic reports.
- Scraping: server request logs can reveal activity that never produces a GA4 event.
- Checkout probes: compare analytics events with payment-system records when investigating failed-payment bursts.
Industry bot percentages describe different datasets and cannot establish your store’s bot share. Use a consistent measurement method and document its limits.
Five metrics bots quietly distort
The reason to care is not data hygiene for its own sake. It is that five numbers your team makes decisions on are being moved, and they are moved in a consistent direction that flatters volume and punishes efficiency.
1. Sessions and users (inflated)
Automation adds sessions with no purchase intent. Traffic looks like it is growing, so nobody investigates. This is the failure mode we covered in why bot traffic is never evenly distributed: a spike that reads as momentum is often just a new scraper finding your catalogue.
2. Conversion rate (deflated)
Orders sit in the numerator, bot sessions pile into the denominator. A store doing 2.4% on clean traffic reports 1.9% when a fifth of sessions are automated. Teams then spend a quarter running CRO tests against a number that was never a conversion problem.
3. Engagement rate and average engagement time (both directions)
Crude scrapers hit one URL and leave, dragging engagement down. Sophisticated ones scroll, wait, and click to defeat behavioural detection, which drags engagement up and makes the traffic look better than your real customers. Neither is a signal you want in a dashboard.
4. Channel and geography reports (misattributed)
Direct and Referral absorb most uninvited automation, which is why sudden Direct growth with flat revenue is a tell. Geography is often louder: sessions from countries you do not ship to, or a concentration in data centre metros such as Ashburn, Virginia. We catalogued the full pattern in seven signs of a bot attack.
5. ROAS, CAC, and experiment results (all downstream)
This is the expensive one. Invalid clicks consume budget and register as sessions, so cost per session looks reasonable while cost per order climbs. Smart bidding then optimizes toward whatever cohort produced the most events, which can include the fake ones. And because bot sessions distribute across your test variants, they add noise without adding signal, which means A/B tests need more traffic and longer runtimes to reach significance. Some never do.
How to find bot traffic in Google Analytics: a six-step audit
You cannot get a perfect number from GA4 alone, but you can get a defensible estimate in about an hour. Run these six checks over a 90-day window.
- Plot sessions against revenue. In a single Exploration, chart sessions and purchases by day. Healthy stores move together. Divergence, especially a step change that starts on one specific date and never reverts, is your first flag.
- Break out sessions by country and city. Add city as a secondary dimension. Look for volume from markets you do not sell to and for clusters in known hosting regions. Compare each segment's conversion rate to your site average.
- Audit Direct traffic. Filter to Direct, then look at landing pages. Real direct traffic lands on your homepage, brand campaigns, and saved product pages. Bot direct traffic lands deep in collection pagination, on search result URLs, or on the same product page thousands of times.
- Check operating system and browser mix. Look at the tail, not the top. Unusual Linux desktop share, ancient browser versions, or a single Chrome build overrepresented by orders of magnitude are all worth a look.
- Compare funnel steps to orders. Build a funnel from page_view to add_to_cart to begin_checkout to purchase. When add_to_cart volume climbs but purchase does not, you are likely looking at cart-probing automation rather than a checkout usability issue.
- Reconcile GA4 against your platform and your CDN. Shopify analytics, your GA4 property, and your edge or CDN logs will never match exactly, but the shape of the gap is informative. Server-side logs see requests that never execute JavaScript, which is precisely the population GA4 is blind to. Our walkthrough of how to tell if Shopify traffic is bots covers that reconciliation in detail.
One caution on step 6. Not everything non-human is unwanted. Search crawlers, AI assistants you want citing you, monitoring, and partner integrations all need access. Sorting the two groups is its own exercise, which is why we wrote a separate primer on good bots versus bad bots. Verify claimed identity rather than trusting the string: Google publishes a reverse DNS and IP verification process for exactly this reason, and anything claiming to be Googlebot that fails it is not Googlebot.
Build a clean reporting baseline
Once you know roughly how much of your traffic is automated, the next job is making sure your reports stop lying to you. Four moves, in order of effort.
Confirm the default exclusion is on. It is enabled by default per data stream, but it gets missed on new streams and on properties migrated in a hurry. Check it, then stop thinking of it as protection.
Filter internal and developer traffic. Your own team, your agency, your staging monitors, and your QA scripts all generate sessions. Define internal traffic by IP in Admin, then apply the data filter. Set it to Testing first so you can see what it removes before it becomes permanent, because active filters are not retroactive.
Add an unwanted referrals list. This handles referral spam and self-referrals from payment gateways and app subdomains, which quietly break attribution on every order that passes through them.
Build a real-users segment in Explorations. This is the highest-value step and the one most teams skip. Create a session-scoped segment that requires at least one meaningful engagement event, then exclude the geographies and city clusters your audit flagged. Report your CRO and experiment numbers against that segment, and keep raw sessions for traffic-cost analysis. Two numbers, clearly labelled, beats one number nobody trusts.
Why filtering is not the same as fixing
Here is the limit of everything above. Filtering changes what you see. It does not change what happens.
A scraper you have excluded from GA4 is still hitting your origin. It still consumes server capacity, still competes with real shoppers for resources during peak, still pulls your pricing and inventory for a competitor's repricing engine, and still costs you money on every request. Card testing you filtered out of reports still generates gateway fees and chargeback risk. Click fraud you segmented away has already spent the budget.
There is a performance cost too, and it lands on the customers you do want. When automated request volume spikes, response times degrade for everyone, and slower pages convert worse. That is the compounding version of the problem: bots inflate your denominator, and separately they make your numerator smaller by slowing the site down. We have written about that interaction in site speed, bot traffic, and conversions.
So treat clean analytics as diagnosis, not treatment. The treatment is deciding, at the edge and before the request reaches your store, which automation gets served, which gets challenged, and which gets refused. That decision has to happen in milliseconds, and it has to distinguish a paying customer on a shared mobile IP from a scraper on the same network. Rate limiting alone will not do it, and neither will blocking by country, which tends to cost real revenue.
AI crawlers and shopping agents: the 2026 wrinkle
Two recent shifts make GA4's blind spot bigger.
First, AI crawler volume is now material. GPTBot, ClaudeBot, PerplexityBot, Google-Extended, and a long tail of smaller agents index product catalogues continuously. Most declare themselves, so the well-behaved ones are filterable. But their infrastructure cost is real even when their analytics cost is zero, and imitators using those names without the matching IP ranges are a growing share.
Second, agentic browsing is showing up in ecommerce traffic. When a shopping agent browses for a customer, it is neither a bot you want to block nor a session you want to count as a human visit. In GA4 today it looks like a strange human: unusual engagement pattern, no scroll depth, sometimes no viewport at all.
This breaks naive bot classification. The old binary of "human good, bot bad" no longer maps to intent, because a meaningful slice of automated traffic in 2026 has a real buyer at the other end of it. Our take is in AI shopping agents and ecommerce. The operational implication: stop maintaining one number for traffic. Maintain three buckets, humans, agents acting for humans, and unwanted automation. Any model that collapses those into a single sessions metric will keep producing decisions that do not survive contact with revenue.
Frequently asked questions
Does GA4 automatically remove all bot traffic?
No. Google documents automatic exclusion of known bots using its research and the IAB list. That does not establish complete coverage of all automation. Google does not expose a count of the excluded traffic.
How much of my ecommerce traffic is probably bots?
It varies by store, period, and measurement method. GA4 sessions and server requests are different measures. Investigate your own traffic and document how it was classified instead of applying an industry percentage to your store.
Why did my conversion rate go up after filtering bot traffic?
Excluding non-converting sessions can increase the reported conversion rate without creating additional orders. Compare the same date range, session definition, and filtering rules before attributing a change to better shopper conversion.
Can bot traffic affect my Google Ads performance and ROAS?
Yes, in two ways. Invalid clicks spend budget directly. Less obviously, if automated sessions generate events that your bidding strategy treats as signal, smart bidding can optimize toward acquiring more of them. Google filters some invalid traffic and issues credits, but detection is not complete, and the analytics distortion persists in your own reporting either way.
Should I block AI crawlers to clean up my analytics?
Do not block useful crawlers just to tidy a dashboard. Verify crawler identity and evaluate its purpose, request volume, and infrastructure impact. Declaring an AI user agent alone does not prove that GA4 excludes every visit from that crawler.
What is the difference between filtering bots in GA4 and blocking them at the edge?
Filtering is a reporting change applied after the request has already been served. Blocking at the edge is an infrastructure change that stops the request before it reaches your store, which is what actually recovers server capacity, protects pricing and inventory data, reduces gateway abuse, and keeps peak-traffic response times stable for real shoppers.
See what is actually in your traffic
Most brands find their bot problem the slow way: a quarter of flat conversion rate, a CRO roadmap built on a broken baseline, and a peak season where the site got slow for reasons nobody could name. Clean analytics tells you the size of the problem. Edge enforcement removes it.
Nostra runs both sides from the edge. Edge Protect classifies and stops unwanted automation before it reaches your store while letting search crawlers, AI assistants, and real shoppers through, and our Edge Delivery Engine keeps the pages customers do see fast under load. If your sessions and your revenue have stopped moving together, start with a traffic breakdown or book a demo and we will show you the split on your own domain.
