Search Signal #03
Cloudflare's crawler defaults flip tomorrow, and blocking Training also blocks Googlebot. Plus: People Also Ask hits 97% AI, and ChatGPT built its own index.
Vikas Kalwani
Founder, SAASY LINKS
01. A Cloudflare setting can block Googlebot from tomorrow
From 15 September Cloudflare applies new bot defaults across three categories: Search, Agent and Training. Newly onboarding domains and existing free-tier customers get Training and Agent crawlers blocked by default on ad-monetised pages. The sharp edge is in Cloudflare's own post: "multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training." Most restrictive rule wins, enforced at network level rather than through robots.txt.
Why it matters: A client who ticks Training to keep content out of model training can silently cut off Googlebot, because Google has not split its crawler. This is the only item this week with a deadline attached.
Action: Open Cloudflare Security settings for every client on Cloudflare today and confirm Search is allowed independently of Training. Existing paid customers are not flipped, but free-tier and new domains are.
02. People Also Ask is now 97% AI-generated, from 12% a year ago
Two independent trackers agree. AlsoAsked measured 97% of People Also Ask boxes answered by AI Overviews in the first week of September, up from 86% in August, across 19.2 million English-language queries. Allintitle measured 100% as of August. Fourteen months ago it was 12%. Mark Williams-Cook: "It seems highly likely we are on track for 100% shortly."
Why it matters: People Also Ask was a cheap, durable link surface for mid-funnel informational content, and it has effectively closed. Both figures came via LinkedIn posts with no methodology write-up, so treat the exact percentage as soft and the trend as solid.
Action: If any content plan still has "win the PAA" as a line item, reprice it before renewal season. It now buys an AI Overview citation at best.
03. The best evidence yet that ChatGPT runs its own search index
Tomek Rudzki documented an OpenAI retrieval stack with vertical indexes across general web, PDF, YouTube, news, arXiv, Wikipedia, local, finance, legal, medical, shopping and images. Between 21 May and 21 July he captured server-side events exposing a result_source field reading "Labrador" alongside external providers. A test site with one billion pages saw ChatGPT crawl 6 million pages at 35,000 requests an hour by early September.
Why it matters: "Optimise for Bing because ChatGPT uses Bing" has been standing advice for two years and is now at best partially true. Single researcher, vendor-affiliated, and OpenAI has confirmed none of it, so hold it loosely.
Action: Pull OAI-SearchBot and GPTBot out of your server logs directly. Bing Webmaster Tools is no longer a usable proxy for ChatGPT retrieval, and any deliverable built on that assumption needs rewriting.
04. The top 100 publishers spent $113m in one month buying back their own traffic
Similarweb data reported by Adweek: the top 100 publishers spent $113 million on paid search in July 2026, up 41% year over year and 274% over three years, buying 23.7 million visits. Spend is growing faster than the visits it buys, so cost per visit is rising. Forbes alone accounted for $72.2 million, roughly 66% of the total. The New York Times spent $11.3 million, more than double its year-earlier figure.
Why it matters: The clearest price tag anyone has put on organic decline. These are panel-based estimates and one advertiser dominates, so the median publisher is not visible in the number.
Action: Use this in the CFO conversation. The alternative to earned visibility is not zero cost, it is a paid search bill compounding faster than the traffic it returns.
05. Wikipedia lost about 100 million monthly referrals to AI Overviews
A difference-in-differences study on Wikimedia clickstream data, December 2023 to December 2024, treating English Wikipedia as treated and German and French editions as controls across roughly one million article pairs. Result: a 5.45% decline against German and 4.82% against French, about 100.27 million fewer referrals a month. Not peer reviewed, now on its sixth revision, and the authors' own English-Japanese comparison returns 16.53% under a different method.
Why it matters: The first credible causal estimate rather than a correlation, and it is smaller than the numbers in circulation. AI Overviews cost real referrals at single-digit percentages, not the 40% collapse vendor decks imply.
Action: Keep this one to hand for when a client arrives with a scary chart. It sets expectations honestly without denying the effect.
06. Search Console had a bad week, and June data is gone for good
On 8 September, for about an hour, Search Console showed indexed pages including homepages as "Crawled, currently not indexed". John Mueller confirmed "a short blip that was resolved quickly in the night". Nothing was posted to the status dashboard. Then on 11 September a chunk of June data vanished from the page indexing report, and Mueller was blunt: "We don't back-fill indexing data."
Why it matters: Two reporting failures in one week on the tool every client screenshot comes from, neither posted to the status dashboard.
Action: Set up a weekly Search Console export so a Google-side gap never becomes a client-side argument. Separately, Mueller asked publicly what AI position reporting should look like, which is a rare narrow window to influence it.
07. Common Crawl read 584,107 llms.txt files. Two thirds are templates.
Common Crawl analysed 584,107 llms.txt files from its July 2026 archive. About 68.27% are templated or plugin-generated, with Wix alone accounting for 41.34%. 22.56% contain no links at all, despite a curated URL list being the format's entire purpose. 10.61%, some 136,578 files, were robots.txt served at the llms.txt path. Their own framing: "the file grants nothing and blocks nothing, and no crawler is obliged to read it."
Why it matters: The strongest evidence to date that llms.txt is mostly template noise. Non-profit research with nothing to sell, which is rare in this space.
Action: Keep llms.txt at the bottom of the roadmap. It costs almost nothing and does no harm, but it does not belong in a pitch deck. Useful ammunition for saying no to a client who just read a LinkedIn post about it.
08. Answer engines change their brand picks by persona
Profound collected 71,147 AI responses across ChatGPT, Claude and Gemini between 11 and 24 August, prefixing base prompts with persona attributes covering gender, age, income and occupation. Across personas, responses shared only about 1 in 4 brand mentions and 1 in 5 citations. Vendor research, consumer categories only, and personas here are prompt prefixes rather than accounts with history, so this measures prompt conditioning rather than true personalisation.
Why it matters: Discount it heavily and the direction still holds: one prompt, one run, one framing is not a visibility measurement.
Action: Add a persona axis to your prompt sets before a client asks why two tools disagree. If you report share of voice from a single prompt set, you are reporting one slice of a distribution nobody has measured.
09. Zyppy's ranking factors survey puts backlinks second
131 SEO professionals gave 13,665 ratings across 103 candidate ranking factors. The results: relevance and search intent 57.1%, backlinks 54.8%, content quality 47.6%, then trust and authority 36.5%, behavioural click signals 29.4% and brand presence 27.0%. Cyrus Shepard's line: "The rumor of the death of backlinks has been greatly exaggerated." This is opinion data from a convenience sample, with no significance testing.
Why it matters: It measures belief, not Google, and should never be cited as evidence about the algorithm.
Action: Read it as a map of what prospects have already absorbed before they reach you. Two years into the AI search cycle, links are still the second thing experienced SEOs name unprompted.
10. Brussels forced Google to stop stripping 90% of unique queries from EU search data
Under a Commission decision adopted 16 July, Google's proposed anonymisation, which removed 90 to 100% of unique queries, was rejected and replaced with a process removing only 10 to 20%. The obligations run deep: daily sharing with seven-day latency, access up to five years, record-level click, view and dwell-time data, and independent audits. Indicative pricing is EUR 1.50 to 9.00 per thousand queries. Licence agreements go out 17 September.
Why it matters: Long-tail query data is the moat, and the Commission has now priced it. If a rival engine or AI product actually licenses this, the query-level asymmetry that has held since roughly 2010 starts to close in Europe first.
Action: Watch who signs on 17 September. That list is the real story, not the pricing.
11. Google published new documentation on EEA search layouts
Google told Reuters its EEA results overhaul produces "the largest reduction in quality of service at the world's most popular internet search engine in its 29-year search history". That is Google's characterisation inside a live regulatory fight. The useful half is quieter: on the same day Google published a Regional differences in Search experience page documenting aggregator units, supplier units, ecosystem carousels and job site features across the EEA, Türkiye and South Africa.
Why it matters: The documentation names eligibility criteria for aggregator and supplier units by region and query type. Almost nobody has read it.
Action: If you have clients selling into the EEA, Türkiye or South Africa, read the aggregator features page this week. It is a concrete and currently uncontested placement surface.
12. Amazon started selling ads inside ChatGPT
US advertisers can now extend Amazon DSP campaigns into ChatGPT inventory through a managed-service pilot. Ads render as text and image units beneath ChatGPT responses with sponsored labels, priced on CPC and CPM. Amazon owns buying and campaign management while OpenAI keeps control of delivery and placement. No spend, volume or performance figures disclosed. Separately, The Information reported OpenAI has told ad partners it will not approve ads for competing image and voice generation tools.
Why it matters: A second demand-side platform is now buying the space directly beneath the answer, in the same zone as organic citations.
Action: If you run paid for any client selling into the US, request access to the pilot. The paid surface keeps getting more measurable while the organic one still reports impressions and no clicks.
Also worth knowing
- Google's ranking systems stayed quiet for a second week. No confirmed update and nothing on the status dashboard.
- Your rank tracker data may have broken in late August. Google's goto click redirects reached near-total rollout, so trackers now have to follow the redirect. Check your tool's data continuity before you explain a client dip.
- Bots are now more than half of all web requests. Cloudflare puts bots and AI agents above 50%, GoFish reports client bot traffic up 80% year over year, and agencies report a 20% CPM increase from cleaning bot contamination out of retargeting pools.
- Google documented a partner-only full-web Search API. JSON over REST or gRPC, up to 20 results per request, requiring a partner agreement. No pricing or eligibility published. The Custom Search JSON API sunsets January 2027, so every SERP-data vendor has roughly fifteen months and no announced route.
- Asking in the local language changes who gets cited by 3.5x to 13.5x. An audit of 1,920 health queries across 12 countries found querying in a country's official language rather than English sharply raises locally sourced citations. Health vertical only, preprint, but the lesson transfers: English-only content competes in the wrong citation pool.
- On the calendar. 15 September: Cloudflare's crawler defaults take effect. 16 September: Judge Brinkema's ad tech remedies opinion clears confidentiality review. 17 September: Google distributes European search data licence agreements. 29 September: Google's reply brief is due in the search remedies appeal.
SAASY LINKS
We build the visibility this newsletter reports on
Editorial links, AI citations, and organic pipeline for B2B SaaS. Senior-led, 90-day sprints, no long-term contracts.
Book a Free Strategy Call