Er is al jaren een interessante dynamiek aan de gang bij Wikipedia:
- eerst de eeuwige klaagzang over “de buitenwereld” die niet weet hoe het er “bij ons” aan toegaat, en hele roedel editors die dan maar in de loopgraven kruipen en zich een vestingsmentaliteit aanpassen
- dan de eeuwige klaagzang over het dalende aantal editors, min of meer samen met de eeuwige klaagzang dat het bijna allemaal witte mannen zijn
- dan de eeuwige klaagzang over “Google pakt onze inhoud gratis af”
- dan, bij sommigen, jarenlang klagen dat niemand lijkt wakker te liggen van de gebruikerservaring
- …en dat niemand lijkt wakker te liggen van het alsmaar dalend aantal gebruikers
- …en dat niemand het lijkt erg te vinden dat ze allerlei technische boten missen — zoals video bijvoorbeeld
- …en dat er alsmaar meer en meer administratie aangenomen wordt bij de Foundation, en dat het alsmaar schaamtelijker wordt op wat voor berg geld ze zitten en hoe weinig ze er mee doen.
“Google Zero” was eventjes een heel doembeeld: wat gaat Wikipedia doen als ze niet eens meer in Googleresultaten zullen staan?
Alex Stinson had deze bijzonder boeiende reactie op de mailinglijst, die ik in extenso copieer. Lees en huiver.
I don’t think we have a shared framing of what we are competing for in a “Google” Zero landscape dominated by ChatBot/AI search style RAG citations (i.e. https://en.wikipedia.org/wiki/Retrieval-augmented_generation).
For the last 9 months, I’ve been examining how civil society content should be showing up in AI search, and there is a missing perspective in our “we will build it and they will come” approach to Wikipedia. I don’t think we can wait for the big tech companies or European regulatory bodies to adopt a different idea of how RAG should work. Here is my take from what I have been engaging with in the AIO/AEO/SEO space:
Wikipedia doesn’t have content that is SEO/AEO optimized
Part of the problem, even if we did have an MCP server is that the models (at least in my tracking), are pushing many citations away from “factual” websites, towards “authoritative” websites. This authoritative content includes:
- Expert original, analysis that makes strong claims based on facts (i.e. blog posts by authoritative companies or recently published ScienceDirect articles)
- Content that has been updated recently, with the biggest “hot takes” (i.e. I have monitored a couple of prompt pools where citations shift to newer content after 2-3 months)
- Content that helps users make a decision between different choices (i.e. review websites, etc)
This is following Google’s longer-term push towards “human centered and useful” content (sometimes called E-E-A-T an abbreviation of experience, expertise, authoritativeness, and trustworthiness, in SEO world). https://developers.google.com/search/docs/fundamentals/creating-helpful-content
To win in an AI optimization battle — its less about Wikipedia doing well in the keyword search indexes that led to our content being visible (which is why we have a reputation as a “fact checking” website) and more about “winning” in the criteria for what makes a good RAG citation — and our content format, is the exact opposite of the EAAT criteria:
- Wikipedia is not authoriative, but rather points to other authorities
- We ground our content in anonymity instead of named experts or instutional process/opinoin
- We rarely do original analysis instead summarizing the experience and expertise of others,
- Alot of our content is out of date, and self-aware of its gaps (i.e. maintenance tags), so also is likely to be undermining its own trustworthiness
All the data points to us being used, but without an official roundup we are all talking in the dark about different assumed reputation losses
RAG unlike Google Search Indexing, seems to be using Wikipedia for a fraction of a fraction of responses, favoring these other kinds of sources:
RAG/AI search optimization focuses more on intent than keywords, and we aren’t very effective at serving intent, and we don’t know where our optimization options are
What we need is an understanding of “which actual user reader behavior are we seeking to serve?”. In the past we were extremely lazy, because keyword search always delivered Wikipedia as “a first”. Now we need our content to be more optimized for the kind of user curiosity driving their use of a chatbot/search tool:
- What percentage of prompts or AI searches are informational vs opinion forming? Are we even a competitor for grounding opinion based questions or only the informational ones?
- How many of the interactions are two or three steps down a chain of more “specific” interactions with the chatbot and thus no longer need “general knowledge” information from Wikipedia, but rather the kinds of stuff that we rely on our citations to provide ?
- How much are the AI companies optimizing for “sales” or “addiction” rather than for leading users to reliable content? (I was tracking a series of informational topics about food that (on ChatGPT and Google), kept wanting me to continue the conversation by inviting me to go to local hamburger restraunts). Do we even have a reasonable chance to be in those searches?
- How much is geolocation forcing more and more responses into “local” sources rather than “global” websites? In one dataset I tracked, in Global South countries citations were overwhelmingly to Facebook and Instagram despite more authoritative academic, news and Wikipedia-type sites in the same searches from the UK.
We may need to radically change the “readable signals” on our content pages, meaning changing the Manual of Style, Editing Practices, and AI enabled enrichment.
If we are trying to market Wikipedia’s content into AI interfaces, we also can’t do what most AI optimization/marketing agencies would suggest: writing listicle/FAQ type content that closely matches the user-queries that folks are giving the IA models (i.e. analysis like: https://neilpatel.com/marketing-stats/trust-signals-ai-engines-reward-most/ ).
We would then have to experiment with other content types, that _no longer look like the encyclopedia_. Or we would need to be reconfiguring the Encyclopedic content to expose enrichments to paragraphs or sections within the encyclopedia that pretty radically change editorial assupmtions and our Manual of Style (i.e. instead of simple 1-2 word section headings, like “History” we may need intent-focused headings like “What is the history of [x topic]?).
If we want to compete in the shifting AI search landscape — we would need a lot more data from the Foundation on where we are succeeding or not, and then consider _radically different_ ways of exposing our content in terms of treating RAG systems as a user that needs correct paths to Wikipedia pages.
However, this doesn’t necessarily need to change the human reader experience , but would need to be about configuring the content (beyond an MCP server or Enterpise APIs) for an AI audience/consumer experience — which I haven’t seen addressed in any WMF publications or community conversations. Without a firm theory of “What kind of consumer is an AI search agent/RAG index?” and “How does our content need to serve that AI audience?” the editing community won’t be able to adjust its editing practices or weigh in on feature recommendations that make our content “AI useful”.
As I have written elsewhere, I think there is a inherent audience for editing/using the Wikis organically: https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed — but its a different question than competing “with other information sources” for AI as an audience.
De reactie van Todd Allen is emblematisch voor alles dat fout gaat in die hele wereld:
We sure do know what a good AI strategy looks like for Wikipedia. No AI on Wikipedia.
We have not succeeded by being FaceGramTwitTube. We have succeeded by not being like them.
So, same here. No “latest and greatest”. No AI on Wikipedia. Ever, for any reason, period. Wikipedia is written by people for people.
Zucht.