Publishing on a web where a third of new pages are AI-written
How to keep your pages worth citing now that AI-assisted drafting is normal: what to check, what to disclose, and which detection advice to ignore.
on this page · 0 / 0 checked
You drafted a page with Claude last month. A client has now asked, politely, whether a human wrote it. A freelancer sent you 2,000 words, you pasted them into a detector, it came back “87% AI”, and you have no idea what to do with that number or with the freelancer. Underneath both problems sits the same suspicion: that the open web is filling up with fluent, confident, interchangeable prose, and that your pages are now part of the pile.
The suspicion is roughly correct, and it is measurable. What follows is what to change in how you publish because of it. This guide is for solo operators and small teams who publish their own pages and want them to stay worth reading and worth citing. It is not for anyone trying to beat a detector, and not for anyone planning to ship 500 pages a month; the first is not achievable and the second is explicitly against the rules of the search engine you depend on [2].
What the one-third figure actually measures
Pew Research Center analysed almost half a million English-language webpages from the past five years using an AI detection tool called Open Pangram, a machine learning model that checks text for signs of AI authorship by looking at patterns in language [1]. In a random sample of 10,000 pages collected in July 2026, 10% showed significant signs of AI authorship. Filter that sample down to pages published after ChatGPT’s release in November 2022, and the share rises to over one-third [1].
The domain split is the part worth pinning to your wall. Commercial .com pages sat at roughly 10%, .org at 4.6%, and .edu and .gov at about 1% each [1]. Every one of those domains has identical access to the same models. What differs is incentive. Sites that publish to win search visibility have a reason to publish more than a small team can write by hand, and sites that publish to fulfil a mandate do not.
Pew also tracked the stylistic drift. Since 2023, em dashes have appeared about twice as frequently across the sampled pages, Oxford commas rose 63%, words AI models like to use such as “delve”, “interplay” and “testament” more than doubled in usage, and negative parallelism of the “it’s not just X, it’s Y” shape nearly tripled [1]. None of that convicts any individual page. Plenty of careful writers use Oxford commas on purpose. In aggregate, across hundreds of thousands of documents, it is a fingerprint of who is holding the pen.
Detection scores cannot settle an argument about a single page
Pew is careful about this, and you should be too: AI detection models are not perfect, and they sometimes misclassify individual documents written by humans as showing signs of AI authorship, and the reverse [1]. That is why the finding is presented as a population-level trend rather than a verdict on any one URL.
The clearest number on how bad single-document detection gets comes from the vendor with the most to gain from it working. OpenAI’s own AI text classifier correctly identified 26% of AI-written text as “likely AI-written” while incorrectly labelling human-written text as AI-written 9% of the time [6]. As of 20 July 2023 the classifier is no longer available, due to its low rate of accuracy, with OpenAI saying it was researching more effective provenance techniques for text instead [6].
Take that seriously in practice. A 9% false-positive rate means that if you screen 20 pieces from human freelancers, you should expect to accuse roughly two of them wrongly. There is no accompanying number telling you which two. So do not use a detector score as evidence in a conversation with a writer, a client, or yourself. If the question is “did a machine draft this”, you cannot answer it from the text. Change the question to one you can answer from the text: are the claims in it true, are the sources real, and does it say anything that could only have come from this writer.
Search engines grade the page, not the process
The rule that matters for distribution is not about authorship at all. Google’s spam policies define scaled content abuse as when many pages are generated for the primary purpose of manipulating search rankings and not helping users, and list “using generative AI tools or other similar tools to generate many pages without adding value for users” as one form of it [2]. The trigger is scale plus absence of value, not the tool.
Google’s guidance on people-first content says the same thing from the other direction: if you use automation, including AI-generation, to produce content for the primary purpose of manipulating search rankings, that is a violation of the spam policies [3]. The self-assessment questions on that page are the most useful free editing checklist available, and two of them cut deepest for AI-assisted work. Does the content provide original information, reporting, research, or analysis. Is the content mass-produced by or outsourced to a large number of creators [3].
The same page is where Google puts the who, how and why framework. It asks whether it is self-evident to your visitors who authored your content, whether the use of automation including AI-generation is self-evident to visitors, and whether you are creating content primarily to help people rather than to attract search traffic [3]. That is a production standard you can hold yourself to without anyone measuring your prose.
It is also worth deleting one popular anxiety. Google’s own documentation states that there are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary, and that you do not need to create new machine readable files, AI text files, or markup to appear in these features [4]. A page must be indexed and eligible to be shown in Google Search with a snippet [4]. Anyone selling you a separate “AI answer engine optimisation” package is selling you the work you were already meant to be doing.
Disclosure moved from courtesy to obligation
Article 50 of the EU AI Act came into force on 2 August 2026 [5]. Two parts of it affect anyone publishing text. Providers of systems generating synthetic content must ensure the outputs are marked in a machine-readable format and detectable as artificially generated or manipulated [5]. That obligation sits on the model vendors, not on you, but it is the direction of travel for provenance metadata on the pages you publish.
The part that sits on you is the deployer obligation. Deployers of systems generating text which is published with the purpose of informing the public on matters of public interest must disclose that the text has been artificially generated or manipulated [5]. There is a carve-out: the requirement does not apply where the AI-generated content has undergone a process of human review or editorial control and where a natural or legal person holds editorial responsibility for the publication [5].
Read that carve-out closely, because it is the operating model, not a loophole. A named human being takes responsibility for what ships. If you can point at a person who read every line, checked every number, and would answer for a mistake, you are inside the exemption and you are also, not coincidentally, running the process Google’s guidance describes [3]. If you cannot point at that person, you are publishing undisclosed synthetic text into a market where disclosure is now law in one of your larger jurisdictions.
Drafting is nearly free, so checking is the entire job
The economics are why the Pew number looks the way it does. Claude Sonnet 5 costs $2 per million input tokens and $10 per million output tokens; Claude Opus 5 costs $5 and $25 [7]. At OpenAI, gpt-5.6-terra is $2 per million input and $12 per million output at short context, and gpt-5.6-sol is $4 and $20 [8]. At the Sonnet 5 rate, 10,000 output tokens cost 10 cents [7]. The marginal cost of the words has collapsed to something you cannot feel on an invoice.
Nothing has happened to the cost of the other half. Confirming that a statistic exists, that a price is current, that a quote was said by the person it is attributed to, and that a linked page still resolves takes a human a few minutes per claim, and that cost scales linearly with how much you publish. This is the actual constraint on a small publisher, and it is the thing to budget for explicitly rather than discover at the end of a month.
Two habits make the checking cheaper without making it fake. First, draft against pasted primary material rather than against the model’s memory, so verification is comparison rather than research. Second, cap your publishing volume at your verification capacity rather than at your drafting capacity, which is now effectively unlimited. If the number below comes out larger than the time you have, publish fewer pages.
The parts a model cannot supply are the parts that get cited
Strip a generic AI-assisted page down and you find claims with no owner. A percentage with no study attached, a price with no date, a best practice attributed to nobody. The model produced those because the request did not contain anything more specific, and it will produce them again next time.
The material that survives contact with a sceptical reader, and with an answer engine looking for something quotable, is material that could not have been generated: a number you measured in your own business, a source you read and linked with the date you read it, a method described precisely enough to be repeated, a case where the obvious advice failed and you say why. Google’s AI Overviews and AI Mode show supporting links, and a page has to be indexed and snippet-eligible to be one of them [4]. Being the page that carries the checkable specific is how you end up as the link rather than as one more restatement.
None of this requires you to write by hand. It requires you to supply the specifics before drafting and to verify them afterwards. The draft in the middle is the cheap part, and after Pew’s numbers, the least distinctive part.
pages × claims × minutes, converted to hours. Computed in the page; nothing is sent anywhere.
What still goes wrong
The headline number is softer than it sounds. Pew’s study covers English-language pages only, and rests on one detection model whose per-document judgements Pew itself describes as imperfect [1]. The one-third figure is a defensible estimate of a trend, not a census. Treat anyone who quotes it as a precise fact, including anyone quoting it back at you from this page, as quoting an estimate.
Disclosure is genuinely unsettled. Article 50’s editorial-responsibility carve-out has not been tested by any regulator, and its scope depends on what counts as text published to inform the public on matters of public interest [5]. A marketing page for your own service is probably outside that; a page explaining a change in the law probably is not. If your exposure is real, this is a question for a lawyer in your jurisdiction, not for a guide.
The uncomfortable part is that doing all of this correctly does not guarantee traffic. Verification is a cost you pay whether or not the page performs, and a well-sourced page can lose to a worse one for reasons that have nothing to do with either. What the discipline buys you is narrower and more durable: pages that stay true after the model changes, that you can defend when someone challenges them, and that a person or a machine looking for a source can actually use. On a web where a third of new pages were drafted by something that cannot check its own work, that is a smaller club than it used to be.
- 01Pew Research Center — How Much of the Internet Is Written With AI?pewresearch.org
- 02Google Search Central — Spam policies for Google web searchdevelopers.google.com
- 03Google Search Central — Creating helpful, reliable, people-first contentdevelopers.google.com
- 04Google Search Central — AI features and your websitedevelopers.google.com
- 05EU AI Act — Article 50, transparency obligations (Regulation (EU) 2024/1689)artificialintelligenceact.eu
- 06OpenAI — New AI classifier for indicating AI-written textopenai.com
- 07Anthropic — Claude pricingplatform.claude.com
- 08OpenAI — API pricingdevelopers.openai.com