← All tips

Spotting bulk-produced copy: checking 400 delivered texts

The invoice is paid, the texts are online, and nobody in the house has read all 400. The question comes weeks later, when the traffic does not pick up. It gets expensive twice over: since 2024 Google treats content produced en masse to the same pattern as “scaled content abuse”, regardless of who wrote it, and since 2 August 2026 Article 50 of the AI Act (EU) 2024/1689 requires disclosure where AI was involved and the text is meant to inform the public.

How to go about it in JMX

  1. Delimit the area. After the crawl, create a segment for the delivery under Settings → Segments: path starts with /advice/. The first matching row applies.
  2. Declare nothing yet. Under Settings → Content origin, the area stays empty for this pass. The series score goes quiet everywhere an origin is declared — immediately and permanently.
  3. The Issues tab, AI disclosure category. The rule “bulk-produced copy with no declared origin” (AI.TEXT.UNDISCLOSED) reports as a notice, from a series score of 60 out of 100 and only from 250 words per page. Next to every URL sits the evidence: “series score 68 of 100 across 412 words (measured in de)”, with the individual signals behind it.
  4. Look at the site as a whole. The same category carries “series pattern: many pages uniform, none declared” (AI.SITE.SCALED_PATTERN) — from 30 percent of flagged content pages and at least 10 flagged pages. One page is taste. A third is a decision.
  5. Separate artifacts from suspicions. In the same list sit leftovers from the chat window — the apology with which a model explains its limits, a placeholder for the company name that was never replaced — and invisible characters in the body text, the leftovers as an error, the characters as a warning. They claim nothing: the sentence is on the page and therefore in the Google index.
  6. Check for templates. The near-duplicate rule compares via SimHash and reports up to Hamming distance 6. Where no sentence is the same, it does not catch — that is what the Semantics view is for, the “Pages that mean the same” view, from a cosine of 0.92. With Ollama that costs nothing, and no page text leaves the machine.
  7. Output the list. “Export” in the Issues view writes CSV or Excel, evidence column included. That is the table that goes to the agency.

What to watch out for

A series score of 72 does not mean “72 percent AI”. What is measured is uniformity: narrow sentence lengths, paragraphs of equal length, formulaic connectives, few checkable numbers, and above all the absence of any long sentence. No method reliably recognises machine-generated text, and a text carefully written to a scheme by a human hand meets the same criteria. That is why the rule sits at notice level. This is not evidence, it is a question.

The edges of the measurement are sharp. It is calibrated for German and English, and stays quiet otherwise. Below 250 words it does not judge at all: short product texts are not unremarkable, they are unchecked — anyone counting 400 rows has to know how many of them were even up for checking. Pages with noindex stay out too. There is no triage view sorted by series score; you filter by category in the issue list.

Plain language, news items and terse product descriptions are deliberately short and uniformly built, and a check that measures uniformity reports them reliably — not a fault of the method but the reason the declaration sits above the estimate. Enter such areas once as “written without AI” — but in that order: measure first, then enter.

The agency will want to land on “just a heuristic”. You do not have to argue: bulk-produced copy stays bulk-produced copy no matter whose hand built it, and Google does not ask about the origin when it evaluates.

TippsKI-Kennzeichnung