Equation 58 · A Chatbot Confessed to Being Built by a Company That Never Trained It
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
This is where the strongest objection to the entire construction has to be stated plainly, because no cleverer statistic makes it go away. By 2024, a large and growing share of the public internet — benchmark leaderboards, scraped forum answers, resold instruction datasets, entire sites of AI-generated filler — already consisted of text some earlier model had produced, and recursive training on a mixture of real and machine-generated content is documented to change what later models can represent at all, not merely what they happen to repeat [ 10 ] . A closely related line of work modeling “self-consuming” training loops, in which each generation of models trains partly on the previous…
Read the full surrounding passage
This is where the strongest objection to the entire construction has to be stated plainly, because no cleverer statistic makes it go away. By 2024, a large and growing share of the public internet — benchmark leaderboards, scraped forum answers, resold instruction datasets, entire sites of AI-generated filler — already consisted of text some earlier model had produced, and recursive training on a mixture of real and machine-generated content is documented to change what later models can represent at all, not merely what they happen to repeat [ 10 ] . A closely related line of work modeling “self-consuming” training loops, in which each generation of models trains partly on the previous generation’s synthetic output, shows that even a modest, realistic fraction of synthetic data recirculating through a training pipeline measurably shifts a model’s output distribution across successive generations [ 11 ] . If GPT-4’s outputs had already diffused broadly enough through public data by the time DeepSeek V3 was trained — through resold instruction sets, leaked chat logs, or simply other labs’ own GPT-4-derived synthetic data folded into a shared training pool — then no clean -class model may exist for this specific trait anywhere in 2024 or 2025, because the “unexposed” background rate a fair test needs may itself already be contaminated. On this account, DeepSeek V3’s self-identifications are not evidence of a traceable channel between two companies; they are evidence that the ambient concentration of one company’s synthetic exhaust in the shared training commons had, by the time anyone thought to test for it, already become high enough to leave a mark on a stranger who never went looking for it.
Sources cited in the surrounding passage
- [10] AI models collapse when trained on recursively generated data ↗
- [11] Self-Consuming Generative Models Go MAD ↗
These citations give research context. Read each source to check which claims it supports.
Return to A Chatbot Confessed to Being Built by a Company That Never Trained It