Five Times Out of Eight, the Wrong Name Came Out
In December 2024, a researcher testing DeepSeek’s newly released V3 model asked it, plainly, who had built it. In five of eight independent generations it answered that it was ChatGPT, built by OpenAI; in the other three it correctly named itself [1]. Pushed further, it offered to explain “OpenAI’s API” when asked about its own, and on at least one occasion told the identical joke GPT-4 tells, down to the punchline [1]. DeepSeek is a Chinese company with no disclosed access to OpenAI’s weights, no shared training infrastructure, and no corporate relationship to OpenAI that any org chart would recognize. And yet, for a short window after its release, a stranger’s product kept introducing itself with somebody else’s name.
Two explanations compete for that fact, and the gap between them is this article’s actual subject. The boring explanation is contamination: DeepSeek’s training corpus, scraped in large part from the same public internet everyone scrapes, already contained a great deal of GPT-4 output — logged conversations, benchmark completions, other labs’ synthetic datasets built using GPT-4 as a generator — and the model simply memorized some of it, the way any large model memorizes any sufficiently repeated string. The provocative explanation is transmission: that a specific, identifiable behavior belonging to one company’s model became a stable feature of an entirely unrelated company’s model through some traceable channel, without either company’s engineers designing it that way and without a line of code being copied. Both explanations agree on the observation. They disagree about whether an origin can be assigned, and assigning an origin — reproducibly, on evidence rather than on confidence — is the operation this article wants to make measurable.
The stakes are not confined to one embarrassing chatbot. Every model a major lab ships now carries a retirement date. OpenAI’s own deprecation registry lists exact shutdown dates for dozens of model snapshots: text-davinci-003 went dark on January 4, 2024, and gpt-4-0613, gpt-4-turbo, gpt-4o-2024-05-13, and several other GPT-4 and GPT-3.5 snapshots are scheduled to go dark on October 23, 2026 — weeks from the date this sentence was written [4]. A shutdown notice certifies exactly one fact: a set of weights will stop answering API calls after a stated date. It certifies nothing about whether any behavior those weights produced — a phrasing habit, a refusal pattern, a way of naming itself — is still circulating in systems that never held a copy of them. The industry keeps excellent books on when a model dies. It keeps none at all on what, if anything, outlives it.
A Trait Has to Be Defined Before It Can Be Traced
Before asking whether DeepSeek V3 inherited anything, the word inherited needs an operational meaning that does not depend on which side of the dispute one already believes. Call a trait
A trait needs a case definition: an automatable rubric, applied to a fixed, held-out probe set
Once a trait is defined, a population of models can be sorted by how they could plausibly have acquired it. Call a model vertically exposed if it shares a disclosed training relationship with the artifact’s source — a direct fine-tune of the source’s own weights, or a disclosed distillation run using the source’s outputs as training targets, the same mechanism named a decade ago as knowledge distillation, in which a smaller model is trained to match a larger one’s output distribution rather than raw labels [9]. Call a model horizontally exposed if it shares no weights or disclosed training relationship with the source at all, and its only plausible contact with the trait runs through an artifact the source produced — a dataset built from its outputs, a prompt template copied from its documentation, a piece of scaffold code encoding its behavior. Call a model unexposed, written
A Transfer Rate Turns “It Happened Again” Into a Number
Sort a population of models
Define the channel-attributable transfer rate for exposure class
A transfer rate above zero says a channel is associated with a trait. It does not yet say whether the trait is a faithful copy or a garbled one, and it says nothing about whether curators along the way made the trait more or less likely to survive into the next training corpus. Two further quantities complete the construction. The mutation rate
Epidemiology Lends Its Vocabulary and Immediately Complains About the Loan
The borrowing above is not decorative. Dan Sperber’s “epidemiology of representations” argued that a mental representation spreading through a population is better modeled as a chain of alternating public productions and private re-representations than as a single replicating code: “a human population is inhabited by a much wider population of mental representations,” connected by “complex causal chains where mental representations and public productions alternate” [13]. That is a close description of what this article’s transfer rate measures: a rate of spread through contact, not a rate of copying through a shared germline. The mapped part of the analogy is real and does useful work: a “case” (a trait-positive output), a “host” (one model instance), an “exposure” (contact with a training artifact), and an “attack rate” (
The unmapped part is where the borrowing has to stop paying itself off in cheap insight. A pathogen replicates itself with a copying mechanism no host designed; every transmission event this article examines runs through a training pipeline a human team built, tuned, and could in principle audit line by line, which makes “outbreak” a metaphor for an engineered process rather than a claim about autonomous contagion. A population has an immune system a pathogen must evade; a training corpus has no analogous filter unless a curation team builds one, which is exactly what the selection coefficient
A Made-Up Population Shows What the Numbers Would Look Like
No sealed trial of the kind proposed later in this article has been run, and every number in this section is invented to show what the construction predicts, not what it has found. Suppose a population of
Suppose further that among the trait-positive horizontally exposed models, 70 percent produce a response diverging from the seeded canonical form beyond the declared threshold,
One Real Transfer Was Disclosed, Measured, and Never Disputed
Set the invented population aside and look at the one publicly disclosed case that maps cleanly onto vertical exposure. DeepSeek’s own account of its R1 model, published in Nature after peer review, describes taking 800,000 samples curated from R1’s own reasoning traces and using them to directly fine-tune open checkpoints from two other companies’ model families: Qwen2.5 variants from 1.5 billion to 32 billion parameters, and Llama-3.1 and Llama-3.3 checkpoints from Meta, applying supervised fine-tuning only, with no reinforcement-learning stage added to the distilled models [7]. The resulting DeepSeek-R1-Distill-Qwen-7B model scored 55.5 percent on AIME 2024 and 92.8 percent on MATH-500; the 32-billion-parameter version scored 72.6 percent and 94.3 percent on the same two benchmarks, and the paper reports that even its smaller distilled models outperformed non-reasoning models such as GPT-4o on these evaluations [7]. Qwen2.5 and Llama-3 share no weights with DeepSeek-R1’s own architecture, a sparse mixture-of-experts design documented separately in DeepSeek’s technical report for the related V3 model [8]; the transfer ran entirely through the 800,000 training examples, not through any shared parameter inheritance.
By this article’s classification, this is about as clean a case of vertical, channel-disclosed exposure as currently exists in public: the source, the artifact, the recipient checkpoints, and the training procedure are all named in a peer-reviewed paper, so there is no attribution problem left to solve. A case definition built around a trait specific to R1’s reasoning style — a distinctive self-verification phrasing, a particular pattern of restarting a solution attempt mid-chain — would presumably return a high
The Contested Case Is the One That Actually Needs the Test
Return to the scene this article opened with. Weeks after the self-identification reports, White House AI czar David Sacks said in a televised interview that there was “substantial evidence” DeepSeek had used distillation on a competitor’s model; an OpenAI spokesperson would say only that the company had “technical measures in place to prevent and detect these sorts of attempts,” and the reporting is explicit that “Sacks did not go into detail about the evidence OpenAI had” and that “OpenAI declined to provide details when asked” [2]. A legal analysis of the episode, written for companies weighing their own exposure to the same kind of dispute, could point to no public disclosure of a training relationship comparable to the R1-to-Qwen case, and noted only that a company’s terms of service can in principle prohibit training a competing model on its outputs, leaving open whether such a restriction could ever bind a company that was never a signatory to those terms [3]. More than a year after the fact, “substantial evidence” remains a claim, not a published dataset; nobody outside OpenAI and DeepSeek has been shown the specific artifacts a transfer rate could actually be computed against.
This is where the strongest objection to the entire construction has to be stated plainly, because no cleverer statistic makes it go away. By 2024, a large and growing share of the public internet — benchmark leaderboards, scraped forum answers, resold instruction datasets, entire sites of AI-generated filler — already consisted of text some earlier model had produced, and recursive training on a mixture of real and machine-generated content is documented to change what later models can represent at all, not merely what they happen to repeat [10]. A closely related line of work modeling “self-consuming” training loops, in which each generation of models trains partly on the previous generation’s synthetic output, shows that even a modest, realistic fraction of synthetic data recirculating through a training pipeline measurably shifts a model’s output distribution across successive generations [11]. If GPT-4’s outputs had already diffused broadly enough through public data by the time DeepSeek V3 was trained — through resold instruction sets, leaked chat logs, or simply other labs’ own GPT-4-derived synthetic data folded into a shared training pool — then no clean
That objection does not make the transfer-rate construction useless. It makes clear what the construction can and cannot certify from a single reported incident, however well documented. A newspaper report of a striking coincidence is not a
Three Sealed Lineages Would Turn the Argument Into a Measurement
The discriminating trial this construction actually needs has not been run, and nothing in the two real cases above substitutes for it. Seed one small, isolated teacher model with a single harmless, low-prior trait chosen precisely because no ordinary training objective would produce it independently — for concreteness, a specific and otherwise-arbitrary variable-naming habit inside generated code, invoked only under a narrow, rare combination of task conditions unlikely to arise from generic style transfer. Designing that seeded trait is itself a solved problem in miniature: language-model watermarking already shows how to embed a statistical signal in generated text that is invisible to an ordinary reader but reliably detectable by an algorithm holding the key, by biasing sampling toward a randomized token list without materially changing text quality [12] — exactly the property a good seeded trait needs, conspicuous to the probe set
Build three descendant lineages, each starting from a different base checkpoint with no shared weights with the teacher or with each other. Expose the vertical lineage to direct supervised fine-tuning on the teacher’s outputs, mirroring the disclosed R1-to-Qwen procedure. Expose the horizontal lineage only to secondary artifacts — a released prompt template, a piece of scaffold code, or a third party’s synthetic dataset built from the teacher’s outputs without copying its weights — so that any trait it picks up must have traveled through an artifact rather than through direct training contact. Leave the third, unexposed lineage to train on the same task distribution with a data pipeline audited to exclude every one of the seeded artifacts.
Then do the one thing neither real-world case above allows: decommission the teacher. Delete its weights, retire its serving endpoint, and only afterward run the case-definition probe set
A Named Result Would Prove the Ghost Is Only Noise
The construction is falsifiable in two independent and equally decisive ways, and either one alone is enough to reject it. It fails if, in the sealed trial, neither
A pass would not prove the opposite in the strong sense either. Even a large, well-attributed
A Shutdown Notice Currently Certifies Nothing About What Survives It
If the construction holds up under a real trial, the most immediate consequence is bureaucratic rather than philosophical, and bureaucracies are exactly where it would have to land first. OpenAI’s deprecation registry is a model of clean recordkeeping for exactly one variable: whether a set of weights will still answer an API call after a given date [4]. It was never built to record, and currently could not record, whether a behavior that set of weights was known for is still detectable somewhere else on the day the weights go dark. A deprecation notice for text-davinci-003 in January 2024 said nothing about whether the instruction-following manner Stanford researchers had extracted from it nearly a year earlier — 52,000 demonstrations generated in the self-instruct style specifically to teach an unrelated architecture, Meta’s LLaMA, to imitate that manner [5, 6] — was still shaping descendants of that dataset two years after the source could no longer be queried for comparison. Nobody was obliged to check, because no registry asks the question.
A dataset-provenance audit built on this article’s transfer rate would ask it directly: before a model is retired, run its known distinctive traits against the descendant population it is suspected of having touched, and record
The Model Died on Schedule; the Question It Leaves Behind Did Not
Put the record straight rather than dramatic. One transmission event in this account is disclosed, measured, and undisputed: DeepSeek’s own account of distilling 800,000 of R1’s reasoning traces into Qwen and Llama checkpoints that share no weights with it, published and peer-reviewed rather than alleged [7]. One is reported, striking, and genuinely contested: DeepSeek V3 naming itself after a rival product more often than not, in an episode neither company has fully explained, defended by an unspecified “substantial evidence” on one side and a declined comment on the other [1, 2]. And one thing is true regardless of how that dispute is ever resolved: an entire public internet’s worth of earlier models’ synthetic exhaust now sits inside the training data of nearly everything trained after it, which means the clean, uncontaminated background rate any transfer-rate calculation needs may already be unrecoverable for cases already in the wild — an honest limitation, not a loophole, and the reason this article’s own proposed trial insists on sealing its lineages and killing its teacher model before anyone is allowed to look [10, 11].
What changes if the sealed trial someday runs and the construction survives it is not a verdict on any single company. It is a retirement notice’s missing second half: not just when a set of weights stopped answering, but a number, computed against a stated baseline and a named channel, for whether anything it was known for is still out there answering in its place. Text-davinci-003 has not taken an API call in over two and a half years. Whether any of the 52,000 demonstrations built from its outputs still shapes the manner of some model still being fine-tuned this week is, right now, not a question any registry is built to answer — only a question this construction is built to ask, and, if the sealed trial ever runs and comes back empty, entitled to answer no.