← Back to article

Equation 3 · A Chatbot Confessed to Being Built by a Company That Never Trained It

What does this equation mean?

∅\varnothing

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Once a trait is defined, a population of models can be sorted by how they could plausibly have acquired it. Call a model vertically exposed if it shares a disclosed training relationship with the artifact’s source — a direct fine-tune of the source’s own weights, or a disclosed distillation run using the source’s outputs as training targets, the same mechanism named a decade ago as knowledge distillation, in which a smaller model is trained to match a larger one’s output distribution rather than raw labels [ 9 ] . Call a model horizontally exposed if it shares no weights or disclosed training relationship with the source at all, and its only plausible contact with the trait runs through an…
Read the full surrounding passage
Once a trait is defined, a population of models can be sorted by how they could plausibly have acquired it. Call a model vertically exposed if it shares a disclosed training relationship with the artifact’s source — a direct fine-tune of the source’s own weights, or a disclosed distillation run using the source’s outputs as training targets, the same mechanism named a decade ago as knowledge distillation, in which a smaller model is trained to match a larger one’s output distribution rather than raw labels [ 9 ] . Call a model horizontally exposed if it shares no weights or disclosed training relationship with the source at all, and its only plausible contact with the trait runs through an artifact the source produced — a dataset built from its outputs, a prompt template copied from its documentation, a piece of scaffold code encoding its behavior. Call a model unexposed , written ∅\varnothing , if no such channel, disclosed or suspected, connects it to the source. DeepSeek-R1’s later distillation into Qwen and Llama checkpoints, discussed below, is vertical by this definition, since the training relationship is disclosed and direct. DeepSeek V3’s self-identification is at best a candidate for horizontal exposure, since no training relationship between DeepSeek and OpenAI has been disclosed by either party — which is exactly why it is the harder case, and exactly why it needs a measurement rather than a press statement.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to A Chatbot Confessed to Being Built by a Company That Never Trained It

See this formula across 4 published contexts →

Browse the mathematical compendium →