Equation 3 · A Chatbot Confessed to Being Built by a Company That Never Trained It
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Once a trait is defined, a population of models can be sorted by how they could plausibly have acquired it. Call a model vertically exposed if it shares a disclosed training relationship with the artifact’s source — a direct fine-tune of the source’s own weights, or a disclosed distillation run using the source’s outputs as training targets, the same mechanism named a decade ago as knowledge distillation, in which a smaller model is trained to match a larger one’s output distribution rather than raw labels [ 9 ] . Call a model horizontally exposed if it shares no weights or disclosed training relationship with the source at all, and its only plausible contact with the trait runs through an…
Read the full surrounding passage
Once a trait is defined, a population of models can be sorted by how they could plausibly have acquired it. Call a model vertically exposed if it shares a disclosed training relationship with the artifact’s source — a direct fine-tune of the source’s own weights, or a disclosed distillation run using the source’s outputs as training targets, the same mechanism named a decade ago as knowledge distillation, in which a smaller model is trained to match a larger one’s output distribution rather than raw labels [ 9 ] . Call a model horizontally exposed if it shares no weights or disclosed training relationship with the source at all, and its only plausible contact with the trait runs through an artifact the source produced — a dataset built from its outputs, a prompt template copied from its documentation, a piece of scaffold code encoding its behavior. Call a model unexposed , written , if no such channel, disclosed or suspected, connects it to the source. DeepSeek-R1’s later distillation into Qwen and Llama checkpoints, discussed below, is vertical by this definition, since the training relationship is disclosed and direct. DeepSeek V3’s self-identification is at best a candidate for horizontal exposure, since no training relationship between DeepSeek and OpenAI has been disclosed by either party — which is exactly why it is the harder case, and exactly why it needs a measurement rather than a press statement.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to A Chatbot Confessed to Being Built by a Company That Never Trained It