Two lineages, not one

It is tempting to date AI security to the arrival of a single famous tweet in September 2022, when a researcher first showed the world that a language model could be talked out of its instructions in public. That moment matters and this article gives it its due weight. But treating it as an origin misreads the history. Two separate lineages of research and disclosure run into that moment from different directions, one of them nearly a quarter-century old and the other a decade old, and neither began with anything resembling a chatbot.

The first lineage is a story about software architecture: what happens when a system reads something it should have treated as inert, and executes it instead. That story was written down, in detail, in 1998, in a hobbyist security magazine, describing a flaw in ordinary web forms that had nothing to do with artificial intelligence.

The second lineage is a story about perception: what happens when a trained statistical model is shown something a human would call unremarkable, and confidently gets it wrong. That story was written down in 2013, in a paper about image classifiers, that also had nothing to do with chatbots.

ADVERTISEMENT

What follows is an attempt to trace both lineages honestly, mark the point where they collide and get a shared name, and follow the documented, dated chapters that come after: the arms race between jailbreak attempts and safety training, the formation of red-teaming as an institutional discipline inside AI laboratories, the arrival of government guidance written specifically for AI systems, and the recent, still-unfolding turn in which AI tools begin finding real vulnerabilities and defending real networks rather than only creating new risk. Every claim below is attached to a specific, dated, independently checked source, because a history that cannot be checked is not a history — it is a story about a history, and this field already has too many of those.

The vulnerability class that already had a name

On 25 December 1998, a security researcher writing under the handle rain.forest.puppy published an article in issue 54 of Phrack, the long-running hacker magazine, under the plain title “NT Web Technology Vulnerabilities.” Buried in a discussion of Microsoft’s web and database stack was a description of a flaw that would eventually be given its own name and its own decades-long literature: a web form that builds a database query by directly concatenating whatever a user typed, so that a value like rfp' select * from table1 -- does not just fill in a blank in the intended query, it appends a second command that the database dutifully executes alongside the first [1]. The article did not use the term “SQL injection” — that phrase came later, as the security community found a name for what Forristal had shown — but the mechanism described is exactly it: the system had no way to tell the difference between the data a user was supposed to supply and an instruction that data could be made to contain.

A late-1990s web-security research desk with a beige CRT monitor and a dot-matrix printer caught mid-feed drawing a fresh sheet through its platen, a printed zine issue lying open beside the keyboard
Figure 1. The first widely circulated public account of a system executing what it should have treated as inert data was written down in a hobbyist security magazine, not a corporate disclosure or a court filing.

That distinction — between a channel that is supposed to carry only inert values and a channel that can, through an oversight in how a string gets assembled, carry commands instead — is the single idea that this entire article is organised around. It is not a metaphor borrowed from computer security to help explain AI security. It is the same defect, observed for the first time in a different substrate.

The reason this history matters for a history of AI security specifically is that the durable fixes for the 1998 version of the problem were architectural, not detective. Parameterised queries, which became the standard defence taught in every secure-coding curriculum that followed, do not try to inspect a submitted string for malicious content. They instead carry the query template and the user-supplied values down genuinely separate channels to the database engine, so that no arrangement of characters typed into the value field can ever be interpreted as part of the query’s structure. The fix does not get better at recognising attacks. It removes the shared channel the attack depended on.

Keep that fact in view. It is the yardstick against which everything from 2022 onward in this history has to be measured, because — as later sections in this series on AI security document in more technical detail — the equivalent separation does not currently exist inside a language model’s input, and the reasons why are a central, unresolved fact about the field this history describes.

ADVERTISEMENT

A separate discipline: fooling a classifier on purpose

The second lineage begins with a different question, asked by a different research community, in a paper with no obvious relationship to the first. In December 2013, Christian Szegedy and colleagues at Google and NYU circulated “Intriguing Properties of Neural Networks,” reporting two findings about the image classifiers of the day. The first was about how information is represented inside the layers of a trained network. The second, and the one that mattered for security, was a demonstration that these networks’ input-output mappings are “fairly discontinuous” — that an attacker could compute a deliberate, visually imperceptible perturbation to an image and cause a well-trained classifier to misclassify it with high confidence, and that this same perturbation often fooled a second network trained on different data, meaning the vulnerability was not a quirk of one particular model but something closer to a shared property of how these systems generalise [2].

An early-2010s machine-learning lab bench with a beige-and-grey workstation and a printed contact sheet of near-identical small images, one freshly circled in pencil mid-stroke
Figure 2. Years before anyone tried to redirect a chatbot, researchers had already shown that a trained classifier could be made to fail with total confidence by a change too small for a person to see.

The following year, Ian Goodfellow, Jonathon Shlens, and Christian Szegedy went further than documenting the phenomenon; they offered an explanation for it and, in doing so, made it tractable. Their 2014 paper “Explaining and Harnessing Adversarial Examples” argued that the primary cause of this vulnerability was not some exotic nonlinearity in how neural networks compute, but almost the opposite: their vulnerability came from being “too linear” in high-dimensional space, which meant a small, carefully directed nudge to every input pixel could accumulate into a large change in the network’s output. That explanation produced the fast gradient sign method, a simple, cheap recipe for generating an adversarial example in a single step rather than an expensive search, and the same paper showed that training on these adversarial examples could measurably reduce a model’s error rate on them [3].

It is worth sitting with how little these two papers have to do with the 1998 SQL injection disclosure on the surface. Nobody in 2013 or 2014 was talking about instructions, prompts, or data channels. The attack surface was a fixed-size numerical input, a photograph represented as pixel intensities, and the “vulnerability” was a mathematical property of a trained function: it would confidently misclassify an input that was, to any human observer, unchanged. This is why adversarial machine learning is properly understood as a separate research lineage from injection attacks, not a subcategory of them. It concerns the model’s perception, not the model’s instruction-following. For most of the 2010s, the two literatures developed independently, at different conferences, read by different people, and neither anticipated the other would eventually be needed to understand the same production system.

2022: the two lineages collide and get a name

The collision happened because a new kind of AI system arrived that had properties from both worlds at once. Large language models accept a single undifferentiated stream of text as input — meaning, unlike a bounded numeric image tensor, the input channel itself could carry commands as easily as content — and they are trained to generalise their behaviour statistically across that input, the same underlying training paradigm that produced the discontinuities the 2013 and 2014 papers had studied. When these systems were wired into applications that concatenated a developer’s instructions with a user’s untrusted input into one prompt, the 1998 pattern reappeared almost exactly, but now inside a system with no parameterised-query equivalent to fall back on.

The demonstration that made this concrete and public happened in September 2022. On 12 September 2022, Simon Willison published a post titled “Prompt injection attacks against GPT-3,” crediting the discovery to a demonstration by fellow researcher Riley Goodside days earlier and proposing a name for the pattern: “This isn’t just an interesting academic trick: it’s a form of security exploit. I propose that the obvious name for this should be prompt injection” [4]. Willison’s post explained the mechanism in terms directly continuous with the 1998 lineage: developers were building GPT-3 applications by concatenating untrusted user input directly into a prompt, so a user could supply text that the model would treat as a new instruction overriding the developer’s original one, up to and including leaking the developer’s own hidden prompt back out. The naming stuck immediately, in large part because the analogy to a vulnerability class security engineers already understood made the danger legible on sight.

A 2022 chatbot-research desk with a laptop caught mid-scroll, a spiral notebook of draft prompts with lines crossed out, and a coffee cup set down beside a webcam whose indicator light is lit
Figure 3. The name given to this vulnerability in 2022 credited a discovery made in public, on a live model, days before the term existed to describe it.

The academic literature caught up quickly. On 17 November 2022, Fábio Perez and Ian Ribeiro submitted “Ignore Previous Prompt: Attack Techniques For Language Models” to arXiv, giving the phenomenon its first systematic treatment as a research object rather than a demonstrated trick. Their paper introduced PromptInject, a framework for composing adversarial prompts, and formally separated two attack goals that still structure the field’s vocabulary: goal hijacking, in which an attacker misaligns a prompt’s original objective toward printing an attacker-chosen target phrase, and prompt leaking, in which an attacker instead extracts part or all of the original, supposedly hidden prompt [5]. The paper went on to win a best paper award at the NeurIPS 2022 ML Safety Workshop, a small but telling sign that the machine-learning research community was already prepared to receive this as a security finding worth formal study, rather than dismiss it as a curiosity of chat interfaces.

ADVERTISEMENT

What makes September and November of 2022 a genuine hinge point in this history, rather than merely a convenient naming event, is that the two older lineages had finally been forced into the same sentence. The vulnerability being named was structurally the 1998 one — a failure to separate data from instruction — but it was occurring inside a system whose behaviour was governed by the same statistical, learned, non-symbolic generalisation that the adversarial-examples literature had spent a decade characterising. Prompt injection did not need to be invented from nothing in 2022. It needed the specific kind of system — a general-purpose, natural-language, instruction-following model deployed inside real applications — to exist for the older insight to become newly, urgently applicable.

The arms race, catalogued

Once the vulnerability had a name, exploiting it became a public, competitive, extremely fast-moving pastime, conducted largely on social media through late 2022 and into 2023, as users traded increasingly elaborate role-play and persona-based prompts designed to talk deployed chat models out of their safety training. What changes the character of this period from folklore to documented history is that, within roughly a year, the phenomenon received the same treatment SQL injection had received decades earlier: academic security researchers stopped merely collecting examples and started building a taxonomy of why the attacks worked.

The clearest published account of that shift is Alexander Wei, Nika Haghtalab, and Jacob Steinhardt’s July 2023 paper, “Jailbroken: How Does LLM Safety Training Fail?” The authors proposed two structural failure modes rather than a grab-bag of clever phrasings. The first, competing objectives, arises when a model’s training to be helpful and capable is placed in direct tension with its training to be safe, and an attacker constructs a prompt where the two objectives cannot both be satisfied, betting the model resolves the conflict in the attacker’s favour. The second, mismatched generalization, arises when safety training fails to transfer to a domain where the model’s underlying capabilities still work fully — the model was never taught to refuse in that particular encoding or framing, even though it is perfectly able to comply in it. Guided by these two mechanisms, the authors designed attacks that, by their own account, succeeded on “every prompt in a collection of unsafe requests from the models’ red-teaming evaluation sets” against systems including GPT-4 and Claude v1.3, and they argued explicitly that scaling models further would not, on its own, resolve either failure mode, because safety mechanisms would need to keep pace in sophistication with the underlying model rather than trail behind it [6].

A 2023 red-team lab wall with pinned printed transcripts and a stopwatch caught mid-count, a short server rack behind running a live evaluation batch
Figure 4. By 2023 the contest between jailbreak attempts and safety training had become something academic security researchers timed, catalogued and published, rather than a folklore of clever phrasings.

That last point is the one worth carrying forward, because it reframes what “the arms race” actually is. It is not a contest that a sufficiently clever prompt occasionally wins against an otherwise solid wall. It is two training processes — one optimising for capability, one optimising for refusal — running in the same model and never guaranteed to agree with each other at every possible input, which means the space of disagreement is available to be searched by anyone willing to look for it. That is a structural description, not a temporary state of the art, and it is the reason the following two sections — institutional red-teaming and government guidance — appear in this history at all: once a failure mode is understood to be structural rather than a bug to be patched away, the response has to be organisational as much as technical.

Red-teaming becomes an institution

The practice of deliberately attacking a system before an adversary does is not new to computer security generally, but its formalisation specifically around deployed AI models is dated and traceable. On 19 September 2023, OpenAI announced the OpenAI Red Teaming Network, described in contemporaneous reporting as “a contracted group of experts to help inform the company’s AI model risk assessment and mitigation strategies” — a move from ad hoc, one-off adversarial testing arrangements toward a standing, named programme of external domain experts who could be called on at multiple stages of a model’s development and deployment [7]. The significance of this announcement is less about any single laboratory’s internal practice and more about what it signals had happened to the field’s self-understanding in the year since prompt injection was named: adversarial testing of a language model had gone from something a curious outsider did on a public chat interface to something a laboratory formally staffed, contracted for, and treated as a standing line item in how a model gets shipped.

This is the point in the history where “AI red-teaming” stops being a borrowed phrase from military and penetration-testing jargon and becomes its own recognisable discipline, with its own recruiting pipeline, its own vocabulary of jailbreaks and failure modes drawn from the 2022 and 2023 literature just described, and its own institutional home inside the organisations building the models rather than only among the outside researchers probing them.

Governments write it down

Formal government guidance followed the same year, in a cluster of publications that is easy to treat as a single undifferentiated event but is worth separating into its actual dated pieces, because each documents a different institutional actor arriving at the same conclusion independently.

On 14 November 2023, the US Cybersecurity and Infrastructure Security Agency released its first Roadmap for Artificial Intelligence, organised around five lines of effort spanning CISA’s own responsible use of AI, the assurance of AI systems, protection of critical infrastructure from AI-enabled misuse, international coordination, and workforce expertise — CISA’s leadership framed the document against the backdrop of an executive order on AI governance signed the previous month, with Director Jen Easterly stating that AI “holds immense promise in enhancing our nation’s cybersecurity” while also presenting “enormous risks” as “the most powerful technology of our lifetimes” [10].

Less than two weeks later, on 27 November 2023, the UK’s National Cyber Security Centre, jointly with CISA, the US National Security Agency, the FBI, and cybersecurity authorities from more than a dozen additional countries, published the Guidelines for Secure AI System Development. The document is explicit that it targets “providers of any systems that use artificial intelligence (AI), whether those systems have been created from scratch or built on top of tools and services provided by others,” and it structures its guidance around four phases of a system’s life: secure design, secure development, secure deployment, and secure operation and maintenance [9]. The scale of the co-signing list — described by the releasing agencies as the first document of its kind agreed to across the G7 and a wider international group — marks this as the moment AI-specific security guidance stopped being any single country’s initiative and became a coordinated multinational baseline.

The technical standards-setting side of this institutionalisation has its own, longer paper trail. NIST’s taxonomy of adversarial machine learning attacks and mitigations did not appear fully formed in 2023; it descended from an interagency report first drafted in 2019, was reissued as a public draft in March 2023 for comment, and was published in its first final form as NIST.AI.100-2e2023, “Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations,” on 4 January 2024, authored by Apostol Vassilev, Alina Oprea, Alie Fordyce, and Hyrum Anderson [12]. The document is organised, tellingly, around exactly the two lineages this article has traced: a taxonomy of attacks on predictive, classification-style machine learning systems descending directly from the Szegedy and Goodfellow line of research, and a separate taxonomy of attacks on generative AI systems, including prompt injection, descending from the 2022 line. Seeing both lineages formally catalogued side by side, in one federal standards document, is as close as this history gets to an official confirmation that they are understood as two related but genuinely distinct research traditions that a single field of “AI security” now has to hold together.

The frontier: AI turns to defence

The final chapter of this history so far is the one still being written, and it inverts the story told up to this point. Everything before this section is about AI systems as the thing being attacked. The frontier is the emergence, on a documented and dated basis, of AI systems as tools used to defend and to find vulnerabilities in other software — including, in at least one documented case, software that has nothing to do with AI at all.

The first major production deployment of this idea from a large vendor came early. On 28 March 2023, Microsoft introduced Security Copilot, described by the company as combining a large language model with “a security-specific model from Microsoft” and informed by “more than 65 trillion daily signals” of threat intelligence, aimed at letting security analysts submit natural-language prompts to accelerate incident investigation and threat hunting [8]. Whatever one makes of any single vendor’s own characterisation of its product — and vendor claims should always be read as claims — the announcement date matters here as a marker: a major security vendor judged, only about six months after prompt injection was named as a new class of AI risk, that the same underlying technology was ready to be sold as a defensive tool to the people who would have to manage that risk.

Government-sponsored research followed a more adversarial, competitive path. On 9 August 2023, DARPA announced the AI Cyber Challenge, describing it as pursuing “artificial intelligence-enabled cyber reasoning systems that can automatically find and fix software vulnerabilities in real-time and at scale in widely used, critical code,” structured as a competition with tracks for funded small businesses and self-funded open participants, running through qualifying, semifinal, and final rounds at DEF CON conferences in 2024 and 2025 [11]. AIxCC is worth separating carefully from red-teaming as described earlier in this history: it is not about testing an AI model’s own safety behaviour, but about using AI systems as the tool that finds and patches vulnerabilities in other software — a direct, government-funded, competitive successor to the automated vulnerability-discovery tradition that fuzzing and static analysis had occupied for decades, now with a trained model doing part of the reasoning.

A modern vulnerability-discovery bench with a device under automated test, its single status LED caught turning from unlit to lit, beside a workstation printing a short bug report
Figure 5. The most recent chapter of this history is the first in which the software doing the finding, rather than only the software being found vulnerable, is itself a trained model.

The clearest single data point for that successor claim, and a fitting point to close this history on, is dated 1 November 2024. Google’s Project Zero team, working with Google DeepMind on a tool named Big Sleep, reported finding an exploitable stack buffer underflow in SQLite, a database engine used across a vast range of production software, describing it as “the first public example of an AI agent finding a previously unknown exploitable memory-safety issue in widely used real-world software” [13]. The team’s account is specific about how the tool was directed — it was given the starting point of a previously patched SQLite vulnerability and asked to search for similar, still-unpatched variants in the current codebase — and specific about what conventional tooling had not managed: the same flaw was not found by fuzzing the affected code for 150 CPU-hours. SQLite’s developers were notified and shipped a fix the same day the issue was reported, before any public release shipped the vulnerable code [13].

Read next to the 1998 Phrack article this history opened with, that November 2024 disclosure closes a loop worth naming explicitly. In 1998, a piece of software’s failure to distinguish data from instruction was found and written up by a human researcher, working alone, describing a defect in code that had nothing to do with machine learning. In 2024, an AI system was used to find a memory-safety defect — a different, older class of vulnerability than the one this history has spent most of its time on — in a piece of software that likewise has nothing to do with machine learning. The subject being defended has become, in part, the same kind of system that introduced the newest attack surface in the first place.

What the history actually shows

Laid end to end, the documented chronology resists a single tidy moral, and it is worth resisting the temptation to supply one anyway. But a few honest observations survive close reading of the dated record.

The first is that AI security did not invent a new category of vulnerability out of nothing in 2022; it rediscovered, inside a new and much less structurally defensible substrate, a category that computer security had already named, studied, and partially solved for older systems by the late 1990s. The word “prompt injection” is new. The underlying failure — a system unable to distinguish data it should treat as inert from instructions it should act on — is not, and the fact that the classical fix for that failure, a hard separation of channels, has no clean equivalent inside a language model’s input is the reason this particular chapter of computer security has stayed open for longer than many practitioners initially expected.

The second is that the two founding lineages — injection and adversarial perturbation — remain genuinely distinct bodies of research even after 2022 forced them into contact, and NIST’s own taxonomy keeps them as separate categories rather than merging them, which is a useful corrective against treating “AI security” as one undifferentiated topic.

The third is that the institutional response — red-teaming networks, CISA and NIST guidance, multinational agreements — arrived remarkably quickly by the standards of prior computer-security history, within twelve to eighteen months of the vulnerability class receiving its name, even though the underlying technical problem it responds to remains, by the field’s own published accounts, unresolved rather than closed.

And the fourth is the one this history’s final chapter puts in sharpest relief: the tools now being used to find and fix vulnerabilities, including vulnerabilities entirely unrelated to AI, are themselves increasingly the same kind of system that this whole history is about securing. Whether that turn changes the balance between attackers and defenders, or simply moves the argument onto new and unfamiliar ground, is a question this article leaves open rather than answers — it is a question for the frontier this history has only just reached, not for the history itself.