Three ways to ask what a gene does
Ask a molecular biologist how to find out what a gene does, and the honest answer is that it depends which kind of evidence the question needs, and that no single method supplies all three kinds at once. Three broad approaches currently do most of the work, and they trade against each other on the same three axes every time: how directly the result licenses a causal claim, how many genes or variants can be tested per experiment, and how much money and time the experiment costs.
Classical genetics — knocking a gene out or knocking its expression down, one gene at a time, usually in a whole organism or a single cell line — has been the field’s gold standard for causal necessity since long before genome sequencing existed. CRISPR-based functional genomics screens parallelize the same logic across thousands of genes at once inside a pool of cells, trading some of the per-gene rigor of a hand-built knockout for enormous throughput. Genome-wide association studies (GWAS) invert the whole strategy: instead of perturbing anything, they mine variation that already exists across a large human population and look for statistical association between a genetic variant and a trait. This article compares the three head to head, and does the same for how each interacts with gene regulation, genome variation, epigenetic marks, protein-interaction networks, and the single-cell readouts that increasingly sit downstream of all three. The recurring theme is not which approach is best. It is that intervention and observation answer different questions, throughput and rigor trade against each other within intervention, and treating an association as though it were an intervention is the most common and most avoidable interpretive error in the field.
Classical genetics: strong causal claims, one gene at a time
A knockout — deleting a gene entirely, usually in a mouse or a cell line — or a knockdown — reducing its expression via RNA interference or an antisense oligonucleotide without deleting the DNA — both test the same logical structure: remove or reduce the thing, and see what breaks. Because the intervention is applied directly by the experimenter rather than inferred from a natural population, the causal chain from genotype to phenotype is about as direct as biology gets. If a mouse lacking a gene reliably develops a specific defect, and reintroducing the gene rescues it, that is about as strong a causal claim as the field can produce for a single gene.
The price is throughput. A single knockout mouse line takes months to generate and characterize, and a knockdown experiment, while faster, still tests one gene (or a small handful) per experiment. Testing a genome’s worth of roughly twenty thousand human protein-coding genes one at a time this way is not realistically feasible for any single lab, and it never has been — large-scale mouse knockout consortia exist precisely because no single lab could do it alone. Classical genetics also runs into biological confounds that a screen or an association study can sidestep differently: a gene knocked out from conception can trigger developmental compensation, where other genes partially take over its role, so the adult phenotype can understate the gene’s normal function. Conditional and inducible knockout systems exist specifically to remove a gene only in a particular tissue or at a particular time, which is more work but avoids that compensation problem.
What classical genetics cannot easily do is scale to whole networks. It is the right tool when the question is “does this one specific gene, in this one specific context, cause this one specific phenotype” and the answer needs to be as airtight as an experiment can make it. It is the wrong tool for surveying which genes, among thousands of candidates, matter at all for a given process.
CRISPR screens: causal breadth at population-genomics speed
The 2012 demonstration that Cas9 could be programmed with a single guide RNA to cut a specified DNA sequence [1] made it possible to generate a targeted double-strand break, and therefore a loss-of-function mutation, at almost any genomic locus on demand. Two years later, pooled genome-scale screening turned that single-locus tool into a population-scale one: a library of guide RNAs targeting thousands of genes is packaged into a lentiviral pool, delivered to cells at a low multiplicity of infection so that (with high probability) each cell receives only one guide, and the population is then subjected to a selective pressure — a drug, an infection, a sort on a fluorescent reporter — after which the surviving or selected cells’ guide identities are read out by sequencing [2]. The logic is still “remove the gene, see what breaks,” but it now runs in parallel across the genome inside one pooled population of cells rather than one animal at a time.
This buys enormous throughput at some real cost to precision. A pooled screen’s signal per gene depends on how well its guides work: an inefficient or off-target guide produces a false negative or a spurious hit regardless of the gene’s true role, which is why guide-design rules that score predicted on-target cutting efficiency and predicted off-target risk are now a standard, load-bearing part of screen design rather than an afterthought [3]. Library representation is the other constraint that a screen designer cannot skip: if the pool of cells transduced per guide is too small, a real hit can be lost to sampling noise before selection even begins, which is one reason screens are typically run with hundreds of cells per guide rather than the bare minimum. None of this makes a screen’s causal claim as airtight, gene by gene, as a validated knockout mouse — pooled screens are read out at population scale and a top hit is routinely confirmed afterward with an individual, arrayed knockout precisely because the pooled result is a strong candidate, not a certainty.
What a screen buys back for that loss of individual rigor is scope. A single well-designed genome-scale screen can rank essentially every protein-coding gene for its contribution to one phenotype in one experiment, at a cost and turnaround measured in weeks rather than the years a knockout consortium would need to test the same number of genes one at a time. That is the throughput axis moving hard in the screen’s favor, purchased by giving up some of the per-gene certainty classical genetics offers.
A useful way to see the tradeoff quantitatively is coverage. If a screen library contains
where
GWAS: no perturbation at all, and a different kind of evidence
A genome-wide association study asks a structurally different question. Instead of perturbing anything, it genotypes a large number of individuals who already differ in a trait of interest — a disease, a blood marker, a behavioral measure — and tests millions of naturally occurring genetic variants for statistical association with that trait. Ten years after the method’s early results, a field-wide retrospective catalogued both its productivity and its limits: thousands of robustly replicated variant-trait associations, a substantial fraction of trait heritability still not captured by identified variants, and translation into mechanism or therapy that lags well behind the pace of discovery [4]. Reference resources like the 1000 Genomes Project, which catalogued tens of millions of variants across twenty-six populations with phased haplotypes [5], are what make a GWAS statistically tractable at all — they supply the background variation against which a study’s own genotyped variants can be imputed and interpreted.
GWAS’s central strength is scale of observation rather than scale of intervention: it can survey genetic contributions to a trait across tens or hundreds of thousands of already-living people, capturing variation that includes rare regulatory alleles, gene-gene interactions, and gene-environment interactions that no laboratory perturbation study could economically stage. Its central weakness is exactly what that strength gives up. An associated variant is, by construction, correlational: it says a particular allele co-occurs with a trait more often than chance in this sampled population, not that the allele causes the trait through the nearest gene, or through any gene at all. Population stratification, linkage disequilibrium that ties an associated variant to a nearby causal one without identifying which is which, and reverse causation in traits that are themselves influenced by disease all sit between an association and a mechanism. Most GWAS hits also fall outside protein-coding sequence, in regulatory DNA, which is one reason resources that map the regulatory genome — chromatin state, transcription-factor binding, and open-chromatin regions across cell types [6] — and resources that map how genetic variants change gene expression across tissues [8] have become necessary companions to a GWAS hit rather than optional extras: without them, a variant sitting in a non-coding region has no obvious gene to blame.
On cost, GWAS is the cheapest of the three per data point once population cohorts and biobanks exist, because the genetic variation being studied already occurred for free across evolutionary and demographic history; the study’s expense is almost entirely in recruiting, phenotyping, and genotyping enough already-existing people to detect a small effect against the noise of a diverse genome and a diverse population, not in creating any biological material. That is a genuinely different cost structure from a screen or a knockout, where the expense is in generating the perturbation itself.
Gene regulation and genome variation seen through three lenses
Gene regulation illustrates the complementarity cleanly. A knockout of a transcription factor answers whether that factor is necessary for a target gene’s expression in the specific tissue and developmental window tested. A CRISPR screen using a catalytically dead Cas9 fused to an activator or repressor domain (CRISPRa or CRISPRi, rather than a cutting Cas9) can test, across thousands of regulatory elements or genes at once, which of them causally raises or lowers expression of a reporter or an endogenous transcript when dialed up or down — a scale no knockout campaign could match, still built on the causal logic of directly perturbing the system rather than observing it. GWAS approaches regulation from the opposite direction: it finds a naturally occurring variant statistically associated with a trait, and only afterward, using resources like ENCODE’s map of regulatory elements [6] and GTEx’s map of tissue-specific expression quantitative trait loci [8], can a candidate regulatory mechanism be proposed for it. The GWAS route can be genome-wide in its search but only correlational in its conclusion; the screen route is causal but has to be told in advance which elements or genes to test; the knockout route is causal and confirmatory but slow and narrow. None of the three, alone, closes the loop from variant to mechanism to phenotype; in practice a well-supported regulatory claim usually needs evidence from at least two of the three.
Genome variation itself is characterized almost entirely by observational methods — population sequencing projects like 1000 Genomes exist to catalogue what variation already exists in human populations [5] — and then handed to intervention-based methods for functional follow-up. A variant flagged as common, rare, or population-differentiated by a reference panel is a hypothesis about which position in the genome might matter; whether it actually changes anything is a question for a reporter assay, a base-edited cell line, or a CRISPR screen built around that specific variant, not for the population catalogue itself.
Epigenetics and protein networks: where perturbation and observation blur together
Epigenetic marks — DNA methylation, histone modifications, chromatin accessibility — sit awkwardly between the observational and interventional camps, because they are themselves regulatory states that can be either measured or manipulated. Large observational consortia map which marks sit where across the genome in which cell types, essentially treating chromatin state the way GWAS treats DNA sequence variation: a correlational catalogue that needs a perturbation to test causality [6]. But those same marks can also be the direct target of an intervention: CRISPR-based epigenome editing tools can add or remove a specific methylation mark or histone modification at a chosen locus and ask, causally, whether that specific mark is sufficient to change expression — which is the same logical move a knockout screen makes, just aimed at a chromatin mark instead of a coding sequence.
Protein-interaction networks present a similar blend. A yeast two-hybrid screen or a mass-spectrometry pulldown observes which proteins physically associate inside a cell, which is itself correlational evidence about a network’s structure — co-occurrence in a purified complex does not by itself prove that the interaction matters functionally in vivo. Establishing that an interaction is functionally important, rather than merely present, again typically requires a perturbation: mutating the interface, or knocking out one partner and watching whether the other’s downstream activity changes. The network-mapping step and the causal-testing step are complementary in exactly the same way GWAS and a validation screen are complementary — one surveys structure at scale, the other tests whether a specific piece of that structure actually does the causal work claimed for it.
Single-cell methods: a shared readout layer
The most significant recent development cutting across all three approaches is single-cell sequencing, because it changes what “phenotype” means in a screen or a population study. A pooled CRISPR screen historically read out a coarse phenotype — cell survival, a sorted fluorescence bin — that collapses an entire cell’s response down to one number. Perturb-seq couples pooled CRISPR perturbation to single-cell RNA sequencing so that each individual cell’s guide identity and its full transcriptome are captured together; a genome-scale implementation of this approach profiled the transcriptional consequences of knocking down essentially every expressed gene, one cell at a time, across millions of cells [7]. That turns a screen’s output from a single selectable readout into a rich, gene-expression-level phenotype for every perturbation, closing much of the gap between a screen’s throughput and a knockout’s depth of characterization — though not the gap in per-perturbation certainty, since a single-cell dropout, doublet, or low-count profile still injects its own measurement noise into the picture, distinct from the guide-efficiency and library-coverage noise discussed earlier.
This same single-cell layer increasingly connects back to the observational side of the field as well: reference atlases that catalogue which genes are expressed, and by which cell types, become the annotation layer against which both a screen’s transcriptional hits and a GWAS variant’s candidate target cell type get interpreted — a GWAS hit that only makes expression sense in one rare cell type is a very different kind of finding than one expressed everywhere, and only a single-cell-resolution reference can distinguish the two. In that sense single-cell sequencing is not a fourth competing approach; it is a shared measurement technology that both the perturbation side and the observational side of the field now route through, in the same way that a benchtop sequencer, rather than a specialized instrument for each experiment type, has become the common endpoint for a knockout’s RNA profiling, a screen’s guide counting, and a GWAS cohort’s genotyping.
Causal strength, throughput, and cost, compared directly
Laid side by side on the three dimensions this article opened with, the tradeoffs are consistent rather than incidental.
On causal strength, classical knockouts and knockdowns supply the most direct evidence of necessity for a single gene in a specific, well-controlled context, because the experimenter applies the perturbation and observes the consequence directly. CRISPR screens supply causal evidence too — a hit gene really was perturbed and really did change the measured phenotype — but at a coarser per-gene confidence, contingent on guide quality and adequate library coverage, which is why top screen hits are routinely followed up with an individual arrayed validation. GWAS supplies the weakest direct causal claim of the three by design: an association is evidence consistent with causation, not evidence of it, and closing that gap requires exactly the kind of functional follow-up — a knockout, a CRISPR-based reporter assay, an expression-quantitative-trait-loci lookup — that the other two approaches specialize in providing.
On throughput, the ranking essentially reverses. A knockout campaign tests genes in the single digits to low hundreds per year across a well-resourced program; a genome-scale CRISPR screen tests the entire protein-coding genome, tens of thousands of genes, in one experiment lasting weeks; a GWAS, because it perturbs nothing and instead surveys variation that already exists, can in principle test millions of genetic variants against a trait in a single analysis run, limited mainly by cohort size and genotyping cost rather than by any experimental step that has to be repeated per variant.
On cost, the ordering follows the same logic but is worth stating plainly because it is often assumed rather than checked. Classical genetics is the most expensive per gene tested, because generating and phenotyping a single validated knockout — animal husbandry, breeding, genotyping, and characterization — is itself a multi-month-to-multi-year undertaking. A pooled screen amortizes most of its cost (cloning a library, generating a stable cell pool) across thousands of genes at once, so its cost per gene tested is far lower even though its absolute cost per experiment can be substantial. GWAS’s expense is front-loaded into cohort recruitment and phenotyping rather than into the genetic analysis itself, which is why once a large biobank-scale cohort exists, testing an additional trait against the same genotyped variants is comparatively cheap — the marginal cost of one more association test is far below the cost of generating one more knockout mouse.
None of this yields a single winner, because the three approaches are answering different questions rather than competing to answer the same one. A knockout is the right tool when a specific causal claim about a specific gene needs to be as certain as an experiment can make it. A screen is the right tool when the question is which genes, among many candidates, matter for a process, with follow-up validation expected for whichever ones come out on top. GWAS is the right tool when the question is which genetic variation matters across a real human population, with the understanding that a hit is a lead for causal work, not causal work’s substitute.
Reading a result in context, not in isolation
The practical discipline this comparison points to is straightforward to state and easy to violate under publication pressure: identify, for any given genomics result, which of the three evidentiary regimes produced it, and do not let the language used to describe the result imply a stronger claim than that regime supports. A GWAS hit reported as “the gene for” a trait has already overstated an association as a mechanism. A screen hit reported without arrayed validation, or without acknowledging that guide efficiency and library coverage set a real floor on what the screen could have detected, understates the uncertainty in a population-scale pooled result. A knockout’s phenotype generalized to human disease without noting the developmental-compensation and species-difference caveats overstates what a single, tightly controlled animal experiment can license.
The field’s actual output on any well-studied gene today is rarely built from one of these three approaches alone; it is built from a GWAS or population survey nominating a candidate, a screen testing that candidate alongside thousands of others for a specific cellular phenotype, and a knockout or targeted perturbation confirming the causal claim in a controlled system, all now increasingly stitched together by single-cell readouts that let a transcriptional signature from a screen be compared directly against a transcriptional signature from a natural population’s tissue samples. Treating that pipeline as three complementary layers of evidence, rather than three competing methods with one correct answer, is the discipline the comparison itself recommends.