Single cells sorted into droplets, a methyl group locking a promoter shut, and a guide RNA proving a gene's job by deleting it — the mechanics behind three claims biology can actually back up.

At the droplet junction, a cell and a single uniquely barcoded bead are about to be sealed into one oil droplet together — the physical step that turns bulk RNA extraction into per-cell resolution. — Image prompt and art direction by Brecht Corbeel; generation pending.
This briefing walks through three concrete mechanisms in modern molecular biology: how droplet-based single-cell RNA sequencing physically captures transcriptional heterogeneity one cell at a time, what DNA methylation does chemically to silence a promoter, and how a genome-scale CRISPR knockout screen turns correlation into a causal claim about gene function. Each section separates verified fact from vendor framing from open interpretive question, grounded in the primary papers that established the methods.
Three claims get made constantly in biology reporting: “single-cell sequencing reveals hidden cell types,” “an epigenetic switch turns a gene off,” and “CRISPR proved this gene causes that disease.” Each rests on specific physical machinery. This briefing opens each one and separates fact from vendor framing, interpretation, and open question.
A bulk RNA extraction pools transcripts from thousands of cells into one tube, so any measurement is already an average and a rare cell is invisible inside it. Droplet-based single-cell RNA sequencing (scRNA-seq) solves this by never letting the pooling happen.
In the method Macosko and colleagues published as Drop-seq, a microfluidic chip meets three streams — cells in buffer, uniquely barcoded beads in lysis buffer, and oil — at a T-junction. Poisson loading statistics govern the encounter: cells and beads are dilute enough that most droplets contain neither, and among droplets that do capture a cell, capturing exactly one bead alongside it is the common outcome [1]. Inside that droplet the cell lyses, its mRNA hybridizes to poly(dT) primers on the bead’s surface, and because every primer on a bead carries the same barcode, every transcript captured there is tagged with one identifying sequence before the droplets are broken. Once barcoding is done, the emulsion is broken and beads are pooled for bulk library prep. Per-cell structure survives the pooling because it was written into the sequence, not because cells stayed physically separated.
Zheng and colleagues’ GemCode platform, underlying commercial 10x Genomics instruments, runs the same logic with a second identifier layered in: each primer also carries a random unique molecular identifier (UMI), so PCR amplification bias can be corrected by collapsing duplicate reads sharing a UMI back to one original molecule. Their instrument processes tens of thousands of cells per sample in minutes, at roughly 50% capture efficiency [2]. A “cell” in the resulting count matrix is not something watched under a scope — it is a barcode that recruited a plausible number of reads.
Fact: droplet encapsulation plus per-bead barcoding converts a pooling extraction into a per-cell-resolved dataset, demonstrated at tens of thousands of cells per run [1] [2]. Vendor-style overclaim: “single-cell resolution” is sometimes marketed as seeing every transcript in every cell; capture efficiency is well under 100%, so a gene’s true absence and true expression below detection are indistinguishable in the matrix alone — called “dropout.” Analysis: the barcode carries the whole epistemic burden; if lysis or emulsion breaking lets transcripts drift into the wrong droplet, the error is invisible unless independently checked, which is why doublet-detection and empty-droplet filtering are now standard steps on top of the base chemistry.
Where single-cell sequencing answers “which genes are on, in which cell,” DNA methylation is one mechanism that decides the answer for a given cell type — worth being precise about, rather than treating “epigenetic switch” as a black box.

Figure 1. Bisulfite conversion chemistry, run here in a thermocycler, is what lets a methylated cytosine be told apart from an unmethylated one by sequencing afterward. — Image prompt and art direction by Brecht Corbeel; generation pending.
In mammalian genomes, DNA methyltransferase enzymes add a methyl group to cytosine, almost always where it is followed by a guanine — a CpG dinucleotide — producing 5-methylcytosine. CpG dinucleotides cluster in stretches called CpG islands, many sitting at gene promoters. Jones’s review separates two mechanisms often conflated in casual writing. First, a methylated cytosine projects into the DNA major groove and can directly obstruct transcription factors whose recognition motif overlaps it. Second, and shown to matter more broadly, methylated CpGs recruit methyl-CpG-binding domain proteins, which recruit chromatin-remodeling complexes that compact local chromatin into a configuration inaccessible to transcription machinery generally, not only to the one blocked factor [3]. Promoter methylation is therefore usually a stable, heritable repression signal rather than a fast toggle: it is copied to the new strand after replication by maintenance methyltransferases recognizing hemimethylated CpGs — the physical basis for the mark persisting across divisions without new signal each time.
The same review notes methylation is not uniformly repressive by position: methylation within an actively transcribed gene body correlates with, and in some contexts is required for, elongation and alternative splicing, while methylation of repetitive elements suppresses their transcription as a genome-defense function distinct from promoter silencing [3]. One word — “methylation” — covers several distinct regulatory jobs depending on where in the gene it sits.
Fact: promoter CpG methylation is chemically linked to stable repression through direct steric interference and, more consequentially, recruitment of repressive chromatin complexes [3]. Analysis: because the mark is copied at replication, it functions as cellular memory — how a liver cell’s daughters stay liver cells without re-reading regulatory logic each division — but a report of only “gene X is hypermethylated,” without saying where (promoter, body, or repeat), has not yet said whether that predicts repression, activation, or nothing. Prediction, with horizon and disconfirmation: within five to ten years, clinical methylation panels for cancer subtyping likely shift toward whole-region pattern classifiers, since single-site changes show weaker outcome association than block-level patterns; disconfirmed if approvals keep favoring single-marker assays regardless of added value.
Neither mechanism above can, by itself, say a gene’s expression causes a phenotype rather than merely accompanying it. That is the problem genome-scale CRISPR knockout screening was built to address.

Figure 2. Each well in a pooled CRISPR screen carries cells with one gene deleted and one guide-sequence tag — the tag that lets a later sequencing read identify which deletion caused which outcome. — Image prompt and art direction by Brecht Corbeel; generation pending.
Shalem and colleagues built a pooled lentiviral library — the GeCKO library — carrying 64,751 distinct sgRNAs targeting 18,080 human genes, roughly three to four guides per gene, plus non-targeting controls. Delivered at low multiplicity of infection into Cas9-expressing cells, each cell in the pool stably integrates one guide and therefore one gene knockout — the population becomes a library of single-gene loss-of-function perturbations, each physically linked to a sequence tag identifying which gene was cut [4]. The population then goes through selection: in Shalem’s melanoma experiment, cells were exposed to vemurafenib, a BRAF inhibitor, and survivors were sequenced to see which guides became over-represented. A guide’s abundance rising under selection means the gene it targets, when deleted, conferred resistance — a direct causal readout, since the intervention preceded and is mechanistically linked to the observed phenotype, which a correlation between baseline expression and resistance cannot establish on its own [4]. The screen recovered NF1 and MED12, already known to affect resistance, alongside novel hits NF2, CUL3, TADA2B, and TADA1 — internal validation before the new hits are trusted [4].
This design only works if the guide-to-phenotype link is trustworthy per guide: an sgRNA can cut inefficiently, or at unintended sites resembling its target, corrupting the inference for that gene even if the library design is sound. Doench and colleagues addressed this by generating large-scale empirical activity data across thousands of guides, fitting predictive rules for which sequence features give high on-target and low off-target activity, then packaging those rules into design algorithms for improved libraries [5]. The rules are empirical regularities mined from data, not first-principles derivations, so they get periodically retrained rather than fixed once and reused.
Fact: a pooled CRISPR screen establishes causality because it perturbs the gene first, under a stable heritable tag, and reads the phenotype afterward — the step an association study cannot substitute for [4]. Vendor-style overclaim: a hit is causal for the specific phenotype under the specific condition applied (one drug, one cell line); “gene X causes resistance” stated unqualified overstates one context. Analysis: the guide-design layer [5] is load-bearing for the whole claim — a screen run with an older guide set and one with a current optimized set are not directly comparable hit-for-hit. Scenario: a strong hit with no prior support is not automatically wrong — Shalem’s novel hits later held up — but it warrants orthogonal confirmation before being treated as established.
A droplet-based scRNA-seq experiment can show a gene’s expression varies sharply across cells of a nominally uniform population; a methylation assay can show a plausible chemical reason a subset of those cells keeps the gene off; only a CRISPR intervention on that gene, in that cell type, shows whether the gene is actually load-bearing for the downstream phenotype. Treating any one as proof of the others’ conclusion is the most common way a real result gets overstated on the way to a headline.
Prediction, with horizon: within roughly a decade, combined platforms — CRISPR perturbation read out directly through single-cell RNA sequencing on the same cells, so a knockout’s transcriptional consequence is measured cell-by-cell rather than inferred from bulk survival — likely become the default design for mechanistic follow-up, on the indicator that such combined protocols already dominate new method publications relative to either technique alone. Undercut if per-cell readout costs fail to fall enough to displace simpler bulk screens for routine triage.
Originally published at https://absolutedigitalpublishers.com/articles/how-molecular-biology-and-genomics-actually-works.