Decoding the genomic Rosetta Stone
Explain sequence-to-expression reporter measurements.
Loading video…
# Decoding the genomic Rosetta Stone
Watch the video first. Use this companion to revisit the reasoning and its evidence limits.
Testing variants one by one would take years. Massively parallel reporter assays test thousands at once, every variant driving a reporter, all grown together. In the earlier Sort-Seq method, FACS sorted cells into bins by fluorescence, then the DNA in each bin was sequenced. The sorter never reads DNA; it sorts by brightness.
Reg-Seq, introduced in 2020 by Ireland and colleagues in the Phillips Lab, replaces sorting with counting. Each variant makes an RNA with a unique barcode, and its count in RNA relative to DNA measures expression. The first study covered more than a hundred E. coli promoters in twelve growth conditions.
Now the analysis. At each position, ask: does knowing whether this letter was mutated help predict whether a read came from RNA, that is, from expression? If so, the position likely carries regulatory information. That dependence, in bits, is the mutual information. Plotted along the promoter, it forms an information footprint.
Peaks say where. The expression shift says which way: where mutations raise expression, a repressor site is likely; where they lower it, an activator or the polymerase site.
But a footprint is a hypothesis. It shows that a stretch of DNA matters, not which protein binds it. A sequence resembling a known motif isn't automatically a working site. And mutations at one position needn't act alike: some letters weaken binding far more than others.
To find the protein, the site becomes bait: DNA containing it pulls binding proteins out of a cell extract, and mass spectrometry identifies the catch. Here, mass spectrometry identifies proteins; it doesn't measure expression. Matching sites to known motifs is a second route.
Once site and factor are known, the last video's tools return: sequence, to binding site, to transcription factor, to thermodynamic model, to quantitative prediction.
The newest version of this pipeline is a 2026 preprint by Röschinger and colleagues, not yet peer reviewed. With genome-integrated reporters, they studied more than a hundred genes, many of unknown function, in thirty-nine growth conditions. They recovered forty-nine known binding sites, found thirteen new ones, and assigned five of those to a specific transcription factor.
Identifying the factor is the hardest step. For two salt-induced genes, clear footprints appeared, yet mass spectrometry never found the protein. And the reporters span 160 letters, so distant sites are missed.
A second 2026 preprint, from Vincenzo Vitelli's group in Chicago with Phillips Lab members, rethinks the analysis. Its information blueprint groups correlated mutations into larger units, hyperletters, that best explain expression: coarse-graining applied to DNA. It shows one promoter can use different architectures in different conditions.
## Evidence guide
PHILLIPS-LAB PRIMARY RESULT: Reg-Seq links sequence changes to measured expression. MODEL PREDICTION: Pan’s synthetic-expression study tests the theory of the experiment. CURRENT PREPRINT: Röschinger’s expanded measurements and Gökmen’s collaborative informational blueprints remain unreviewed in the supplied audit and current lab list. Footprints do not identify a protein or prove function alone.
Sources: [ireland2020], [pan2024], [roschinger2026], [gokmen2026]. See the course bibliography and claim audit.