Beyond abundance: What HeLa's genetic drift reveals about how cells buffer and adapt

Protein abundance alone often does not explain a phenotype. We combined signaling activity, transcription factor activity, and protein complex assembly state instead, and used a panel of genetically diverse HeLa cell lines to test whether this gives a more mechanistic picture of cell state.
Like

Share this post

Choose a social network to share with, or copy the URL to share elsewhere

This is a representation of how your post may appear on social media. The actual post will vary between social networks

Explore the Research

Nature Publishing Group UK
Nature Publishing Group UK Nature Publishing Group UK

Complex assembly and activity states as multifaceted protein attributes explaining phenotypic variability - Molecular Systems Biology

The state of a cell depends not only on protein abundance, but also on the biochemical and cellular activities of proteins, which are largely invisible to abundance profiling alone. Here, we introduce a multi-omics framework that infers context-specific protein activities from transcriptomic, phosphoproteomic, and protein correlation-based protein-protein interaction data, integrating modality-specific algorithms via network diffusion. Applying it to a panel of phenotypically diverse HeLa cell lines, whose genetic drift provides a natural perturbation system, we make three findings. First, physical separation of monomeric and assembled protein fractions by protein correlation profiling provides direct evidence that complex assembly buffers variation in gene copy number and transcription, a mechanism previously only inferred from bulk measurements. Second, using Let7 perturbation data, CRISPR gene dependency scores, and subcellular localization, we orthogonally validate that inferred protein activities capture functional regulation linked to cellular phenotypes inaccessible from abundance data alone. Third, differential analysis of context-specific activity profiles identifies molecular mechanisms underlying phenotypic divergence, including a WIPF1/WIPF2--Arp2/3 axis governing invadopodium formation and infection susceptibility, and an immunoproteasome switch linked to immune adaptation.

Every systems biologist knows this frustration: transcripts or proteins go up or down, but that number alone tells you very little about what is actually happening on a mechanistic level. Abundance is easy to measure, but the biochemical and cellular activity of a protein, its signaling state, its regulatory role, its assembly into complexes, is what actually connects it to a phenotype. Over the past years, the Califano, Liu and Aebersold labs have built tools that try to infer these activities rather than just measuring abundance: VIPER for transcription factor activity (Alvarez et al, 2016, Nature Genetics), VESPA for kinase and phosphatase activity (Rosenberger et al, 2024, Nature Communications), and SECAT for protein complex assembly and stoichiometry (Rosenberger et al, 2020, Cell Systems). Each of these was developed and validated on its own. What we had not tested was whether combining them on the same biological system gives more mechanistic insight than any single layer, or whether the layers mostly tell you the same thing.

This question connects back to an earlier study from Yansheng Liu and colleagues during our time in the Aebersold lab at ETH Zurich. There, they had profiled 14 HeLa cell lines collected from 13 different laboratories (Liu et al, 2019, Nature Biotechnology). These are nominally the same cell line, but genetic drift over independent lineages had turned them into a naturally occurring perturbation panel, with matched genomic, transcriptomic, proteomic, and phenotypic data, including doubling time and Salmonella infection susceptibility. That study already showed that the relationship between copy number, transcript abundance, protein abundance, and phenotype is far from linear. But since it worked entirely at the abundance level, it could only partially explain why the two most divergent sublines, HeLa Kyoto and HeLa CCL2, differ so much in invasiveness and infection response. With VIPER, VESPA, and SECAT now available, we had the tools to revisit the same system at much higher resolution.

Graphical summary of the study.
A multi-omics framework infers protein activities from protein correlation profiling, phosphoproteomic data, and transcriptomic data. Applied to a panel of HeLa cell lines, it shows how protein complex assembly buffers genetic variation and drives phenotypic divergence.

Around the same time, Moritz Heusel and colleagues in the Aebersold lab adapted SEC-based protein correlation profiling to the newly emerging DIA/SWATH-MS platform (Heusel et al, 2019, Molecular Systems Biology). The added sensitivity and quantitative consistency made the approach much more practical to run at scale. We then built SECAT to analyze these profiles at the level of protein complex dynamics rather than static complexes (Rosenberger et al, 2020, Cell Systems), boiling each protein's complex behavior down to a handful of interpretable attributes. We always suspected this reduction would make SECAT particularly useful for multi-omic integration, since it turns something as messy as complex stoichiometry into a number you can put next to a transcript or a phosphosite.

For this study, we finally put that idea to the test. We added protein complex profiling to HeLa Kyoto and HeLa CCL2, physically separating monomeric from complex-assembled protein fractions across 420 LC-MS/MS runs by SEC-SWATH-MS, and processed the resulting profiles with SECAT. We then combined this with msVIPER (transcriptional activity) and msVESPA (signaling activity) applied to the existing transcriptomic and phosphoproteomic data from the 2019 study.

The first result gave us fairly direct evidence for something that had previously only been inferred from bulk correlations: that complex assembly protects proteins against gene dosage variation. A subunit that fails to find an assembly partner is recognized by cellular quality control and degraded, while its assembled counterpart is comparatively stable. Because SECAT quantifies the monomeric and assembled pool of the same protein separately, we could test this directly rather than by proxy. For proteins detected in both states, the monomeric fraction tracked copy number and mRNA abundance considerably more closely than the assembled fraction (Spearman's rho 0.32 vs. 0.20 for copy number, 0.57 vs. 0.48 for mRNA). And proteins that showed this buffering pattern were far more essential by DepMap CRISPR dependency scores than proteins that only ever occur as monomers (p = 1.35 x 10-29, median dependency 0.065 vs. 0.032).

The second result is the one I find more interesting, because it makes a genuine case for combining modalities rather than choosing one. Signaling activity, transcription factor activity, and complex state barely overlap in which proteins they flag as differentially active between Kyoto and CCL2: VESPA and SECAT share only 4 proteins out of more than 1,400 combined hits. That is not a failure of any single method. Each one reads out a genuinely different regulatory layer, and the overlap worth looking for is not at the level of individual proteins, but at the level of the biological processes they converge on. When we ran network diffusion across all eight layers against Reactome pathways, two coherent mechanisms emerged that no single omics layer had surfaced on its own: a cytoskeletal remodeling axis running through WIPF1, WIPF2, and the Arp2/3 complex, which we could connect to differences in invadopodium formation and Salmonella infection susceptibility between the two sublines, and a shift in proteasome subunit stoichiometry (PSMB8/PSMB9) consistent with a switch to the immunoproteasome. Both were essentially invisible in total protein abundance; PSMB8 was not even detected as differential in the unfractionated proteome data at all.

Schematic comparing the two HeLa cell lines.
In the invasive, infection-susceptible CCL2 line, high Arp2/3 and WIPF1 stoichiometry drives invadopodium formation and matrix degradation, alongside immunoproteasome use. In the less invasive, infection-resistant Kyoto line, high WIPF2 blocks invadopodium initiation and cells rely on the constitutive proteasome instead. Neither side of this picture is visible from protein abundance alone.

We are not claiming that activity and complex state analysis replaces abundance measurements, or that this is the only way to integrate multi-omic data. Latent factor approaches such as MOFA and causal network approaches such as CARNIVAL or CORNETO solve different, complementary problems. What we think this study adds is fairly direct physical evidence for a buffering mechanism that was previously only inferred, and a demonstration that stacking functionally interpretable, network-derived attributes across omics layers reveals mechanisms, an immunoproteasome switch, a cytoskeletal regulatory axis, that stay hidden within any single layer, abundance included.

This connects to a broader debate in the field right now. Recent work on single-cell transcriptomic foundation models found that these models plateau well before they run out of training data, with no clear scaling law relating pretraining set size to downstream performance (DenAdel et al, 2026, Nature Methods). A related benchmark found that none of several published foundation models for predicting the effect of a genetic perturbation actually outperformed simple linear baselines (Ahlmann-Eltze et al, 2025, Nature Methods). We suspect one reason is that these models are trained almost entirely on raw transcript abundance, the same signal we show here is a poor proxy for what a cell is actually doing. We think functionally grounded attributes like the ones in this study, protein activity states rather than raw abundance, could be a useful input for these models to learn from, not just an alternative to them. That is a direction we want to test directly next.

Please sign in or register for FREE

If you are a registered user on Research Communities by Springer Nature, please sign in

Follow the Topic

Computational and Systems Biology
Life Sciences > Biological Sciences > Biological Techniques > Computational and Systems Biology
Proteomics
Physical Sciences > Chemistry > Analytical Chemistry > Mass Spectrometry > Proteomics
Molecular Biology
Life Sciences > Biological Sciences > Molecular Biology
Cell Biology
Life Sciences > Biological Sciences > Cell Biology