Tayfun Özçelik, Barış Kayaalp and Kerem Çil
Bilkent University, Department of Molecular Biology and Genetics
Every genome carries thousands of missense variants. Most arise from a single DNA change that alters one amino acid. Some are harmless, others alter protein function or help explain disease, and many remain difficult to interpret. Reading DNA is now fast; understanding it remains hard. The question is: what does this variant mean for biology, disease, and a person or family?
I have been thinking about this gap for a long time (Figure 1). In 2007, Nature Genetics asked several of us what we would do with a $1,000 genome1. My answer was optimistic, but cautious: personal genomic medicine would depend not only on sequencing capacity, but on our ability to interpret what we found. Nearly two decades later, sequencing is no longer the bottleneck. But a genome without interpretation is like an ancient inscription before its Rosetta Stone: the signs are visible, but the language has not yet been understood.
FuncVEP grew from that unresolved problem, not from a desire to add another missense predictor to a crowded field. We wanted to discover disease genes in large cohorts, and gene discovery depends on choosing the right variants. Poorly interpreted alleles can make biological signal disappear into noise.
The work developed inside a larger effort at Bilkent to improve variant interpretation. AAVC, developed by Barış Kayaalp and our team, addressed automated ACMG-based classification2. But missense variants remained difficult: a single amino-acid substitution may have little consequence, destabilize a protein, disrupt an interaction, alter activity or create a new molecular effect.
Our first intention was pragmatic. We wanted to identify the best-performing missense predictor to incorporate into the framework. But comparing predictors required a benchmark independent of clinical classifications many tools had already seen in training. We therefore turned to systematically curated functional data from experimentally studied variants.
The results were unexpectedly disappointing. When we evaluated existing predictors against this direct functional evidence, none performed as well as we had hoped; even the best achieved an accuracy below 80%. That observation changed the direction. Instead of asking which predictor to use, we asked whether our functional evidence could provide a better foundation for prediction.
It also made us look more carefully at what different classes of predictors were learning. Clinical classifications are enormously valuable, but they are not direct measurements of molecular consequence. They are influenced by penetrance, age of onset, phenotype definition, diagnostic accuracy, and other factors. Clinical labels can therefore be an imperfect target when the goal is specifically to learn the functional consequences of a missense variant.
There was also a major concern: circularity. As we examined how predictors were developed and evaluated, its importance became increasingly clear. Computational predictions can contribute to clinical classifications, those classifications can later enter training datasets, and successive generations of predictors may begin to reinforce information derived partly from earlier predictions.
We therefore asked whether comparatively direct functional evidence could serve not merely as a benchmark, but as the learning target itself. Functional assays are also imperfect, but they measure effects on proteins, cellular processes or molecular phenotypes without first requiring a clinical interpretation of the variant. At the same time, molecular effect is not synonymous with disease in a person; genetic background, compensation, inheritance, environment and phenotype all influence the clinical outcome.
Building FuncVEP was as much about discipline as prediction. We needed the right evidence, protection against leakage, and generalizability beyond a narrow benchmark. We also kept asking the human-genetics question: does this make biological sense?
One interesting finding was that a model trained on functional data achieved better predictive performance on clinical data than models trained on a subset of clinical data itself. This result suggested that labels shaped by accumulated clinical annotation may not always provide the cleanest target for learning molecular effects.
Peer review also shaped the final work3. Reviewers pushed us hard on circularity, leakage, and robustness, sharpening the analysis and our rationale for functional training.
The work happened in the ordinary rhythm of a university lab: long discussions, sanity checks, weekends writing, and results that had to earn our trust. Barış brought a physician-scientist instinct for clinical variant interpretation, genetic epidemiology and genotype-first discovery. Kerem brought mathematical and computational precision for model development. Together, we tried to keep the tool tied to human genetics and biology.
One satisfying outcome was a precomputed atlas of possible human missense variants scored by FuncVEP. Medicine increasingly needs reference maps: ways to ask, before a variant appears in a clinic, what its likely functional effect may be.
FuncVEP is not meant to replace clinical judgment or become a magic label machine. In clinical genetics, evidence still needs context: gene, mechanism, inheritance, phenotype, frequency, segregation and functional data. A good computational tool should make that process stronger, not pretend to abolish it.
For us, better missense interpretation is not an endpoint. It is an enabling layer. This view grew from a genotype-first way of thinking we had proposed4. After benchmarking FuncVEP, we used it with AAVC across inborn-errors-of-immunity genes to ask whether improved interpretation could support genotype-first discovery in population cohorts. It did: the framework helped uncover new gene–phenotype relationships and estimate inherited burden at population scale5. Seeing interpretation become medical discovery was more satisfying than any benchmark score.
The same logic also changed the scale of the question. Variant interpretation is often discussed one patient at a time; large datasets let us ask how many people carry inherited genotypes relevant to risk, protection, diagnosis, surveillance or prevention. As this layer becomes clearer, precision medicine can move toward earlier recognition and risk stratification.
Looking back, the lesson is modest but powerful: uncertainty can be reduced when the right evidence is used carefully. Functional data, clinical data and computational models are all imperfect. If we build with respect for their strengths and weaknesses, variant interpretation becomes more reliable. FuncVEP came from that conviction. The work is technical, but the motivation is simple. Genomic medicine will mature only when we learn not just to read DNA, but to understand it better.
References
- Özçelik T. Farewell to abnormal genes? Nature Genetics Question of the Year: The $1,000 genome. 2007.https://www.nature.com/collections/zpsrgkcbkx
- Arda İnan R, Kayaalp B, Safieh F, Ece Kars M, Stein D, Cooper DN, Stenson PD, Konu Ö, Casanova JL, Itan Y, Nazlı Başak A, Özçelik T. AAVC: An automated framework for high-accuracy ACMG-based variant classification. Genet Med. 2026;102624.
- Kayaalp, B., Çil, K., Conil, C. et al. Prediction of human missense variant effects from functional evidence. Nat Genet (2026). https://doi.org/10.1038/s41588-026-02727-3
- Özçelik T, Onat OE. Genomic landscape of the Greater Middle East. Nature Genetics. 2016;48(9):978-9.
-
Kayaalp B, Kars ME, Itan Y, Başak AN, Casanova JL, Özçelik T. Inherited burden for disease predisposition in diverse populations. npj Genom Med. 2026;18;11(1):18.
Figure 1. Half-chromatid mutation, stylized as a mosaic. Original artwork by Tayfun Özçelik, featured as the cover art of Nature Genetics, Volume 34, Issue 4 (August 2003).