Shift Bioscience says better benchmarking could make AI “virtual cells” more useful for finding ageing drug targets
A new Nature Biotechnology paper argues that some genetic perturbation models were judged with poorly calibrated metrics, and the Cambridge company plans to use its fix to screen for rejuvenation and fibrosis targets.
Shift Bioscience, a Cambridge, UK biotech working on cell rejuvenation, has published a paper in Nature Biotechnology describing a new way to calibrate the metrics used to evaluate AI models that predict how cells respond to genetic changes. The company, known simply as 'Shift', says the framework will now underpin large-scale lab and computational screens for new drug targets, starting with fibrosis.
What are genetic perturbation models?
Genetic perturbation models are a type of AI “virtual cell”. They are trained to predict how a cell’s gene expression changes when a gene is switched on or knocked down. If they work, researchers can test thousands of interventions computationally before committing bench time to the most promising ones.
That “if” has been a sticking point. Earlier studies questioned the reliability of these models, and some failed to beat simple baseline approaches. For anyone considering building a drug discovery programme on them, that is a serious problem.
What did Shift find?
According to the company, the issue may partly lie in how the models were scored. Its framework accounts for both the biological and technical signals in a dataset, and it showed that underperformance in some earlier benchmarks could come from miscalibrated metrics that were not sensitive enough to detect real model skill. The work builds on research Shift first reported in November 2025.
In essence, the claim is not that the models are now better. It is that the ruler used to measure them was bent, and that a straighter one gives a more honest picture. Dr Brendan Swain, CSO and Founder of Shift Bioscience, said:
“Our findings show that by using well-calibrated metrics and the right dataset, virtual cell models can generate biologically meaningful insights.”
Why does this matter?
Reliable evaluation is the unglamorous foundation of any AI-driven discovery effort. Without a trusted benchmark, a model’s output is hard to tell apart from noise, and teams cannot compare approaches or know when to trust a prediction. A credible calibration method, published in a peer-reviewed journal, helps the whole field regardless of whose models it is applied to.
For Shift, the practical payoff is a target discovery programme. The company will run in vitro and in silico screens for inhibition targets that could address both rejuvenation and age-related disease, following the discovery of SB-101, which it describes as its first dual-purpose target. Fibrosis, a driver of ageing and age-related disease, is the starting point. Swain said the company is applying the framework “directly in our target identification program.”
What remains unproven?
A better benchmark does not guarantee better targets. The paper addresses how models are assessed, not whether the targets they nominate will hold up in cells, animals or patients. The press release gives no clinical timeline, and “rejuvenation” remains a hard biological claim to demonstrate. The real test will be whether targets selected with this framework validate experimentally. Swain’s claim of a “clearly defined route towards clinical development” is the company’s view, and it will need data behind it.
Still, fixing how the field measures virtual cell models is a sensible and necessary step, and the results of Shift’s screens will show how much it matters in practice.

Author
BioFocus Newsroom


