The US government and some of the world’s largest AI companies are launching a major initiative for artificial intelligence in biology. The goal is to create AI models that can predict how human cells will respond to drugs and other changes—before any real-life experiments are needed.
Google DeepMind, Meta, and the pharmaceutical AI company Isomorphic Labs are joining forces with US government agencies and the research organization Biohub in a venture worth approximately $1.8 billion in total.
The project, called the Virtual Biology Initiative, aims to create vast amounts of standardized biological data that can be used to train the next generation of AI models. According to Biohub, it is the largest coordinated effort to date to generate biological data tailored specifically for artificial intelligence.
Want to experiment on cells digitally
The main problem for current AI in medical research is largely the lack of data. Language models can be trained on the massive amounts of text already available on the Internet. There is no comparable dataset describing how human cells react to millions of different changes. This project seeks to change that.
The researchers want to measure how different types of cells respond to various biological and chemical interventions, and then use the material to develop models that can predict what will happen with new interventions.
“An accurate predictive model of biology could dramatically accelerate scientific discoveries by enabling researchers to conduct experiments digitally,” says Biohub’s Head of Research, Alex Rives.
The long-term vision is a kind of “virtual cell” that reacts closely enough to a real cell so that researchers can digitally test hypotheses and potential drugs before moving on to laboratory trials.
Google and Meta invest $300 million
The money comes from multiple sources. Google DeepMind, Isomorphic Labs, and Meta are jointly investing $300 million. The US Department of Energy will contribute more than $500 million over five years for lab work, data collection, AI analysis, and computing power.
The US health agency, NIH, will also bring in databases and other research infrastructure that has been created through more than $500 million in previous federal investments. As early as April, Biohub also set aside another $500 million for the initiative.
Nvidia is also involved, providing computing power, software, and technical expertise. Among the research institutions participating are the Broad Institute, Human Cell Atlas, and the Wellcome Sanger Institute.
Decades of research compressed into five years
The ambitions are high. Biohub’s Head of Research, Alex Rives, tells Reuters that current databases contain hundreds of millions of cells, whereas sufficiently advanced models may require data from billions and eventually trillions of cells.
The first major dataset is scheduled to be ready in about a year. The goal is to achieve functional and reliable predictive models within five years—a job that would otherwise, according to the project, take decades.
The material will eventually be made open to researchers. However, according to Reuters, the private companies funding parts of the project will get a temporary head start, allowing them to work with certain data before it is released publicly. Data funded through public sources will not be subject to this limitation.
Similar races are already underway elsewhere. OpenAI has launched a program of more than $125 million for biological and medical datasets, while Anthropic is building its own laboratory operations for biological research.
