Key points
- Big Tech and US agencies join Biohub's AI biology effort
- Part of the total was committed before this week
- Corporate funders get an early look at the data
Biohub, the research nonprofit started by Meta chief executive Mark Zuckerberg and Priscilla Chan, called an expanded $1.8 billion effort "the largest coordinated commitment to generating AI-ready biological data to date." Two US agencies are joining, along with Meta Platforms (META) and two Alphabet (GOOGL) companies, Google DeepMind and Isomorphic Labs.
The goal is to build "virtual cells": AI models that predict how cells respond to changes such as a new drug. Biohub announced the expansion of its Virtual Biology Initiative on October 7.
Where the $1.8 billion comes from
The Department of Energy will invest more than $500 million over five years in lab measurement, modeling, and computation through its Genesis Mission. Meta, Google DeepMind, and Isomorphic Labs are putting in $300 million together. Nvidia (NVDA) will support the work with computing infrastructure, software, and technical expertise.
The rest was committed earlier. The National Institutes of Health will coordinate datasets built with more than $500 million in earlier federal funding, which Biohub will standardize for AI training. The total also counts the $500 million Biohub committed in April.
Corporate funders see the data first
The datasets will eventually be public, but the companies paying for them get a head start, Biohub head of science Alex Rives told Reuters. "With commercial funders we have embargo periods where there's a period of time where the groups can work on the data, and then it becomes available as a public scientific resource," he said.
Rives said the government-funded work running alongside it won't carry those restrictions, and Biohub plans to approach drug companies and philanthropies next.
The partners aim to compress decades of work into five years
Rives said the work would normally take decades, and the partners aim to compress it into five years. The first dataset should be ready in about a year.
Today's cell datasets run to hundreds of millions of cells, while an accurate model will need billions and eventually trillions, Rives said. "We need to capture the language of biology, we need to capture the language of the cell. And that doesn't exist today," he said.
Anthropic has expanded its biology work with a wet lab, and the OpenAI Foundation started a grant program of more than $125 million for biological and medical datasets, Reuters reported.
"Biology has been just sort of a clever discovery-based science until this point," Chan told Reuters. "We have always held this as a community asset, not just for one group, so that it can build upon itself over time."



