
Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)
Quick Answer
Xaira's X-Cell model leverages the X-Atlas dataset to enhance drug discovery by utilizing 30x more information, overcoming limitations of smaller datasets.
Quick Take
The shift from autoregression to diffusion and CRISPR-based experiments enables predictive modeling of gene expression changes in human cells, aiming to revolutionize AI-driven drug development.
Key Points
- X-Cell model trained on X-Atlas dataset enhances predictive capabilities for gene expression.
- 30x more information improves model performance beyond 1.5B parameters.
- CRISPR-based experiments generate extensive raw data for X-Atlas.
- New approach shifts from autoregression to diffusion for better results.
- Promotions of Ci Chu and Bo Wang highlight strategic importance of data-rich drug development.
DeepSignal Analysis
What happened
Xaira's X-Cell model utilizes the X-Atlas dataset to enhance drug discovery by providing significantly more information than previous datasets. The model shifts from autoregression to diffusion and incorporates CRISPR-based experiments, aiming to improve predictive modeling of gene expression changes in human cells.
Key evidence
- The X-Cell model is based on the X-Atlas dataset, which offers approximately 30 times more information than smaller datasets, addressing limitations in predictive modeling.
- Models trained on CELLxGENE have been shown to describe relationships between cell types and states but struggle to predict changes in RNA expression effectively.
- Xaira's approach includes CRISPR-based experiments that run millions of tests in parallel, generating the raw data necessary for the X-Atlas dataset.
Why it matters
The advancement in data richness is crucial for AI-driven drug development, as it allows for better predictive modeling of gene expression changes. This could lead to more effective drug discovery processes and a deeper understanding of cellular mechanisms, which are essential for developing targeted therapies.
What to watch
📖 Reader Mode
~3 min readIf test loss flatlines after 1.5B parameters while training loss continues to drop as you scale, that tells you that your model is limited by the amount of information in your data.
Training on a single, smallish data set exposed an information gap: the 3.1B model falls off the scaling trend. Neither parameters nor compute will improve performance past this wall. For predicting changes to gene expression, you need more information rich data.
This is what Chu and Bo’s teams have done, and here is what ~30x the information buys you:
Now we can scale with parameters and training compute! We don’t know how much this effort costed, but we can guess that data collection experiments and infrastructure was a few tens of millions, and compute + headcount + research was a few million. The budget looks like a RL rollout budget, rather than a data rich pre-training one.
We were lucky enough to have the two central figures in this story on our podcast. Taking the lead from Ci Chu and Bo Wang, Xaira Therapeutics is betting that information rich data is the key to AI-driven drug development. Chu was recently promoted to Chief Discovery Officer and Bo to Chief AI Scientist1, underscoring just how strategic Xaira considers this bet.
If you had to figure out how a human cell works, what would you do? A good place to start might be by documenting what genes are expressed (e.g. what RNA is floating around) in different kinds of cells, in different circumstances.
That is CELLxGENE, a database of 168M cells built by Chan Zuckerberg Institute that maps each cell to a count of how many times 20K-30K genes were detected in that cell, plus detailed metadata about every cell. A ~4 trillion-entry matrix.
If the Protein Data Bank (PDB) unlocked structural biology models [link Boltz, BioHub], CellXGene has done the same thing for Virtual Cell models. Like PDB, CELLxGENE has inspired a zoo of AI models of RNA expression; so much so that RNA expression models have become synonymous with Virtual Cell models. Bo Wang built one of the most influential, scGPT, that became the starting point for Xaira’s new model.
Models trained on CELLxGENE describe the relationship between cell types and cell states, but they are not good at predicting what will happen if we make changes to RNA expression. Changes in gene expression are highly correlated, and its is difficult (impossible) to figure out what causes what in most cases.
If you could “turn the dial down” on one gene at a time, however, then you would be able to observe what is upstream and downstream of a given gene2. You could tell if A → B & C or B → A & C or B → A, C → B → … If you did this for all of the genes, then maybe you could train a model that could predict what would happen to a cell if you change a gene (e.g. with a drug or a gene edit). Or maybe you could figure out the least invasive way to change a particular gene’s expression.
This is exactly what Chu and Bo’s teams have done. The data set is called X-Atlas and the model is called X-Cell.
In this episode, we discuss:
Why the team abandoned autoregression for diffusion
The CRISPR-based experiments that run millions of tests in parallel, and generate the raw data for X-Atlas and X-cell
Generalization to real lab experiments in real human cells
Beating the linear baseline that has outperformed previous models
Justifying a kitchen-sink of priors, and how that stacks up vs. data and architecture
Bo also shared with us some of the (major) advantages he has as an academic vs. industry leader, and how his labs keep up with the breakneck pace of AI innovation.
Check out the full episode on YouTube, or your favorite podcasting platform!
These promotions happened after we recorded the episode
There can be cycles in the chain reaction, of course, and there can be second, third, etc. order effects (meaning things that only happen when multiple genes change at once), but the first order effects are a great place to start, and might tell us a lot of what we need to know.
— Originally published at latent.space
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Latent Space
See more →
Ontologies Are So Back: Why AI Agents Are Reviving the Semantic Web
Frank Coyle at AIEWF 2026 emphasized the revival of ontologies as essential 'logical guardrails' for effective AI agents, integrating them with for better reasoning. Neo4j's Emil Eifrem highlighted three ontology types to enhance agent scalability, while Kingsley Idehen discussed the challenges and benefits of maintaining ontologies in AI systems.


![[AINews] SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition](https://substackcdn.com/image/fetch/$s_!8D6O!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fpbs.substack.com%2Fmedia%2FHMuQw2BXUAAJaQd.png)
