SpliceAI2 helps researchers predict splice sites, splice junctions, and complete transcript isoforms directly from DNA sequence
Addressing the tremendous complexity of human genetics
Prior to the Human Genome Project, scientists generally assumed that biological complexity would correlate with gene count, with estimates suggesting that humans might have 50,000–100,000 protein-coding genes1. Instead, the Human Genome Project revealed a surprisingly small number, ultimately settling at around 20,000, giving rise to the G-Value Paradox: the disconnect between gene number and organismal complexity2. In fact, the humble grocery store variety of banana has roughly 36,000 protein-coding genes, 80% more than the apparently more complex human3.
Two decades later, we now understand that gene count tells only a small part of the story. Human complexity arises from a rich regulatory landscape that determines how, where, and when genes are expressed and how a single gene can produce multiple outcomes. RNA splicing is a striking mechanism by which organisms achieve incredible complexity: by allowing one gene to generate multiple distinct transcripts and proteins, it dramatically expands the functional repertoire of a relatively small quantity of genes4.
Wearing many hats: Understanding how a single gene can perform a multitude of functions
Today, the BioInsight AI Lab launched SpliceAI2 (detailed in our paper), a new, pre-trained AI model designed to help researchers investigate alternative splicing. Over 95% of human genes undergo alternative splicing, demonstrating its near ubiquity in influencing an incredible diversity of functions ranging from slight modification of protein activity to completely opposite molecular roles4.
Alternative Splicing in Action
BCL2L1 is a gene involved in the normal process of apoptosis. However, previous research has found that BCL2L1 produces long (BCL-xL) and short (BCL-xS) isoforms due to splicing, both of which have opposite function5. Despite the pro-apoptotic roles of BCL-xS, several cancers promote the BCL-xL isoform, which acts as a cell survival factor by stopping cytochrome c release from mitochondria and blocking pro-apoptotic proteins like Bax5.
Such splicing activity is a fundamental part of development and is essential throughout an individual's lifespan. It also becomes altered in a variety of diseases, including cancer5. Given the incredible diversity of crucial regulatory changes related to alternative splicing, the BioInsight AI Lab has been working for years to develop tools to help researchers investigate how splice variants influence critical biology6.
Introducing SpliceAI2: A more complete understanding of splicing mechanisms
SpliceAI2 succeeds the original SpliceAI6, which the BioInsight AI Lab released in 2019. That model became one of the most widely used tools in translational genomics research and was cited in more than 3,400 publications. The initial SpliceAI model efficiently answers one key question. Will a cell splice in a given location, or not?
SpliceAI2 helps answer three key questions:
- Splice Sites – Which positions in a gene are used as splice sites, and how frequently are they used?
- Splice Junctions – Which splice sites are connected to one another during RNA splicing?
- Complete Transcripts – Given those splice sites and junctions, which full-length RNA transcript isoforms are produced?
The original SpliceAI remains a powerful tool because knowing where splice sites occur is an important part of the regulatory puzzle. However, it does not reveal how those sites connect or predict the RNA transcripts a cell ultimately produces. SpliceAI2 was therefore designed to go beyond splice-site detection by predicting how splice sites connect through splice junctions and reconstructing the complete transcript isoforms generated from a DNA sequence. Together, these capabilities provide a far more comprehensive view of how genetic variation shapes a gene’s capacity to perform often unexpected functions.
Importantly, all that is needed to run SpliceAI2 is a DNA sequence. The model was built with the understanding that DNA sequencing is common in translational research settings, and with the goal of helping researchers avoid the cost and complexity of generating long-read RNA data themselves. It also addresses a key limitation of RNA-based approaches, which is tissue accessibility. Many genes are not expressed at sufficient levels in blood to reliably measure splicing, meaning that researchers may need RNA from difficult-to-obtain tissues to evaluate potential splice-altering variants. By predicting transcript consequences directly from DNA sequences, SpliceAI2 enables transcript-level analysis without requiring access to the tissue where a gene is expressed.
How SpliceAI2 was built: AI is only as valuable as the data its trained on
The gains from SpliceAI2 come largely from the breadth and diversity of its training data. A training dataset more than 100 times larger was used for SpliceAI2 compared to the original SpliceAI training dataset. The updated training data included 314,745 RNA sequencing samples across ten species, covering more than 46 million observed splice junctions after filtering. By pooling hundreds of thousands of RNA sequencing experiments, the model effectively learns from the equivalent of an ultra-deep RNA-seq dataset, capturing not only common splice junctions but also many that are used only rarely. These infrequent splicing events can be difficult to observe in individual studies but collectively provide valuable information about the full spectrum of splicing biology.
Learning from this richer catalog of splice sites and junctions significantly improves the model's ability to predict splicing outcomes. This first set of training data was derived from a highly relevant subset of hundreds of thousands of publicly available datasets. This was crucial for the first two functions of SpliceAI2, which are splice site prediction and splice junction usage prediction.
The third function, predicting complete transcript isoforms, benefited from a different type of training data. Conventional short-read sequencing captures genetic messages in fragments, which is excellent for identifying individual splice sites but can be limiting when determining which splice events belong to the same finished transcript. Long-read sequencing, by contrast, captures transcripts end-to-end. To enable transcript-level prediction, the BioInsight AI Lab trained SpliceAI2 on 330 long-read samples from the public ENCODE project. This addition allowed SpliceAI2 to learn how individual splice events combine to form complete transcripts and enabled the model to reconstruct the single most common transcript 82% of the time for genes it had never seen during training, compared with 78% when trained on short-read data alone.
Model performance: More accurate across every test
To determine how SpliceAI2 stacks up against other analytical tools, the BioInsight AI Lab compared the model against the original SpliceAI, Pangolin, and Google DeepMind's AlphaGenome across three independent benchmarking datasets. The AlphaGenome comparisons were run independently by academic collaborators at the University of Oxford.
SpliceAI2 delivered the highest accuracy and strongest performance across all tested benchmarks. Specifically, it was the most accurate model at identifying variants that create entirely new splice sites, with the largest gains observed for variants located deep within introns. It also outperformed alternative approaches in several areas. First, it more accurately predicted how strongly a splice site is used, rather than simply whether it exists. Next, it identified variants that drive splicing differences in natural populations. Finally, its predictions matched the results of a large-scale experimental study that directly measured the splicing effects of hundreds of thousands of mutations (Figure 1).
Figure 1: SpliceAI2 delivered the highest accuracy and strongest performance across all tested benchmarks
Impact on genetic disease research: Looking beyond the protein coding regions
One of the most impactful findings from the SpliceAI2 study emerged from analyses of the Genomics England 100,000 Genomes Project, a large-scale effort to better understand the genetic basis of rare disease. Researchers examined samples with extensive phenotypic information from 7,504 individuals and asked a simple question. If SpliceAI2 predicts that a variant alters splicing, are those variants more likely to occur in genes already linked to an individual's observed phenotype?
To answer this rigorously, the team repeatedly shuffled phenotype assignments among participants and re-ran the analysis, establishing a baseline expectation for what would occur by chance alone. Variants prioritized by SpliceAI2 were significantly enriched within biologically relevant genes. At matched confidence thresholds, SpliceAI2 identified 17% more disease-associated variants than any other tested splicing model. SpliceAI2 also outperformed the legacy SpliceAI model, finding 33% more disease-relevant splice variants at a 2X confidence interval and 66% more at a 4X confidence interval (Figure 2). The strongest enrichment was observed in genes known to be highly sensitive to loss of function, suggesting that the model is identifying splice-disrupting variants with meaningful biological consequences.
Perhaps even more revealing was where many of these variants were located. Roughly 50% of the cryptic splice variants identified by SpliceAI2 occurred deep within intronic regions, often thousands of bases away from protein-coding sequences. These variants are particularly difficult to detect and interpret because they lie outside the regions typically emphasized in gene-focused analyses, yet they can create entirely new splice sites, alter transcript composition, and ultimately change gene function. The findings reinforce an important lesson from modern genomics, which is that understanding biology requires looking beyond the approximately 2% of the genome that directly encodes proteins.
Figure 2: SpliceAI2 improves identification of biologically relevant splice variants
AI at population scale: SpliceAI2 makes predictions that are backed by evolutionary and functional evidence
Success of SpliceAI2 in predicting impactful splicing variation encouraged the BioInsight AI Lab to test whether the model’s predictive capacity holds up at even larger scale. This time, SpliceAI2 was applied to a massive 627,000 sample genome and 36,764 sample proteome combined cohort, pulled from gnomAD, TOPMed, and UK Biobank to ask two questions:
- Across hundreds of thousands of genomes, are variants predicted to be deleterious actually rare in human populations?
- When present, do predicted deleterious variants reduce the amount of protein produced?
Both questions are based on the field-wide understanding that truly harmful variants reduce fitness, resulting in general rarity in a large population due to negative selection. The hypothesis being that if a model truly identifies detrimental variants, then those variants should be less-frequently observed both in genotype and consequential proteomic abundance in rare cases in which the variant is present.
The results provided independent evidence that SpliceAI2 predictions correspond to biologically meaningful effects. Variants assigned the highest SpliceAI2 scores were strongly depleted across more than 627,000 genomes, approaching the level of depletion observed for protein-truncating loss-of-function mutations. Independent proteomic data reached the same conclusion: individuals carrying higher-scoring variants consistently exhibited lower plasma protein levels, with stronger correlations than observed for other tested prediction models (Figure 3). Together, these findings demonstrate that SpliceAI2 scores capture real biological impact, not just statistical associations. They also show that the model generalizes across population-scale genetic and proteomic datasets.
Figure 3: Variants identified by SpliceAI2 are under negative selection in population-scale data─SpliceAI2 was applied to 627,000 sample genome and 36,764 sample proteome combined cohort, pulled from gnomAD, TOPMed, and UK Biobank, to test for negative selection of predicted impactful variants. A: SpliceAI2 best identifies likely deleterious variants. B: SpliceAI2 predictions best match downstream proteomic depletion of likely deleterious variants.
Modeling tissue-specific splicing: SpliceAI2 captures tissue-specific splicing differences
Another important advance introduced by SpliceAI2 is the ability to move splice variant research beyond impact prediction toward mechanistic explanation. Rather than simply identifying variants likely to disrupt splicing, the model begins to reveal the regulatory factors and sequence features driving those effects. Tissue-specific expression of RNA binding proteins is one of the key mechanisms by which cells accomplish cell type-specific splicing activity, and SpliceAI2 was developed with this crucial process in mind. SpliceAI2 incorporates the measured activity levels of 147 splicing regulators that guide the cellular splicing machinery. This allows the model to evaluate not only DNA sequences, but also the regulatory environment in which that sequence exists. Across nearly 15 million splice site differential usage measurements spanning 48 human tissues, SpliceAI2 accurately captured tissue-specific patterns of RNA splicing.
Importantly, the model was less successful at predicting how the impact of a given variant changes between tissues. Instead, researchers found that the strongest predictor of tissue-specific variant effects was the underlying splicing program already present in each tissue, highlighting the importance of cellular context. When they examined how the model reached its predictions, they discovered that it had independently learned many of the sequence motifs recognized by real splicing regulators and adjusted their importance based on regulator abundance. The model was never explicitly taught these relationships.
Together, these results suggest that SpliceAI2 is learning aspects of the regulatory code governing RNA splicing rather than simply memorizing outcomes. Because of this, the framework extends naturally beyond healthy tissues to disease research, successfully modeling splicing disruptions observed in cancers affecting the splicing machinery and in disorders such as myotonic dystrophy.
Better together: Combined analysis with multiple specialized models doubles identification of impactful variants
Since the human genome project, the explosion of multiomic technologies, analytical methodologies, and population-representative datasets have made one thing clear. It is exceptionally rare for any single variable to explain complex biology. Consequently, the BioInsight Lab has developed multiple specialized, pre-trained AI models with the intent of utilizing their combined strengths to simultaneously understand the many aspects of regulation yielding phenotypes and results.
SpliceAI2 represents a major advancement for one key aspect of cellular regulation. Among biologically relevant findings from analysis of the Genomics England 100,000 Genomes Project during SpliceAI2 development, cryptic splice variants identified by SpliceAI2 contributed an additional 15% of candidate disease-associated variants beyond those readily identified through clearer protein-disrupting mutations. Importantly, integration of other Illumina models, which add additional mechanistic context, approximately 2X the number of interpretable variants relative to use of traditional methodologies alone (Figure 4). Together, these results suggest the need for more integrated, unified analytical capabilities that expand access to a more complete repertoire of advanced annotation toolkits.
Figure 4: Combined analysis with Illumina AI models finds variants missed by traditional methods alone
Expanding access: Bringing advanced genomic interpretation to more researchers
SpliceAI2, PrimateAI-3D, Promoter AI, and many other Illumina analytical tools are intended to be maximally accessible across basic and translational research, and Illumina is creating multiple points of access to help researchers incorporate the model into existing analysis workflows. The BioInsight AI Lab develops the underlying models, while researchers access these innovations through BioInsight applications and software products.
Click here to create a free account to explore the BioInsight Platform
SpliceAI2 will be accessible across multiple tools on BioInsight Platform including:
- DRAGEN Annotation—SpliceAI2 and multiple other advanced annotation tools are now integrated into one intuitive application.
- Emedgene—SpliceAI2 brings even higher confidence to Emedgene, an explainable AI-powered genomic interpretation platform that helps translational researchers prioritize and explain germline variants.
- Illumina Connected Insights—SpliceAI2 further expands AI-driven annotation on the ICI cloud-based somatic variant analysis, annotation, and evidence review pipelines
Source code, trained models, and precomputed predictions for every possible single-nucleotide variant within human gene bodies, plus indels observed in human populations, are also now available on GitHub (github.com/Illumina/SpliceAI2).
Collaborate with BioInsight
As the BioInsight AI Lab improves and develops new advanced annotation tools, the foundational importance of high-quality training data has been exceptionally clear. Every model that the lab has produced performs best on data that resemble its training data, therefore the diversification and expansion of training data are crucially important to further advance capabilities and accuracy.
To accelerate development in this critical field, BioInsight welcomes opportunities for collaboration, both for SpliceAI2 and models in development. Contact BioInsight-Info@Illumina.com to inquire regarding improvement of cohort-specific performance, rare disease research, improved ancestry representation, model evaluation, research into future therapies, or other relevant topics.
About the BioInsight AI Lab
The BioInsight AI Lab develops sequence-based models for interpreting the entire genome, including non-coding regions, where common interpretation frameworks remain largely built around the 2% of the genome that encodes proteins. SpliceAI2 follows SpliceAI6, since incorporated into ClinGen's recommendations for clinical variant interpretation; PrimateAI-3D7, for protein-altering variants; and PromoterAI8, for variants affecting gene expression. Each has been published with its methods and benchmarks open to inspection and released for independent evaluation — including, in this work, an AlphaGenome comparison conducted by our academic collaborators rather than in-house.
Acknowledgements
This work was led by Kishore Jaganathan and Kyle Kai-How Farh, with Jianbin Chen, Xin Liu, and Yan Zhang as joint first authors, alongside colleagues at Illumina and collaborators at the University of Oxford, UCSF, and the New York Genome Center. The blog post was authored by Laurel Mastro, Lisa Eldridge, and Robert Yamulla.
We thank the participants of the Genomics England National Genomic Research Library, and the donors and participants of GTEx, gnomAD, TOPMed, and UK Biobank. We acknowledge the ENCODE Consortium for the long-read data underpinning transcript prediction.
References
- Salzberg SL. Open questions: How many genes do we have? BMC Biol. 2018 Aug 20;16(1):94. doi: 10.1186/s12915-018-0564-x. PMID: 30124169; PMCID: PMC6100717.
- Hahn, M.W. and Wray, G.A. (2002), The g-value paradox. Evolution & Development, 4: 73-75. https://doi.org/10.1046/j.1525-142X.2002.01069.x
- Droc G, Larivière D, Guignon V, Yahiaoui N, This D, Garsmeur O, Dereeper A, Hamelin C, Argout X, Dufayard JF, Lengelle J, Baurens FC, Cenci A, Pitollat B, D'Hont A, Ruiz M, Rouard M, Bocs S. The banana genome hub. Database (Oxford). 2013 May 23;2013:bat035. doi: 10.1093/database/bat035. PMID: 23707967; PMCID: PMC3662865.
- Wang Y, Liu J, Huang BO, Xu YM, Li J, Huang LF, Lin J, Zhang J, Min QH, Yang WM, Wang XZ. Mechanism of alternative splicing and its regulation. Biomed Rep. 2015 Mar;3(2):152-158. doi: 10.3892/br.2014.407. Epub 2014 Dec 17. PMID: 25798239; PMCID: PMC4360811.
- Kim YJ, Kim HS. Alternative splicing and its impact as a cancer diagnostic marker. Genomics Inform. 2012 Jun;10(2):74-80. doi: 10.5808/GI.2012.10.2.74. Epub 2012 Jun 30. PMID: 23105933; PMCID: PMC3480681.
- Jaganathan K, Kyriazopoulou Panagiotopoulou S, McRae J et al, Predicting Splicing from Primary Sequence with Deep Learning, Cell, 2019; 176, 535-548.e24
- Hong Gao et al. The landscape of tolerated genetic variation in humans and primates. Science 380, eabn8153 (2023). DOI:10.1126/science.abn8197
- Kishore Jaganathan et al. Predicting expression-altering promoter mutations with deep learning. Science 389, eads7373 (2025). DOI:10.1126/science.ads7373