Further Gains in Germline Whole Genome Sequencing Accuracy: NovaSeq X Series Software v1.4 and DRAGEN v4.5 Deliver Improved Results, Including in the Most Challenging Regions

In our recent blog on the NovaSeq X Series software v1.4 update, we highlighted improvements in sequencing quality, usability, and workflow flexibility enabled by the NovaSeq X Series software v1.4 release, demonstrating clear gains in FASTQ quality metrics and overall data consistency. We presented whole genome sequencing (WGS) germline small variant calling results for the software v1.4 update using the DRAGEN analysis pipeline v4.4. 

In this blog, we extend that work by evaluating NovaSeq X Series software v1.4 with the latest DRAGEN analysis pipeline v4.5.4, highlighting additional improvement in variant calling accuracy, including in regions that have historically been more difficult to resolve.

From improved sequencing to improved variant calls

A key driver of the improvements observed with the NovaSeq X Series software v1.4 update is continued optimization of the sequencing process, including enhancements to the sequencing recipe, clustering, and primary analysis.  

With NovaSeq X Series software v1.4, refinements to the sequencing recipe—including optimization of thermal and fluidic steps—are complemented by improvements in clustering that promote the formation of more genomically pure and robust clusters. These changes increase signal-to-noise ratios and lead to more consistent base quality across sequencing cycles. 

In parallel, enhancements to primary analysis further improve signal processing performance, particularly in regions that are more challenging to sequence accurately, such as homopolymers.

Together, these improvements reduce systematic error and improve data consistency, resulting in higher-quality inputs for downstream analysis and a stronger foundation for accurate variant detection across the genome.

Evaluating variant calling with the latest DRAGEN pipeline 

In our previous blog, we presented WGS germline small variant calling results for NovaSeq X Series software v1.4 using the DRAGEN pipeline v4.4. Illumina has continued to advance secondary analysis to further enhance variant calling accuracy and consistency.

Recent updates in DRAGEN v4.5.4 build on this progress by enabling sample specific personalized alignment and variant calling by default, optimizing variant detection and improving genotyping to reduce both false positives and false negatives.

To reflect these advances, the analysis presented here processes both NovaSeq X Series software v1.3 and v1.4 datasets using DRAGEN v4.5.4, enabling a consistent comparison while highlighting the combined impact of sequencing and analysis improvements on downstream accuracy.

Study design 

To assess the impact of NovaSeq X Series software v1.4 on WGS germline variant calling we designed the study summarized in Figure 1. 

Figure 1: Study design. HG001-HG005 samples were prepared with a TruSeq PCR-Free 450 library, sequenced on a NovaSeq X 10B flow cell, processed with both software versions v1.3 and v1.4, subsampled at 35x raw coverage, and processed using DRAGEN v4.5.4. The results were benchmarked with NIST T2TQ100-v1.1-v0.019 truth set for HG002 and NIST v4.2.1 truth set for HG001-HG005.

Improved small variant calling accuracy with NovaSeq X Series software v1.4

Figure 2 shows a reduction in total small variant calling errors (false positives + false negatives) when moving from NovaSeq X Series software v1.3 to v1.4 sequencing data[1] and processed with DRAGEN v4.5.4.

For HG002, evaluated against the NIST T2TQ100 truth set, total errors decreased from 78,860 (software v1.3) to 48,923 (software v1.4)—a reduction of approximately 38%. This improvement is driven by decreases in both false negatives (from 53,435 to 32,848) and false positives (from 25,425 to 16,075).

These results show that improvements in NovaSeq X Series software v1.4 sequencing translate directly into more accurate variant detection, substantially reducing both missed variants and erroneous calls in whole genome analysis. 

Figure 2: WGS Germline small variant errors (FP + FN) for HG002 (NovaSeq X 10B, TruSeq PCR-Free 450) processed with DRAGEN v4.5.4, showing a substantial reduction from NovaSeq X Series software v1.3 to software v1.4. 

Importantly, these improvements extend beyond a single sample. Across HG001–HG005 sequenced on the 10B flow cell with a TruSeq PCR-Free 450 library at 35x raw coverage, NovaSeq X Series software v1.4 maintains strong and consistent performance against established truth sets. Evaluated against the NIST v4.2.1 truth set, upgrading from NovaSeq X Series software v1.3 to software v1.4 consistently reduces total errors (FP+FN) across all five samples, lowering the overall error range from 5,146–8,227 errors in software v1.3 down to 3,869–8,145 errors in NovaSeq X Series software v1.4 (representing a 1.0% to 31.3% reduction).

The larger improvement observed for HG002 against the T2TQ100 truth set (~38%, Figure 2) compared to the v4.2.1 analysis (~18% across HG001–HG005, Figure 3) reflects differences in truth set composition. T2TQ100 includes a higher proportion of homopolymer-rich regions, where NovaSeq X Series software v1.4 delivers the most substantial gains, as shown in the next section.

Figure 3: WGS Germline small variant errors (FP + FN) for HG001-HG005 (NovaSeq X 10B, TruSeq PCR-Free 450) processed with DRAGEN v4.5.4, showing a substantial reduction from NovaSeq X Series software v1.3 to software v1.4.

Largest improvement in homopolymer regions

A more detailed view of the data highlights where these improvements are most pronounced. As shown in Figure 4, large improvements are observed in homopolymer regions (≥10 bp), which have historically been more sensitive to sequencing and variant calling error.

In this context, NovaSeq X Series software v1.4 reduces the combined burden of false positives and false negatives by more than 50% compared to software v1.3, demonstrating a substantial gain in variant calling accuracy for HG002.

These improvements are consistent with the underlying enhancements in sequencing performance and in the latest DRAGEN analysis updates. By reducing systematic error and delivering more consistent base quality across sequencing cycles, the NovaSeq X Series software v1.4 update enables more accurate detection of variants in homopolymer regions, where error profiles can otherwise obscure true signal. 

Figure 4: WGS Germline small variant errors (FP + FN) for HG002 (NovaSeq X 10B, TruSeq PCR-Free 450) processed with DRAGEN v4.5.4, showing a substantial reduction from NovaSeq X Series software v1.3 to software v1.4 in homopolymer regions.

Sequencing and analysis together drive performance

Accurate variant calling demands high-quality data and high-quality analysis, and both have to advance together. The NovaSeq X Series software v1.4 update raises the quality and consistency of sequencing data: higher signal, lower systematic error, more consistent base quality. DRAGEN v4.5.4 converts that data into better calls through improved modeling, candidate generation, and filtering. These results reflect what becomes possible when sequencing and analysis improve in parallel.

Summary

The NovaSeq X Series software v1.4 update delivers improvements that extend beyond sequencing metrics and are clearly reflected in downstream results. When processed with the latest DRAGEN pipeline v4.5.4, NovaSeq X Series software v1.4 data shows: 

  • Up to 38% fewer germline small variant calling errors

  • More than 50% error reduction in homopolymer regions (≥10 bp)

  • Consistent accuracy gains across samples

Together, these gains reinforce the value of continuously advancing sequencing and analysis in tandem, enabling high-confidence whole genome sequencing even in regions of the genome that have historically been the most difficult to resolve. 

Interested in improving WGS performance?

Connect with an Illumina specialist to learn how NovaSeq X Series and DRAGEN can support your sequencing objectives and help optimize your workflow.

Speak to a specialist

References

  1. FASTQ generated from the 10B flow cell with a TruSeq PCR-Free 450 library at 35x raw coverage.