Next-Generation Sequencing (NGS)

Introduction & Summary

Next-Generation Sequencing (NGS) refers to a set of high-throughput platforms that allow the simultaneous sequencing of millions of DNA molecules. In biopharmaceutical CMC, NGS is used primarily for deep characterization of AAV vector genome identity, full-length integrity, and packaged DNA impurities. It encompasses both short-read (e.g., Illumina) and long-read (e.g., PacBio, Oxford Nanopore) platforms.

Key Quality Attributes Assessed

Method Evolution: Superseded, Current Standard, and Emerging

  • Legacy Technique: Sanger Sequencing & Restriction Mapping. For decades, the combination of restriction enzyme digests (to confirm the overall map) and Sanger sequencing (to verify the sequence of key regions like the transgene) was the standard for genetic identity testing.
  • The NGS Platform for Characterization. The current industry standard for in-depth characterization of vector genomes is NGS, typically complemented by orthogonal confirmation such as Sanger sequencing. A combination is often used: highly accurate short-read sequencing (Illumina) for deep variant analysis and long-read sequencing (PacBio, Oxford Nanopore) to confirm the full-length integrity across the ITRs.
  • Future: Integrated Long-Read Sequencing and AI. The future is trending towards the routine use of highly accurate long-read sequencing (like PacBio HiFi reads) as the single, definitive characterization method. The next evolution will be integration of advanced bioinformatics and Artificial Intelligence (AI) to streamline data analysis and better predict the impact of sequence variants. These approaches are still exploratory and not yet part of regulatory expectations.

Scientific Principle

NGS platforms share a principle of massively parallel sequencing, but differ in execution:

Short-Read (Illumina)

DNA is fragmented and ligated with adapters. Sequencing occurs via cyclic nucleotide incorporation, with fluorescent base calls captured in real-time across millions of clusters.

Long-Read (PacBio SMRT)

Uses Zero-Mode Waveguides (ZMWs) to monitor DNA polymerase incorporating labeled nucleotides in real-time. Enables sequencing of full AAV genomes and plasmids without fragmentation or PCR amplification.

Explainer Videos

Common Instrumentation & Software

Data Output & Interpretation

  • Output: The primary output is a data file (e.g., FASTQ) containing millions of DNA sequence reads, each with associated quality scores.
  • Interpretation: The reads are aligned to the expected reference sequence of the vector genome. The data is analyzed to confirm full-length integrity (ITR-to-ITR) and high-fidelity sequence identity of the transgene cassette. Any deviations are quantified and assessed. Reads that align but have differences are quantified as product-related impurities, and their levels are monitored to ensure process consistency.

Strengths

  • Long Read Lengths (PacBio, Nanopore): Essential for sequencing the entire AAV genome, resolving complex structures, and sequencing through the ITRs.
  • High Accuracy and Depth (Illumina): Ideal for detecting very low-frequency sequence variants or confirming the sequence of the starting plasmid material with high confidence.
  • Comprehensive Information: Can provide a complete picture of the genetic material present in the drug product, including desired genome and undesired impurities.

Limitations

  • Bioinformatics Expertise: Processing, analyzing, and interpreting large NGS datasets requires significant bioinformatics expertise and computational infrastructure.
  • Platform-Specific Challenges: Short-read methods struggle with highly repetitive and structured regions like the AAV ITRs. Long-read methods have historically had a higher raw error rate, though this is now largely mitigated by consensus sequencing approaches (e.g., PacBio HiFi reads).
  • Primary Role as a Characterization Tool: Due to its complexity, cost, and turnaround time, NGS is primarily used for in-depth product characterization, reference standard qualification, and comparability studies. While it is not typically used for routine, multi-attribute lot release testing, validated NGS-based assays are increasingly being implemented for specific QC applications, such as adventitious agent testing.

Key Validation Considerations

As a characterization method, NGS undergoes formal method qualification, and laboratories must verify suitability under actual use conditions in line with USP <1226>.

  • Accuracy: The accuracy of the final consensus sequence is confirmed against an orthogonally determined sequence (e.g., Sanger sequencing of specific regions).
  • Coverage: Qualification must demonstrate that a sufficient number of reads ("read depth") are generated to provide high confidence in the final consensus sequence across the entire genome.
  • Limit of Detection (LOD): When used for impurity analysis, the method must be qualified to determine the lowest frequency at which a variant or contaminant sequence can be reliably detected.

Method Standardization & Reference Materials

NGS methods require rigorous standardization because results depend heavily on library preparation, sequencing platform, and bioinformatics pipelines. Critical parameters such as read depth, coverage uniformity, error-correction strategy, and reference genome selection must be defined in controlled SOPs to ensure data comparability across studies and sites. Reference materials play a key role: well-characterized plasmid DNA or synthetic constructs are commonly used as positive controls to benchmark sequencing accuracy, while mock or null samples provide negative controls. Regulatory expectations (e.g., USP <1226>, FDA/EMA gene therapy guidances) emphasize demonstrating method suitability in the intended laboratory environment and maintaining lifecycle control of reference materials, including bridging studies when reagents, platforms, or software versions change.

Use in Specific Modalities

  • AAV Gene Therapy: A cornerstone characterization technique. Regulatory guidances, such as the FDA's CMC Guidance for Human Gene Therapy INDs, establish the expectation for a thorough characterization of the vector genome that includes its complete sequence and integrity. Long-read NGS is increasingly used to meet these expectations and is often included as a critical component of BLA submission packages.
  • Plasmid DNA Characterization: Used to provide a complete, closed-circle sequence of the entire plasmid starting material, confirming its integrity beyond what Sanger sequencing alone can provide.
  • Cell Therapy: Used to verify the integrity of the integrated transgene and to search for off-target integration events.

Key Regulatory Guidance