Nanopore sequencing
Introduction & Summary
Nanopore sequencing is a third-generation technique, works by passing a single nucleic acid strand through a biological pore (a "nanopore") and measuring the characteristic changes in an electrical current as each nucleotide passes through. This allows for the direct reading of the sequence without the need for amplification or complex labeling.
For mRNA therapeutics, Nanopore sequencing provides an unprecedented, high-resolution view of the poly(A) tail. By directly measuring the length of thousands of individual RNA molecules in a single run, it offers the most comprehensive profile of poly(A) tail length, distribution, and heterogeneity available today.
Key Quality Attributes Assessed
Method Evolution: Superseded, Current Standard, and Emerging
- Legacy: Traditional methods like gel or capillary electrophoresis after enzymatic digestion could provide an estimate of the average tail length and overall distribution, but could not measure the length of individual molecules.
- Current Standard: Nanopore direct RNA sequencing is a leading single-molecule method for in-depth characterization of poly(A) tails; it complements CE/CGE and other assays used in QC.
- Future: The technology is rapidly advancing to improve raw read accuracy, increase throughput, and develop GMP-compliant software. This may eventually enable its use not just for characterization but also as a comprehensive QC release assay.
Scientific Principle
The technique directly translates a molecule's structure into an electrical signal:
- Library Preparation: An adapter (via an oligo-dT RT adapter) and motor protein are ligated at the 3′ poly(A) end; RNA translocates 3′→5′, so the poly(A) tail is read first.
- Translocation: The motor protein guides a single strand of RNA through a protein nanopore embedded in a synthetic membrane. An ionic current is passed through this pore.
- Current Disruption: As each nucleotide (A, U, G, C) passes through the narrowest part of the pore, it disrupts the flow of ions in a unique way.
- Signal Acquisition: The instrument's sensor measures this electrical disruption in real-time, generating a characteristic signal "squiggle."
Data Analysis:
- Basecalling: A deep-learning algorithm translates the "squiggle" signal back into a sequence of bases (A, U, G, C).
- Poly(A) Tail Measurement: The long, repetitive poly(A) tail creates a distinct, sustained signal. The duration of this signal is directly proportional to the number of 'A's that passed through the pore, allowing for a precise length measurement for that single molecule.
Explainer Videos
Common Instrumentation & Software
Data Output & Interpretation
- Data Output: The primary processed output is a histogram that plots the number of sequencing reads (i.e., individual molecules) against their measured poly(A) tail length (in nucleotides).
- Interpretation: Analysts use the histogram to calculate the mean/median tail length, the breadth of the distribution, and the percentage of molecules with tails within a target specification.
Strengths
- Direct RNA Sequencing: Can read RNA molecules directly, avoiding biases from reverse transcription and PCR steps.
- Long Reads: Uniquely capable of sequencing an entire multi-kilobase mRNA molecule, including its full poly(A) tail, in one continuous read. Tools such as nanopolish polya/tailfindr convert the poly(A) dwell segment into length per molecule after calibration with standards.
- Single-Molecule Resolution: Provides a true distribution based on thousands of individual measurements, not an averaged signal.
- Real-Time Analysis: Data is generated and can be analyzed as the experiment is running.
Limitations
- Higher Raw Error Rate: The base-calling accuracy per nucleotide is lower than short-read sequencing, though this is less critical for measuring the length of a repetitive poly(A) tail.
- Complex Bioinformatics: Requires specialized bioinformatics expertise and computational resources to process the raw signal data and perform the analysis.
- Primarily a Characterization Tool: Currently used for in-depth characterization and investigation, not as a routine, validated GMP release test.
Key Validation Considerations
- Accuracy of Length Measurement: The analysis pipeline must be calibrated and verified using synthetic RNA standards with precisely known poly(A) tail lengths.
- Read Depth: The method must achieve sufficient sequencing depth (number of reads) to ensure the generated distribution is statistically representative of the entire sample.
- Bioinformatics Pipeline: The software and analysis workflow must be version-controlled and qualified to ensure consistent results.
Method Standardization & Reference Materials
- Synthetic RNA standards with defined sequences and poly(A) tails of various known lengths are essential for calibrating the analysis software and confirming the accuracy of the length measurement.
- Well-characterized in-house reference materials are used as benchmarks for comparing different manufacturing batches.
Use in Specific Modalities
- mRNA Therapeutics & Vaccines: Nanopore sequencing is the most powerful technology for characterizing the poly(A) tail, a critical quality attribute that dictates the stability and translational efficiency of the drug.
- General Transcriptomics: Also used in basic research to study native RNA modifications and alternative polyadenylation in cells.
