
By: Yong-Han Hank Cheng
[Editor’s Note: Yong-Han Hank Cheng, a UW M.D.-Ph.D student in the Stergachis Lab, is the first author of the peer-reviewed paper, “Long-read transcriptome analysis using IsoRanker for identifying pathogenic variants in Mendelian conditions,” published in September in the American Journal of Human Genetics. In this blog, he summarizes the paper’s findings. Read the full paper here.]
Long-read RNA sequencing is giving us an increasingly detailed view of the human transcriptome. Instead of reconstructing transcripts from short fragments, long-read sequencing can capture full-length RNA molecules, making it possible to directly observe alternative splicing, novel isoforms, and other transcript-level changes. But this new level of detail creates an important challenge: How do we determine which of the many transcriptomic abnormalities in a patient are actually relevant to disease? Our new study introduces IsoRanker, a computational framework designed to help answer that question.
Framework for disease-gene prioritization
When RNA sequencing reveals an unusual transcript, it can be difficult to know whether that observation is meaningful. Some abnormal-looking transcripts occur naturally in healthy individuals, while others may be consequences of rare genetic variants that disrupt gene regulation or RNA processing to cause disease. IsoRanker was developed to systematically identify and prioritize unusual transcriptomic events from long-read RNA sequencing data.
Rather than focusing on a single type of RNA abnormality, the framework evaluates multiple signals that can indicate disrupted gene function. These signals can then be used to highlight genes whose RNA profiles stand out from those of control samples. The goal is to move the most biologically informative genes toward the top of the list for investigators and clinicians evaluating an unsolved genetic condition.
From disease-gene prioritization to patient impact
One of the clearest examples of IsoRanker’s clinical potential came from a patient with a previously unsolved genetic condition. IsoRanker highlighted abnormal expression and splicing involving HARS1, a gene encoding histidyl-tRNA synthetase, the protein that attaches tRNA with the amino acid histidine. Follow-up analysis showed that the patient carried noncoding HARS1 variants that were the cause of these changes in HARS1.
That finding ultimately led to a diagnosis for the patient, but its impact extended beyond a single case. Since then, five additional patients carrying the same HARS1 variants have been identified and diagnosed with HARS1-associated multisystem ataxic syndrome, helping define a group of individuals affected by the same molecular mechanism.
The discovery has also opened the door to thinking about treatment. Because HARS1 is responsible for attaching the amino acid histidine to its corresponding tRNA, researchers are now investigating whether histidine supplementation could help compensate for reduced HARS1 function. For us, this illustrates what we hope tools like IsoRanker can accomplish: identifying unusual RNA patterns, connecting those patterns to a genetic diagnosis, finding additional patients with the same condition, and pointing toward approaches for treatment.
What IsoRanker can reveal beyond AI prediction
We also compared the RNA abnormalities identified by IsoRanker with predictions from the newest AI model that infers the effects of genetic variants directly from DNA sequence. The model correctly predicted some of the changes we observed with IsoRanker, but it did not consistently capture the full molecular consequences seen in patient RNA. For example, the model did not fully predict degradation of abnormal transcripts via nonsense-mediated decay, or the formation of an unusual fusion transcript caused by disruption of an RNA processing signal. These are precisely the kinds of complex effects that future tools aiming to study the impact of clinically relevant rare genetic variants will need to handle.
Toward a more complete view of genetic disease
Importantly, IsoRanker is not intended to make diagnoses on its own. Instead, it provides another layer of functional evidence that can be considered alongside clinical phenotypes and genomic variants. Tools such as IsoRanker are part of a broader effort at the University of Washington to make functional genomic measurements useful at scale. Specifically, at the Stergachis lab, we are developing technologies and tools that integrate information across the genome, epigenome, and transcriptome to better understand how genetic variants cause disease. By connecting changes in DNA to their downstream molecular consequences, we hope to uncover diagnoses that would otherwise remain hidden.