ASN's Mission

To create a world without kidney diseases, the ASN Alliance for Kidney Health elevates care by educating and informing, driving breakthroughs and innovation, and advocating for policies that create transformative changes in kidney medicine throughout the world.

learn more

Contact ASN

1401 H St, NW, Ste 900, Washington, DC 20005

email@asn-online.org

202-640-4660

The Latest on X

Kidney Week

Abstract: SA-PO0129

The PKD RRC DataHub: A Harmonized Transcriptomic Resource to Accelerate Polycystic Kidney Disease Research

Session Information

Category: Genetic Diseases of the Kidneys

  • 1201 Genetic Diseases of the Kidneys: Cystic (Monogenic)

Authors

  • Soelter, Tabea M., The University of Alabama at Birmingham Department of Cell Developmental and Integrative Biology, Birmingham, Alabama, United States
  • Crumley, Anthony B., The University of Alabama at Birmingham Department of Cell Developmental and Integrative Biology, Birmingham, Alabama, United States
  • Casper, Kieran James, The University of Alabama at Birmingham Department of Cell Developmental and Integrative Biology, Birmingham, Alabama, United States
  • Okunbor, Elizabeth, The University of Alabama at Birmingham Department of Cell Developmental and Integrative Biology, Birmingham, Alabama, United States
  • Lasseigne, Brittany N., The University of Alabama at Birmingham Department of Cell Developmental and Integrative Biology, Birmingham, Alabama, United States
Background

The "Omics Revolution" offers unprecedented insights into polycystic kidney disease (PKD), but reusing public-domain transcriptomic data is hampered by scattered data deposition, a lack of uniform processing, and resource-intensive metadata curation. Furthermore, generic computational cores often lack the domain-specific knowledge required to interpret complex PKD models, creating a significant bottleneck for researchers. To address this, the Informatics and Data Analytics Resource (IDAR) of the PKD Research Resource Consortium (RRC) is developing the PKD DataHub.

Methods

For the initial Minimum Viable Product (MVP; DataHub 0.1), we systematically identified and downloaded 14 public-domain bulk Illumina mouse PKD gene expression datasets from NCBI GEO. To achieve full harmonization, all datasets were processed consistently using the best-practice nf-core/rnaseq bioinformatics pipeline. Metadata, including study characteristics (e.g., model, age, induction, sex) and quality control (QC) metrics (e.g., percentage of mapped reads, sequencing depth), were compiled for each dataset.

Results

DataHub 0.1 was deployed as a user-friendly, point-and-click web portal. The platform enables researchers to query and filter the 14 harmonized mouse PKD datasets using searchable metadata and comprehensive QC metrics. This bypasses the need for advanced bioinformatics domain knowledge and computational analysis skills. After dataset identification, users can retrieve pre-processed datasets via the open-access repository Zenodo. To ensure maximum robustness, reliability, and usability, the software underwent rigorous, multi-phase code review and user testing.

Conclusion

The PKD RRC DataHub provides a highly curated, QC-verified repository that democratizes access to genomic analysis, enabling any researcher to generate hypotheses regardless of coding ability. By significantly reducing the time and analytical barriers to locating, processing, and assessing data suitability, the DataHub speeds up hypothesis generation and validation for wet- and dry-lab PKD researchers. Future iterations will expand the DataHub to include additional public-domain datasets with other sequencing types (e.g., single-cell, long-read) and other species (e.g., human).

Acknowledgment

Data analysis was performed using custom scripts generated with the assistance of Claude Sonnet 4.6 (Anthropic); the authors take full responsibility for the integrity of the generated code and resulting data.

Funding

  • NIDDK Support