Maximizing the Value of Rigorous Research Through NIH-Funded Data and Biospecimen Repositories, with NIH Institute and Center Perspectives
NIH supports an array of data and biospecimen repositories that help ensure valuable scientific resources remain available to qualified researchers for future studies. Many repositories link biospecimens to rich clinical, phenotypic, imaging, genomic, or longitudinal datasets, allowing researchers to validate findings across multiple data types and apply modern analytical approaches. Here, we briefly discuss the policies and principles governing data management and repositories. Then, we spotlight selected examples of NIH-supported repository programs.
When appropriately preserved, curated, and shared, these data and repositories enable and accelerate new discoveries, support validation of research findings, strengthen transparency, and extend the impact of NIH-supported research investments far beyond the original study. As such, these resources play a critical role in advancing scientific rigor, reproducibility, transparency, and responsible stewardship of participant contributions.
Our support for these resources further:
- Facilitates the FAIRness of the data (Findable, Accessible, Interoperable, and Reusable)
- Embeds quality assurance throughout the resource lifecycle, from collection and processing to curation and long-term preservation
- Pairs specimens and datasets with robust metadata, documentation, and provenance information, allowing investigators who were not involved in the original study to confidently interpret and reuse resources
- Extends the value of NIH-funded research and reduces the need to collect duplicate data or specimens by enabling secondary analyses, validation studies, and new scientific questions
- Provides secure environments for analysis, collaboration, and discovery
- Uses research participant data and specimens appropriately to support future scientific advances, including moving towards more human-centered approaches to study mechanisms of human disease (an NIH priority)
Repositories supported across NIH share common commitments to scientific rigor, responsible stewardship, and maximizing the value of participant contributions. They also can vary in their scientific focus, collections, and access models. Example characteristics may include:
- Mission-specific resources tailored to the scientific disciplines and research communities
- Integrated research platforms with analysis workbenches, artificial intelligence tools, reference materials, and/or curated collections
- Flexible participation models
- Connected research ecosystems linking biospecimens with clinical, imaging, genomic, and other complementary data
- Specific governance that considers scientific access and unique needs of the resource and research community
The selected examples below from NIH Institutes, Centers, and Offices (ICOs) (in no particular order) illustrate how different ICOs operationalize these principles through their repository programs. Potential interested researchers should appreciate that each repository has its own scientific scope, submission policies, and access models to follow. Selecting the ICO name will provide a drop-down with additional information.
The All of Us Research Program is dedicated to advancing scientific rigor and resource sharing by pairing physical specimens with one of the most deeply phenotyped cohorts in biomedical research. The All of Us Research Program Biobank embeds rigor and reproducibility across the data lifecycle, from standardized biospecimen collection and centralized processing to comprehensive data curation and researcher access. Through our uniform operating procedures and robust curation, we can ensure consistency and usability across sites and users.
Relatedly, the Researcher Workbench provides more than 23,000 investigators with secure access to rich, multimodal data and analytical tools. Access is governed through institutional agreements, identity verification, and required training, with tiered data protections to safeguard participant privacy. Investigators may also request biospecimens through the X01 opportunity, with resulting data returned to the Workbench to enhance interoperability and maximize long-term research value. The All of Us profile page shares more information.
The NIDDK Central Repository curates and distributes matched clinical data and biospecimens spanning diabetes, digestive diseases, kidney disease, obesity, liver disease, urologic diseases, and related conditions. Rigor, reproducibility, and responsible stewardship are embedded throughout the repository lifecycle through rigorous quality assurance, FAIR data practices, standardized data and specimen submission guidance, written archival and sharing plans, and cross-linkages with NIH-supported repositories. Two complementary infrastructure innovations further expand access to these data and biospecimens:
- The Analytics Workbench provides a secure cloud-based environment for analyzing controlled-access data.
- The Challenge Management Platform supports the full lifecycle of NIH prize competitions.
Submission to the repository is limited to eligible NIDDK-supported studies in accordance with repository policies and program requirements. Together, these repository practices enable investigators worldwide to test new hypotheses, replicate published findings, and conduct complex multistudy analyses using well-documented, reusable data and biospecimens. Access is available to qualified investigators who execute a data use agreement, demonstrate consistency with original informed consent, and meet NIH data security requirements. Overall, these NIDDK resources facilitate proper long-term stewardship of the investment from NIH, investigators, and research participants. NIDDK's profile page shares more information.
NIA stewards multiple biorepositories that advance NIH expectations for rigor, reproducibility, and data quality, maximizing their value to the aging research community. These repositories house biospecimens and data from multiple focus areas, including aging biology, behavioral and social research, geriatrics and clinical gerontology, and neuroscience. Examples include:
- The AgingResearchBiobank, which contains biospecimens, images, and associated clinical and phenotypic data from NIA-supported studies of age-related conditions
- The National Centralized Repository for Alzheimer’s Disease and Related Dementias, which preserves biospecimens and data for the study of dementia, including etiology, prevention, and treatment
NIA provides transparent and secure access to these and other repositories at no cost to researchers, accelerating discovery in aging science. By preserving specimens beyond the lifespan of individual projects, standardizing collection and management practices, and supporting secondary analyses, NIA ensures the long-term value of these resources to the aging research community. NIA's profile page shares more information.
NIMH promotes FAIR, reproducible open science by enabling reuse of data, software, and models resulting from our scientific investments. The examples below illustrate standards that support secondary analyses, rigorous data curation and quality control, required biospecimen sharing, study harmonization, cost efficiencies, and continuous improvement through advisory input. Standardized storage and cataloging procedures further enhance discoverability and reuse, enabling precise experimental replication and strengthening stewardship of taxpayer-funded resources.
- The centralized NIMH Repository and Genomics Resource includes over 3.8 million biospecimens available to researchers, accelerating research on the genetic architecture of mental illness.
- The NIH NeuroBioBank and NIMH Human Brain Collection Core provide qualified investigators with access to post-mortem brain tissue and related biospecimens.
- The NIMH Data Archive supports compliance with NIH Data Management and Sharing expectations while enabling controlled access to human subject-level data to advance secondary analyses and reproducibility across scientific domains.
NIMH's profile page shares more information.
NIBIB supports imaging data repositories and research infrastructure that maximize the scientific value of NIH-funded research by advancing rigor, reproducibility, and responsible data stewardship. For example, the national Medical Imaging and Data Resource Center supports the collection, standardization, and sharing of medical imaging data and artificial intelligence tools. This includes sequestered datasets and Medical Device Development Tools that facilitate algorithm validation and innovation.
NIBIB-funded resources promote standardized data formats, FAIR principles, harmonized metadata, quality assurance, and comprehensive documentation to enable researchers to interpret, validate, and reuse imaging datasets across institutions and studies. Access frameworks incorporate de-identification, data use agreements, and appropriate governance to support broad use by qualified investigators while protecting sensitive information. Investigators are encouraged to plan for long-term stewardship, interoperability, and data sharing from the outset to maximize the impact and reuse of NIH-supported research resources. Altogether, NIBIB’s aim is to increase the value and efficiency of these publicly funded resources for future scientific discovery, reduce unnecessary duplication, and promote collaboration across disciplines and institutions. NIBIB's profile page shares more information.
NIAID supports scientific rigor, reproducibility, and responsible stewardship by centralizing access to high-quality research materials through the Biological and Emerging Infections (BEI) Resources. This centralized biomaterials resource is a trusted source for authenticated infectious agents and derived reagents and related services for the microbiology and infectious disease research community. It allows for strengthened domestic biomanufacturing capabilities and credible, reproducible research. Moreover, its standardized production, validation, documentation, and distribution processes are disease-agnostic, making the program a valuable biomaterial manufacturing and supply chain resource that can support broader NIH priorities. Access is governed through authorized-based registrations and biosafety requirements, while investigators are encouraged to deposit high-value materials to maximize reuse and long-term scientific impact. Collectively, this infrastructure advances NIH-wide goals, ensuring that high-quality biomaterials remain accessible, reusable, and supported by the quality standards needed for impactful science. NIAID's profile page shares more information.
NIAMS advances scientific rigor, reproducibility, and responsible stewardship through transparent governance of its biospecimen resources and by promoting broad access to data generated through the research it supports. Together, these efforts maximize the long-term value of federally funded biospecimens and datasets by enabling qualified investigators to generate new discoveries while ensuring resulting data are shared for future research.
- The Osteoarthritis Initiative maintains a longitudinal natural history database with clinical assessments, radiologic images, and biospecimens. An oversight committee evaluates access requests based on scientific merit, investigator qualifications, data-sharing plans, and the potential to advance osteoarthritis research. Generated data are returned to the resource, extending its value for the broader research community. See also this Highlighted Topic.
- The Archiving and Sharing Skeletal Phenotyping Data Project provides freely accessible skeletal phenotyping datasets. It promotes standardized submission guidance, workflows, and common data elements that improve interoperability, reproducibility, and reuse across independent studies.
More is available in the NIAMS FY 2025-2029 Strategic Plan and on the NIAMS profile page.
NIDCR-supported biospecimen and data repositories pair well-characterized biospecimens with standardized protocols, comprehensive metadata, and rigorous quality assurance practices to maximize their long-term scientific value. The Data-Driven Science Hub serves as a centralized resource that guides investigators to scientific data, biospecimens and other experimental materials, and resources and tools for data science-driven research and training.
Common best practices include standardized specimen collection and handling, use of the NIDCR Collaborative Common Data Element Framework and other data standards, and documentation that supports reproducibility, interoperability, and secondary analyses. These principles are reflected across NIDCR-supported resources like:
- The Sjögren's International Collaborative Clinical Alliance repository provides saliva, blood, tissue, DNA, and other biospecimens linked to comprehensive clinical, imaging, laboratory, and patient-reported data for reproducible studies and biomarker discovery in Sjögren's disease.
- FaceBase offers curated datasets with persistent identifiers, standardized metadata, and FAIR data practices for studies of craniofacial development, congenital anomalies, and related disorders.
- The expanded Human Oral Microbiome Database provides freely accessible curated genomic and taxonomic reference data for microorganisms throughout the human digestive tract, enabling consistent identification and comparison of microbial communities across studies.
NIDCR repository access frameworks are designed to maximize scientific value while protecting participant privacy and ensuring responsible stewardship. Many resources provide open access to nonsensitive datasets while overseeing the access to individual-level human data or biospecimens through controlled-access or sample custodian’s approval, entailing investigator applications, compliance with NIH data security and use requirements, and approval by the NIDCR Data Access Committee for the data access or the biospecimen custodians. Together, these well-documented, interoperable resources reduce unnecessary duplication of specimen collection, enable independent validation of findings, foster new collaborations, and accelerate scientific discovery. NIDCR's profile page shares more information.
NINDS co-funds and co-manages several biorepositories with other ICOs that operationalize rigor, reproducibility, and data quality through standardized biospecimen collection, harmonized data elements, standardized neuropathological assessments, cell line processing, and centralized quality control processes. All NINDS-supported repositories use transparent access policies, open-access biospecimen catalogs, standardized request review processes, and provide broad availability of specimens to the research community. Two examples include:
- The Biospecimen Exchange for Neurologic Disorders banks and distributes cerebrospinal fluid, plasma, serum, DNA, and RNA from NINDS-supported natural history studies and clinical trials, with linked clinical and phenotypic data.
- The NIH NeuroBiobank holds well-characterized post-mortem brain tissue and whole genome sequencing spanning neuropsychiatric and neurological disorders across the lifespan.
Access requests from qualified investigators for brain tissue, blood fractions, and other nonrenewable resources are reviewed for sample size justifications, power analyses, supporting pilot data, and consistency with the donor consent. When needed, the repositories can assist with designing tissue requests that maximize stewardship of donor tissue while meeting the requestors experimental needs. Encouraging and using broad sharing consent language, standardized best practices, and Material Transfer Agreements further facilitate cross-institutional collaborations and sustained sample integrity to maximize the long-term value and reuse potential of collected biospecimens. NINDS' profile page and the Cell/Tissue/DNA page shares more information.
NLM enables public access to NIH-funded research through a data repository infrastructure that spans the data lifecycle from initial submission through public release to support long-term reuse. In particular, the NLM National Center for Biotechnology Information manages a portfolio of biomedical data repositories, including resources for publications, clinical trials, and genetic data. Examples include:
- BioSample collects standardized metadata about biological samples and links to submitted sequence and other data using machine-readable standards that promote data reuse.
- GenBank contains analyzed and curated nucleotide sequence data enriched with gene and protein annotations, functional information, linked publications, and biological source metadata.
- Sequence Read Archive has raw sequence data and associated metadata prior to curation or downstream analysis, enabling independent validation and reproducibility.
NLM’s practices align with FAIR data principles, quality assurance processes, and robust, secure data management and access procedures (including for ensuring protection of collected human data). By properly managing these complementary repositories, we maximize the value of NIH's investment in research data generation and resource stewardship. Together, these resources accelerate scientific discovery, support rigor, reproducibility, transparency, and data quality, and enable long-term reuse of NIH-funded research data. NLM's profile page provides additional information.
The NIGMS Human Genetic Cell Repository currently houses nearly 12,000 unique cell lines, 6,000 DNA samples, and 140 human induced pluripotent stem cell lines representing inherited disorders and global populations. Extensive physical, molecular, and genetic characterization data are publicly available through an online catalog to help researchers identify validated resources. To ensure scientific rigor and reproducibility, the repository follows a certified quality management system, standardized procedures, international best practices, and secure sample tracking, with comprehensive molecular and genetic characterization uniquely identifying each sample. Responsible stewardship is supported through updated informed consent requirements, donor privacy protections, material transfer agreements for qualified researchers worldwide, and a cost-recovery model that reinvests revenue into maintaining the repository's infrastructure. Together, these practices enable broad resource sharing while ensuring the quality, integrity, and long-term sustainability of this valuable NIH-supported resource. NIGMS profile page shares more information.