How LDbase Meets NIH Guidelines

Summary: As part of the National Institutes of Health’s (NIH) push towards supporting open science practices and data sharing, the NIH has put out guidance on how to select a data repository and the characteristics one should look for when choosing where to share data. This document describes how LDbase meets such standards, providing links to resources where necessary.

 

  1. Desirable Characteristics for All Data Repositories.

    The characteristics in this section are relevant to all repositories that manage and share data resulting from Federally funded research:

    1. Unique Persistent Identifiers: Assigns datasets a citable, unique persistent identifier (PID), such as a digital object identifier (DOI) or accession number, to support data discovery, reporting (e.g., of research progress), and research assessment (e.g., identifying the outputs of federally funded research). The unique PID points to a persistent landing page that remains accessible even if the dataset is de-accessioned or no longer available. 

       

      LDbase provides the option to assign unique digital object identifiers (DOI) for all projects and associated files (datasets, codebooks, code, etc.) free of charge to the user.

    2. Long-Term Sustainability: Has a plan for long-term management of data, including maintaining integrity, authenticity, and availability of datasets; building on a stable technical infrastructure and funding plans; and having contingency plans to ensure data are available and maintained during and after unforeseen events.

       

      LDbase was built with sustainability in mind. Keeping ongoing costs low was the plan to start, which is why we don’t do things like check inside datasets stored on LDbase. LDbase is a collaboration with a large team that includes FSU Libraries, which includes LDbase as part of its collection. As such, the FSU Libraries have guaranteed long-term management of LDbase and the data stored there, as part of its collection.

    3. Metadata: Ensures datasets are accompanied by metadata to enable discovery, reuse, and citation of datasets, using schema that are appropriate to, and ideally widely used across, the community(ies) the repository serves. Domain-specific repositories would generally have more detailed metadata than generalist repositories.

       

      All datasets, as well as projects and other associated files (codebooks, code, stimuli, etc.) have a wide range of metadata available to increase the discovery, reusability, and citability of said files (See requested metadata here). Further, projects and associated files are structured such that users can search for terms in the metadata within any given file’s metadata, not just datasets alone. LDbase is domain-specific and provides users who are uploading data or files with extensive guidance, examples, and the ability to select from a pre-populated list of common terms within the field.

    4. Curation and Quality Assurance: Provides, or has a mechanism for others to provide, expert curation and quality assurance to improve the accuracy and integrity of datasets and metadata.

       

      In keeping with our ethos of long-term sustainability, LDbase has been deliberately created and updated to have user features that support data curation and quality assurance with no permanent human oversight by LDbase staff. We use well documented bespoke metadata fields as well as considerable help documentation to assist data depositors. In addition, we publish manuscripts as well as provide an extensive and searchable "data sharing resources" page, we provide monthly Q&As (see the landing page for the next session!), and the researchers and research support staff provide consulting (e.g., Within & Between Consulting), and trainings (e.g., DMDS workshop). Finally, while actively funded by federal funders, we provide free consulting services for data depositors via a hotline email. 

       

    5. Free and Easy Access: Provides broad, equitable, and maximally open access to datasets and their metadata free of charge in a timely manner after submission, consistent with legal and ethical limits required to maintain privacy and confidentiality, Tribal sovereignty, and protection of other sensitive data.

       

      LDbase is an entirely free data repository for both data uploading or downloading/reuse. All policies and procedures are in accordance with the relevant legal and ethical guidelines.

      LDbase does not offer an "embargo" (i.e., restricted access option) option that meets NIH criteria for controlled access sharing (i.e., data that have limitations on their use imposed by laws, regulations, policies, informed consent, and/or agreements; are sensitive; involve risk of harm; or cannot be sufficiently de-identified to established standards). Any NIH researcher that require such controls for sharing their data should consult the NIH list of repositories that allow this option.

    6. Broad and Measured Reuse: Makes datasets and their metadata available with broadest possible terms of reuse; and provides the ability to measure attribution, citation, and reuse of data (i.e., through assignment of adequate metadata and unique PIDs).

       

      All datasets and associated files can be easily downloaded and reused, and further are assigned DOI’s allowing for the ability to measure the attribution, citation, and reuse of said data and files. A unique aspect of LDbase is the project-centered structure, allowing users to upload information and create unique identifiers for the project, datasets, codebooks, code, stimuli, measures, or any other files that are associated with said larger project. As such, LDbase allows for broad and measured reuse of multiple components of a project and encourages reuse and citing of more than just data. LDbase also tracks and publishes downloads and page views.

    7. Clear Use Guidance: Provides accompanying documentation describing terms of dataset access and use (e.g., particular licenses, need for approval by a data use committee).

       

      LDbase does not require any specific process for dataset access and use. This is entirely dictated by data depositors. To assist this, LDbase provides extensive guidance, resources, and FAQ’s related to dataset uploading, access, terms of use, and licensing. Additionally, documentation to support usage are available related to IRB considerations, data management and deidentification, data reuse and combination, and general information related to open science practices.

    8. Security and Integrity: Has documented measures in place to meet generally accepted criteria for preventing unauthorized access to, modification of, or release of data, with levels of security that are appropriate to the sensitivity of data.

       

      LDbase has documented measures in place to provide the necessary security and privacy required for the stored data and information. Documentation is available on LDbase related to these measures, including a security/privacy focused FAQ, terms of service, and documentation related to the specific systems and processes being used to ensure security and integrity.

      LDbase does not offer an "embargo" (i.e., restricted access option) option that meets NIH criteria for controlled access sharing (i.e., data that have limitations on their use imposed by laws, regulations, policies, informed consent, and/or agreements; are sensitive; involve risk of harm; or cannot be sufficiently de-identified to established standards). Any NIH researcher that require such controls for sharing their data should consult the NIH list of repositories that allow this option.

    9. Confidentiality: Has documented capabilities for ensuring that administrative, technical, and physical safeguards are employed to comply with applicable confidentiality, risk management, and continuous monitoring requirements for sensitive data.

       

      LDbase has documented measures in place to ensure confidentiality. Again, documentation is available on LDbase related to these measures, including a security/privacy focused FAQ, terms of use, and documentation related to the specific systems and processes being used to ensure security and integrity.

      LDbase additionally uses Amazon Web System’s (AWS) S3 File system to store files uploaded onto LDbase, which has its own internal slew of procedures and safeguards to help ensure confidentiality. Files are backed up nightly and stored in AWS Glacier, which is designed for the specific purpose of confidentially archiving data. AWS S3 Glacier has been approved for storing even the most sensitive data, including medical images and genomic data, The use of this system further allows for versioning of data to increase the accessibility and confidentiality of data.

    10. Common Format: Allows datasets and metadata downloaded, accessed, or exported from the repository to be in widely used, preferably non-proprietary, formats consistent with those used in the community(ies) the repository serves.

       

      LDbase explicitly recommends non-proprietary formats. LDbase supports an extremely wide range of file formats (both for datasets and additional documentation) including but not limited to those that are commonly used within the field. For file formats that are currently not supported, LDbase allows individuals to request new file formats for support. As such, it is a goal of LDbase to support any and all formats to the extent that file format will provide no barrier to successfully using LDbase for data sharing.

    11. Provenance: Has mechanisms in place to record the origin, chain of custody, and any modifications to submitted datasets and metadata.

       

      Each dataset and related documentation has a field where authors can be added, which can be seen when looking at the dataset/documentation on LDbase, as well as in the recommendation citation for the dataset and related documentation. Every dataset and related documentation that is uploaded to LDbase can be versioned, so that a user can identify which version of data they have used (and can cite the appropriate version). Any user can also request to be notified if a given dataset or document has been updated.

    12. Retention Policy: Provides documentation on policies for data retention within the repository.

       

      Policies related to data retention within LDbase are clearly outlined in the extensive resources and FAQ’s available on the website. Information specific to data retention are available in the Terms of Service and Data Security FAQ.