Data Management Plan

ORCID IDs for senior personnel

  • Katie E. Lotterhos (Northeastern University): 0000-0001-7529-2771
  • Sam Bogan (Skidmore College) 0000-0003-2244-5169
  • Katherine Silliman (NOAA): 0000-0001-5964-3965
  • Rachel Toczydlowski (USDA Forest Service): 0000-0002-8141-2036
  • Serena Caplins (Northeastern University): 0000-0003-1311-6697
  • Alexa Fredston (UC Santa Cruz): 0000-0002-5449-7054
  • Libby Liggins (University of Auckland, New Zealand): 0000-0003-1143-2346
  • Sam Yeaman (University of Calgary, Canada) 0000-0002-1706-8699
  • Pierre De Wit (University of Gothenburg, Sweden): 0000-0003-4709-3438

Description of Expected Research Products

This project will not generate new research data, but may help participants organize existing data and samples. Other types of new data generated during this project include curriculum materials, computer code and pipelines, compute containers, app and data collection templates, and survey data.

  • Existing Samples - Existing samples may include tissue samples, water samples, or whole organisms. Participants will be encouraged to deposit existing samples in a curated sample collection.
  • Existing genomic datasets - Existing genomic datasets include raw reads, derived genomic data (e.g. BAM files or VCF files), or other ascertainment methods (e.g., SNP array). Participants will be encouraged to deposit existing sequence data in International Nucleotide Sequence Database Collaboration (e.g., NCBI) or other genomic data repositories that adhere to FAIR and CARE so that datasets are interoperable (e.g. Aotearoa Genomic Data Repository).
  • Existing seascape genomic metadata - Existing seascape genomic metadata may include spatiotemporal context, oceanographic and environmental data, taxon information, and sample information. Participants will be encouraged to include contextual metadata with physical sample depositions and genomic metadata in GEOME.
  • Curriculum materials - This project will develop new curriculum materials, which includes videos and written tutorials published on the MarineOmics webpage.
  • Compute containers, code, and pipelines - This project will develop new or organize existing code and pipelines with compute containers. Code will be annotated to promote interpretation and detailed headers will be used to explain the purpose and functions of the scripts, input variables needed, etc.
  • App and Data Collection Templates - This project will develop new or organize existing data collection templates, including spreadsheet templates and AppSheet app templates.
  • Survey data - Survey data includes confidential or anonymous participant responses to IRB-approved questionnaires.

Metadata Formats and Standards

This project will develop a framework and foundation for the adoption of international data standards for seascape genomics (meta)data, including:

  • Darwin Core standards are used to describe the spatiotemporal context for biodiversity data. This standard is used by the Global Biodiversity Information Facility (GBIF), the Ocean Biodiversity Information System (OBIS), GEOME, and the Global Genome Biodiversity Network.
  • Metadata standards and vocabularies for the minimum information about any (x) sequence (MIxS) standards are integrated with Darwin Core and have been adopted by INSDC and other genomic data repositories to describe the what, where, when, how, and by whom a genetic sample was collected. Darwin Core and MIxS also have integrated extensions, such as the DNA-derived data extension, which is used to store information related to DNA from samples (an organism or the environment). It can be used to describe the environment the sample came from, how the tissue was collected, parasites or diseases in the sample, DNA extraction and sequencing, and genome attributes of the sample.
  • BioSample Attributes is a set of terms used to describe biological samples from specimens (such as tissue, blood, urine, etc.) that are sequenced or used in experimental assays. Popular sequencing repositories such as GenBank and the Sequence Read Archive use these attributes to describe the type of material that was sequenced.
  • The Global Genome Biodiversity Network (GGBN) Data Standard is a set of terms and controlled vocabularies that builds on Darwin Core and MIxS to describe genomic samples in biobanks. It includes terms for describing cloning, amplification, gel images, storage buffer, collection permits, and copyrights.
  • Docker containers contain metadata in standard json files that describe the name, description, and version of the software.
  • Minimum Information for an Omic Protocol (MIOP) standards describe the what, when, where, and how of sample processing and have been adopted by Better Biomolecular Ocean Practices (BeBOP) protocol templates.
  • The Internet of Samples (iSamples) has vocabularies to record metadata about material samples, such as the material sample type, material type, and sampled feature type, and also persistently link them to other samples and derived digital content, including images, data, and publications.
  • The Marine Metadata Ontology Registry and Repository hosts a thesaurus of standardized vocabularies related to the marine environment.
  • The Natural Environment Research Council Vocabulary Server (NVS) provides controlled vocabularies for oceanographic and similar data. It includes a broad array or terms for physical data (such as climate or forecasts) and biological data (such as abundance of biological entities per unit area or unit volume).

Policies for Access and Sharing

Curriculum materials, compute containers, and data templates will be made publicly available as described in the proposal. We will maintain the privacy and confidentiality of participant’s survey data following IRB-approved protocols. Survey data will be aggregated to protect participants from pseudo-anonymity. For existing data, we will help participants develop data/sample sharing agreements as necessary.

In most cases, products will be made publicly accessible upon publication or within two years of collection, but we recognized that in some cases, individual datasets may need to be embargoed for a longer period of time. We will develop embargo and data policies for special cases such as endangered species.

Authorship Policy

For authorship and reporting guidelines we will refer participants to the Vancouver Recommendations: a set of recommendations for best practice and ethical standards in the conduct and reporting of published research. They recommend that all 4 of the following criteria are required for authorship: (i) Substantial contributions to the conception or design of the work; or the acquisition, analysis, or interpretation of data for the work; AND (ii) drafting the work or revising it critically for important intellectual content; AND (iii) final approval of the version to be published; AND (iv) agreement to be accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. They also recommend that when a large multi-­author group has conducted the work, the group ideally should decide who will be an author before the work is started.

Policies for Re-Use and Re-Distribution

All products generated by this project will be made freely available for research, educational, management, and non-profit purposes in accordance with University/Participating institutional and NSF policies, and in consultation with Indigenous communities who have sovereignty of some data to respect their interests and aspirations for the data. Consistent with normal scientific practices and publication procedures, publication of data shall occur as the project proceeds or after the project ends.

Plans for Archiving, Data Storage, and Preservation

Our plans for archiving and preservation of existing samples and data will follow the FAIR practices outlined in the table below or the recommendations we develop in this project. We will work with investigators to ensure cross-linking among the various DOIs associated with their data. All the curriculum materials, templates, and compute containers that we develop will be linked to the MarineOmics page, where they will be publicly available. The table below summarizes our current plans for data archival, but these are subject to evolve over the course of the project period as we determine best practices.

Step in data lifecycle Current data management practices FAIR practices and overview of Open Science plans
Project ID Not applicable. The foundation project number and grant DOI number will be assigned upon grant approval. We will make sure to cite these numbers in all research products and publications using the following language, “This work was supported by the Gordon and Betty Moore Foundation, GBMF#####, grant DOI.”
Archive sample metadata Published in individual repositories or as supplemental material, but rarely sufficient metadata and often not interoperable or accessible Metadata DOI: Deposition in GEOME and linked to GBIF/OBIS global repositories. GEOME mints Globally Unique Identifiers (GUIDs) for projects within a team. A Globally Unique Identifier (GUID) is a 128-bit number used in computing to uniquely identify information, objects, or entities across systems, databases, and networks without a central registration authority. For biological data, we will use Darwin Core, MIxS, and other standards. For environmental/oceanographic data, use standards from the Marine Metadata Ontology Registry and Repository and Natural Environment Research Council Vocabulary Server; Develop guidelines/plans for enhancing interoperability.
Archive tissue samples Samples stored in freezers of individual research groups Sample DOI: Use Darwin Core, MIxS, and other standards; Deposition in the Ocean Genomic Legacy biorespository and will link to GEOME and other repos. GEOME mints Biodiversity Collection Identifiers (bcid) for samples. BCIDs (Biodiversity Collection Identifiers) are persistent identifiers in ARK (Archival Resource Key) format, using the NAAN (Name Assigning Authority Number) 21547 for genomic and biodiversity data. These identifiers, often structured as ark:/21547/, are used to permanently link, resolve, and cite digital biodiversity samples.
Archive molecular protocols Lab protocols often customized/optimized for specific species, but not published Protocol DOI: New standards for describing protocols were recently developed through the UN Decade of Ocean Science for Sustainable Development (Samuel et al. 2021). The Minimum Information for an Omic Protocol (MIOP) standards and accompanying Better Biomolecular Ocean Practices (BeBOP) protocol templates have been developed to meet these new standards. Protocols use Minimum Information for an Omic Protocol (MIOP) standards and publicly available with DOI through BeBOP.
Archive sequence data Sequence data uploaded to Sequence Read Archive. No standard practices for derived genetic data (e.g., BAM or VCF files). Sequence DOI: Raw reads with associated metadata uploaded to Sequence Read Archive
Archive code / bioinformatics Code published in public git repository; code often not reproducible; no system for bioinformatics information management Pipeline/Code DOI and Container DOI: We will build compute containers for all code on GitHub and archive them with Zenodo or DockerHub, which will provide DOIs
Cross-link all project-related DOIs Not practiced Develop guidelines/plans to enhance cross-talk among all above DOIs associated with a single dataset