O2A SEQUENCES
O2A SEQUENCES provides a single platform for the full metabarcoding lifecycle: it keeps samples, metadata, and sequences linked, runs containerized workflows that produce provenance-tracked ASV tables, generates submission-ready exports for public archives, and delivers AI-ready outputs.
Description
Metabarcoding, the sequencing of marker genes from environmental samples to identify the organisms present, is central to biodiversity and environmental research across the Helmholtz community. Yet much of the value in the resulting sequence data stays unused: workflows are fragmented across incompatible tools, metadata are inconsistent, analyses rarely survive reuse beyond the group that produced them, and preparing data for submission to public archives is tedious and time-consuming. The O2A Data Flow Framework (Observations to Analysis and Archives; https://o2a-data.de/), first introduced in 2015 and now operated as a mature research-data infrastructure, already supports instrument metadata, data ingestion, collaborative analysis, publication and discovery. The newer O2A SAMPLES (https://samples.o2a-data.de/) service extends this ecosystem to physical-sample management and integrates MIxS-aligned metadata structures developed through the HARMONise (https://harmonise.awi.de/) project for metabarcoding samples. What is missing is the sequence data itself, which is not yet connected to this operational chain. O2A SEQUENCES (https://sequences.o2a-data.de/) closes this gap by developing the module and interfaces that let O2A manage samples, metadata, and sequences together, run standardized analysis workflows, and generate submission-ready files for public archives. Provided as software-as-a-service and available for self-hosting, each institution can run on its own infrastructure, O2A SEQUENCES makes metabarcoding data FAIR, AI-ready, and reusable across centers, providing a foundation for sustainable institutional research data management that extends to other environmental data types.