Connect samples, laboratory processing, sequencing, analyses, datasets, publications and subsequent revisions — preserving the path from source material to research conclusion without making one laboratory or database responsible for everybody else’s history.
A single research finding is rarely produced by one system. A sample is collected in the field or clinic. It is processed in a laboratory that may or may not be the one that later sequences it. Sequencing may happen at a core facility or an external provider. The raw output moves through an analytical pipeline, often a specific version, with specific reference data and parameters. The results feed a dataset, which may be deposited in a public or institutional repository. A publication interprets that dataset. Later, the pipeline is updated, an error is found, or a new analysis reinterprets the same underlying data — and a revision or correction follows.
Each of those steps typically happens inside a system built for that one job: a LIMS, a sequencing platform, an analysis environment, a data repository, a journal. None of them was designed to preserve the relationships connecting it to the others.
The BioChain does not replace any of those systems, and it does not perform the analysis. It records the provenance relationships between them — which sample produced which sequencing run, which pipeline version produced which result, which dataset supports which publication, and what changed when a revision was made. That record survives independently of whether the original systems are still running, still licensed, or still remembered by the people who used them.
If a published finding is questioned two years from now, can anyone actually reconstruct how it was produced?
For a single laboratory, that question is often answerable with enough effort. Across several institutions, multiple software versions and a handful of staff changes, it usually is not — not because anyone did anything wrong, but because no one was ever responsible for the connections between systems, only for their own system.
The five steps are the same everywhere. Here is what each one connects in this context.
No. The BioChain sits alongside those systems and records the relationships between the records they hold. You keep using the systems you already have.
No. That is specifically the problem this is designed to avoid. Each participant can continue operating its own systems; The BioChain connects the provenance between them.
That is exactly the kind of detail the provenance record captures — which version produced which result — so the difference itself becomes visible and traceable, rather than a silent gap.
No. Genomics is one of the clearest examples because pipelines and reference data change often, but the same approach applies to any multi-stage research workflow.