
On 29 July 2026, Africa CDC, the African Society for Laboratory Medicine (ASLM) and the European Union launched WASTEWISE — Wastewater Molecular and Genomic Surveillance in African Megacities — in Addis Ababa, funded through the EU4Health programme via the European Health and Digital Executive Agency (HaDEA) and the EU’s Health Emergency Preparedness and Response Authority (HERA). The three-year programme will introduce molecular and genomic wastewater surveillance in Ethiopia, the Democratic Republic of the Congo and Kenya — piloted in Addis Ababa, Kinshasa and Nairobi — using a hub-and-spoke model intended to be replicated across the continent as an integrated early-warning system for epidemic-prone diseases.
At first sight this is a story about laboratory capacity. It is also something considerably larger: wastewater surveillance is beginning to move from an exceptional tool deployed during particular public-health emergencies towards a persistent component of the surveillance architecture itself, and that transition changes the problem. When wastewater surveillance consists of a research project examining one pathogen in one city for a limited period, data provenance can often be reconstructed from laboratory notebooks, research databases and the knowledge of the scientists involved. When surveillance becomes infrastructure — operating continuously, across laboratories, jurisdictions, pathogens, analytical pipelines and years — that assumption becomes increasingly fragile.
The critical question is no longer simply can we detect the pathogen? It becomes can we reconstruct the evidence that turned an environmental sample into a public-health assertion? And ultimately: can we establish what evidence existed when a consequential decision was made? That is an infrastructure problem.
Wastewater surveillance is not new. Environmental surveillance has played an important role in poliovirus eradication for decades, but during the COVID-19 pandemic, wastewater-based epidemiology acquired a much broader public profile as countries demonstrated that sewage could provide useful population-level information about viral circulation without depending entirely on individual diagnostic testing.
The principle is deceptively simple: people infected with a pathogen may shed biological material into wastewater, so samples collected from sewage systems can contain fragments of viruses, bacteria, antimicrobial-resistance determinants and other biological markers originating from a population upstream. That transforms an ordinary component of urban infrastructure into something else — a distributed biological observation system.
WHO now describes wastewater and environmental surveillance as a complementary component of multimodal public-health surveillance, and its current guidance explicitly considers its application to multiple pathogens rather than treating it exclusively as a disease-specific technology. European policy is moving in a similar direction: in 2025, the European Centre for Disease Prevention and Control published a framework specifically addressing the integration of wastewater-based surveillance into infectious-disease surveillance and public-health decision-making across the EU/EEA.
The direction of travel is therefore important. Wastewater surveillance is ceasing to be merely an interesting source of supplementary epidemiological information. It is becoming part of how health systems observe populations.
There is an important distinction between detecting a biological signal and characterising it genomically. PCR might tell investigators that genetic material associated with a pathogen is present; sequencing can potentially tell them considerably more. Depending upon the organism, sampling method, concentration, sequencing strategy and analytical pipeline, genomic information may assist with characterising variants, examining pathogen diversity, recognising unusual lineages, comparing environmental findings with clinical isolates or understanding how pathogen populations change over time.
WHO’s International Pathogen Surveillance Network has consequently established a dedicated community of practice concerned with genomics in wastewater and environmental surveillance. Yet WHO also describes the field as comparatively immature, identifying a fragmented knowledge landscape, limited standardisation, lack of consensus over best practice and uncertainty about how genomic wastewater data should ultimately be used by public-health systems as barriers. These are not peripheral administrative problems — they determine what a genomic result actually means.
Consider a statement that might eventually appear on a public-health dashboard: a particular pathogen lineage has increased substantially in City A during the preceding fortnight. Behind that apparently simple sentence may sit a surprisingly long chain of biological and computational events: a sampling location was selected, a sample was collected, the method and time of collection affected what entered the container, the wastewater catchment represented some population (although that population may itself fluctuate), the sample was transported and stored, biological material was concentrated, nucleic acid was extracted, a library was prepared, sequencing occurred on a particular instrument using a particular chemistry, reads passed through quality-control procedures, bioinformatic software transformed raw sequencing data, reference sequences and databases were consulted, thresholds were applied, a lineage or taxonomic assignment was generated, and that result was interpreted alongside historical measurements and perhaps clinical surveillance. Only then did somebody decide that the change was sufficiently meaningful to report.
The final number is not the evidence chain. It is the endpoint of the evidence chain.
It is useful to represent the problem as four broad stages: SAMPLE → SEQUENCE → EVIDENCE → DECISION. Each stage introduces information that must remain associated with what follows.
At the first stage the biological object must be contextualised. Relevant information may include sampling location, wastewater catchment, date and time, collection method, sampling duration, environmental conditions, flow conditions, transport and storage, responsible organisation and laboratory accession identifiers. These details are not clerical decoration — they affect interpretation. A sequence produced from an inadequately characterised environmental sample is not equivalent to the same sequence produced from a well-characterised sample.
The second stage records how biological material became genomic data. Relevant provenance may include extraction protocol, enrichment or concentration procedure, sequencing platform, reagent and library-preparation method, run identifier, quality-control metrics, raw read files, reference materials, processing software and analytical parameters. The genomic dataset is therefore not simply a file. It is the result of a transformation.
The third stage converts genomic data into an interpretable observation, potentially involving read filtering, assembly, mapping, taxonomic classification, variant calling, lineage assignment, abundance estimation, phylogenetic comparison and integration with historical or clinical data. Here provenance becomes particularly important because computational analysis is not static: software changes, databases change, reference genomes change, classification schemes change, thresholds change. A dataset analysed in September 2026 may produce a subtly different interpretation when reanalysed in 2028. That does not invalidate the original conclusion, but reproducibility requires knowing which analytical environment produced it.
Finally, surveillance information crosses from science into governance. A finding may lead to increased sampling, targeted clinical testing, sequencing of additional isolates, epidemiological investigation, risk communication, hospital preparedness, vaccination decisions, inspection, animal-health investigation, food-chain investigation or international notification. At this point the provenance question becomes more consequential: what was known, when was it known, what evidence supported that conclusion, which version of the analysis was available, who interpreted it, and what subsequently changed? These questions matter precisely because surveillance is intended to influence decisions.
Conventional information systems frequently provide some form of technical lineage. A database may record when a field changed. A laboratory information management system may record the movement of a sample. A sequencing platform may retain run information. A bioinformatics workflow may retain computational logs. A surveillance database may record the final result. Each system therefore contains part of the story — but the biological evidence chain does not necessarily remain inside any one of them.
A wastewater sample may begin in an environmental monitoring system, enter a laboratory information system, be sequenced in another facility, have its raw data moved into object storage, be analysed in a bioinformatics environment, have its results deposited in an international repository, be interpreted within a public-health surveillance platform, and have a final decision recorded somewhere else entirely. Every individual system can function correctly while the end-to-end provenance remains fragmented.
This is the distinction between ordinary technical logging and biological provenance. The relevant question is not merely which computer changed this field? It is: how is this scientific assertion related to the biological object, observations and transformations from which it was derived?
The scale of initiatives such as WASTEWISE makes this particularly relevant. A continental or international surveillance architecture will rarely consist of identical laboratories running identical workflows under a single institutional authority — it is much more likely to be heterogeneous. Different laboratories may use different sequencing technologies. Different jurisdictions may retain data under different policies. Different databases may use different identifiers. Methods may change as technologies improve. Samples may be reanalysed. Results may be shared internationally. One laboratory’s genomic observation may be compared with another organisation’s clinical isolate months later.
The central infrastructure challenge is consequently not necessarily centralisation. Attempting to place every laboratory, sequence, clinical record, environmental observation and decision into one enormous database could introduce new problems of governance, access, sovereignty and resilience. A more useful principle is that data can remain distributed while provenance remains connectable. That changes what interoperability means: it is not simply the ability to transfer files, but the ability to preserve the relationships that make those files scientifically meaningful.
Two recent WHO developments are particularly revealing. WHO’s pilot guidance on wastewater and environmental surveillance, published December 2024, identifies challenges involving governance, data use, ethics and long-term sustainability, and argues that wastewater surveillance produces value when it is integrated into decision-making, supported by clear institutional leadership and aligned with national priorities. The WHO Hub for Pandemic and Epidemic Intelligence has separately highlighted wastewater and environmental surveillance as a priority area within its collaborative surveillance work — part of a broader push to strengthen how signals from environmental sources feed into shared, actionable intelligence.
A related regional training programme reported by PAHO/WHO used a telling phrase: “from sequences to decisions.” The central proposition was that producing genomic sequences is only one part of genomic surveillance — genomic data must be interpreted in context and transformed into evidence capable of informing public-health decisions. That distinction is fundamental. The scientific value of genomic surveillance does not reside in the sequence alone. It resides in the chain: biological observation → analytical context → interpretation → decision.
WHO’s global genomic-surveillance strategy similarly treats genomics as part of a wider end-to-end surveillance system encompassing sample collection, diagnostics, analysis and data sharing rather than as an isolated sequencing activity. The evidence problem therefore exists both upstream and downstream of the sequencer.
It is tempting to imagine that preserving raw sequencing data solves the reproducibility problem. It does not: raw data are necessary, but they cannot reconstruct the analytical context on their own. Suppose an environmental sequence is reanalysed three years after the original surveillance event. The investigator might possess the original reads but not know exactly which reference database was used, which release of it, which software version, which filtering thresholds, which lineage definition, which contamination controls, which intermediate results were excluded, or which metadata were available at the time.
The resulting analysis may therefore answer what would we conclude from these reads today? It cannot necessarily answer why did the surveillance system conclude what it did at the time? Those are different questions, and the latter requires provenance.
Good provenance architecture should not make evidence look more certain than it is — quite the opposite, it should allow uncertainty to remain visible. Wastewater is intrinsically complex: a sample represents a mixture, the upstream population may not be known perfectly, pathogens are shed differently, rainfall and wastewater flow may change concentrations, industrial discharges can alter sample composition, different laboratory methods have different recovery efficiencies, and genomic reconstruction from environmental material may be incomplete. A detected relationship between two sequences does not by itself prove an epidemiological transmission pathway.
A well-designed provenance system should therefore preserve relationships without over-interpreting them. It might record that Sequence A was derived from Sample X, that Sequence A was computationally classified as Lineage Y using Pipeline Z, and that Sequence A is genetically similar to Clinical Isolate B according to Analysis C. Those are evidence relationships. They do not require the system to assert that Person B was infected by somebody represented in Sample X. The distinction is enormously important: provenance should preserve scientific reasoning, not manufacture causality.
There is another reason wastewater surveillance is unusually important: it sits at the boundary between environmental surveillance and human health, and when antimicrobial resistance is involved, the boundary expands further. Wastewater may contain organisms or resistance determinants originating from human populations, hospitals, livestock systems, food production, industrial activity or environmental reservoirs, so the relevant investigation may require information from several sectors at once.
This is where wastewater surveillance becomes a genuine One Health problem. In a recent One Health Security analysis, The Sewer Knows First: Wastewater, Early Warning and the One Health Governance Gap, we examined the problem from the institutional side. The argument there was that improving detection does not automatically improve preparedness — a signal becomes useful only if institutions know who receives it, who interprets it, when it escalates and which organisation has authority to act. The article proposes thinking of environmental surveillance as a sensor within a broader One Health architecture rather than as an isolated dataset.
That governance problem and the provenance problem are complementary. One Health Security asks who is responsible for acting upon the signal. The BioChain asks whether the evidence behind that signal can remain connected as it crosses the systems responsible for producing and interpreting it. Both become more important as surveillance becomes more technically sophisticated.
Consider a hypothetical wastewater surveillance event. A municipal treatment plant supplies a composite sample, which receives an identifier; metadata identify the collection period, catchment, sampling method and responsible organisation. An aliquot is transferred to a sequencing laboratory, which performs concentration and nucleic-acid extraction, prepares a sequencing library, and produces raw reads from a sequencing run. Those reads are processed through a defined workflow, which detects a genomic signal compatible with an unusual pathogen lineage. The analytical output becomes a surveillance observation. Public-health investigators compare the observation with clinical sequences; several appear genetically related. The signal triggers additional sampling, and subsequent evidence strengthens the conclusion that the lineage is circulating locally. An alert is issued.
In a provenance-oriented architecture, these are not merely rows spread across several databases. They are connected entities and events: Treatment Plant → Sampling Event → Environmental Sample → Laboratory Receipt → Extraction → Library → Sequencing Run → Raw Reads → Pipeline Version → Genomic Observation → Comparison Analysis → Surveillance Interpretation → Public-Health Action. Each arrow represents something that happened. Each transformation can carry evidence. Each output can refer back to its inputs.
The final alert does not need to contain all of this information. It merely needs to remain cryptographically and semantically connected to it.
Provenance discussions often quickly arrive at the word immutable, and immutability is genuinely useful: if a scientific record can be silently rewritten after the fact, reconstructing history becomes difficult. But immutability alone is not provenance — an immutable database containing poorly described records remains a poorly described database.
The architecture must preserve several things simultaneously: identity (what biological or digital object are we referring to?), relationship (what produced what?), transformation (what process occurred?), time (when did it happen?), authority (who or what asserted it?), integrity (has the record changed?) and context (what metadata are needed to interpret it?). Only when these elements are combined does an immutable history become scientifically useful.
The BioChain is being developed as an evidence and provenance layer for biological data. That distinction matters: it is not intended to replace a LIMS, a sequencing platform, a genomic repository or an epidemiological dashboard, and it does not decide whether a pathogen signal is epidemiologically important. Those systems already exist and, in many cases, perform their specialised functions extremely well. The BioChain addresses the space between them.
Its purpose is to allow biological objects, datasets, analytical transformations and assertions to remain connected as evidence moves through different systems and organisations. In a wastewater-genomics environment, that might mean recording verifiable relationships between an environmental sampling event, the biological sample produced, laboratory processing, sequencing outputs, deposited datasets, analytical pipelines, genomic observations, updated interpretations, related clinical or veterinary observations, and public-health outputs.
The underlying information need not all live inside The BioChain. Large sequence files may remain in specialised repositories, clinical information may remain inside controlled health systems, laboratory records may remain within LIMS, and surveillance dashboards may continue doing what surveillance dashboards do. The provenance layer records the evidence relationships connecting them — a much narrower proposition than constructing a universal biological database, and arguably the more useful one.
There is a danger in describing provenance purely as an audit function. Auditing is certainly one application — if a decision is later challenged, investigators should be able to reconstruct its evidential basis — but scientific provenance provides value before anything goes wrong. It can improve reproducibility, make datasets easier to reuse, clarify which analytical version produced a finding, support collaboration between laboratories, make reanalysis easier, distinguish original observations from subsequent interpretation, reduce ambiguity when data move between organisations, and preserve the intellectual history of a biological conclusion.
That last point becomes increasingly important in the age of computational biology, where many scientific observations are no longer observations in the traditional sense but computationally derived objects. A lineage assignment, resistance prediction or phylogenetic relationship may depend upon a complex pipeline, so scientific infrastructure increasingly needs to preserve not only biological material and raw data but the transformations through which meaning was produced — a point our colleagues at Noviota have made from the representation side: a dataset can remain perfectly machine-readable while losing the distinctions that made it biologically informative in the first place.
The provenance question becomes still more important as machine-learning systems become involved in genomic surveillance. Future surveillance workflows may use AI to prioritise samples, recognise unusual combinations of mutations, identify anomalies across environmental and clinical datasets or determine which observations deserve human attention — increasing analytical capability enormously, but also expanding the evidence chain. A future surveillance conclusion might be derived from sample → sequence → pipeline → model → risk score → analyst interpretation → intervention, at which point reproducibility may require knowing not only the software version but the precise model version, training context, parameters and data inputs used to generate the recommendation. We examined the same expanding evidence chain from the AI-design side in When AI Starts Designing Biology, Who Owns the Evidence Chain? AI does not remove the provenance problem. It makes provenance infrastructure more necessary.
The importance of WASTEWISE and initiatives like it therefore lies not solely in their geographic scale. They signal something deeper: wastewater surveillance is becoming institutionalised. WASTEWISE aims to establish surveillance capacity and coordination structures across major African cities while examining how wastewater data can contribute to early warning and wider One Health surveillance. WHO is developing pathogen-agnostic guidance for multi-pathogen wastewater and environmental surveillance. ECDC is formalising its integration into European infectious-disease surveillance. WHO’s International Pathogen Surveillance Network is explicitly working on genomics within wastewater surveillance.
This is not simply the continuation of a COVID-era experiment. It is the construction of a new layer of biological observation infrastructure, and infrastructure lasts. The samples collected in Nairobi in 2027 may be reanalysed in 2032 using methods that do not yet exist. A genomic signal originally classified as unremarkable may later become important. Reference databases will change, new pathogens will emerge, old datasets will acquire new value, and cross-border investigations may need to revisit evidence collected years earlier. For that reason, the architecture we build for biological surveillance now must do more than produce today’s dashboard — it must preserve tomorrow’s ability to understand how today’s conclusion was reached.
Genomic sequencing has transformed our ability to observe infectious disease. Wastewater surveillance is expanding the population over which that observation can occur. The next challenge is ensuring that increasingly powerful observation does not produce increasingly fragmented evidence.
The desirable architecture is therefore not sample → sequence → dashboard but sample → sequence → evidence → decision, with every important transformation preserved sufficiently that the path can later be reconstructed. The great achievement of modern pathogen genomics is that we can extract biological information from material that would once have told us almost nothing. The challenge of the next decade will be ensuring that we do not subsequently lose the information explaining where that evidence came from, how it was transformed and why we trusted it enough to act. That is not simply record keeping. It is part of the scientific infrastructure of surveillance itself.
No. WHO explicitly frames wastewater and environmental surveillance as part of multimodal surveillance, complementing other sources rather than replacing them. Clinical testing can provide individual-level epidemiological and clinical information that wastewater cannot, while wastewater surveillance can provide population-level signals even where individual testing is incomplete or changing.
Genomics can provide more detailed information than simple pathogen detection. Depending on the organism and analytical method, sequencing can contribute to variant or lineage identification, pathogen-diversity analysis and comparison between environmental and other surveillance observations — though WHO stresses that genomic wastewater surveillance is still developing and faces limitations involving methods, standardisation and data interpretation.
Because the interpretation of genomic data depends on context. The same raw sequencing data may be processed using different software versions, reference databases, quality thresholds and classification systems. Preserving the raw data is important, but reproducing a historical scientific conclusion also requires knowledge of how those data were transformed.
Provenance describes the history and relationships of the evidence — for example, Sample X → Sequencing Run Y → Dataset Z → Pipeline A → Observation B → Public-health report C. The objective is to preserve sufficient information to establish how a scientific assertion arose.
Not exactly. Chain of custody primarily concerns possession, control and handling of an item, often for legal or forensic purposes. Scientific provenance is broader: it includes custody where relevant, but also records transformations, analytical processes, computational derivations, relationships between datasets and the evolution of interpretations.
No, and that would often be undesirable. A provenance layer can reference evidence held in multiple specialised systems while recording verifiable relationships between those objects — sequences can remain in sequence repositories, laboratory information can remain within a LIMS, clinical data can remain in protected health systems, and public-health information can remain in surveillance platforms. The important capability is retaining the evidence graph connecting them.
Not necessarily. Some properties associated with distributed ledgers — tamper evidence, independently verifiable history and cryptographic signatures — can be valuable for provenance, but the scientific problem should determine the architecture rather than starting from a particular technology. Provenance requires trustworthy identity, relationships, transformations, timestamps, authority and integrity, and there are several possible technical mechanisms for providing those properties.
An audit log commonly records what happened within a particular information system. Biological provenance connects events across systems: a LIMS may know that Sample 123 was processed, a sequencing platform knows that Run 456 occurred, a repository knows that Dataset 789 was deposited, a bioinformatics system knows that Pipeline ABC ran, and a surveillance platform knows that Alert XYZ was created. The provenance problem is determining how those records relate to one another scientifically.
Wastewater is environmental evidence about biological processes occurring in a population. Antimicrobial resistance, emerging pathogens and zoonotic disease may require connections with clinical, veterinary, food-system or environmental information, which makes the One Health challenge partly organisational: different sectors need mechanisms for recognising and responding to connected signals. Our accompanying One Health Security analysis, The Sewer Knows First, examines that governance layer in detail.
The BioChain is being designed to provide a provenance and evidence layer rather than replace operational scientific systems. In this context it could preserve verifiable relationships between biological samples, sequencing datasets, analytical transformations, derived observations and subsequent evidence products. The long-term objective is simple: when somebody encounters a biological assertion, they should be able to establish where it came from.
More on One Health
If the institutional side of this problem interests you — who is responsible for acting on a signal, and when — our CEO, Ashley Morgan, also writes the One Health Security blog, covering governance, financing and biosecurity across human, animal and environmental health.
Key takeaways
If your organisation generates, manages, analyses or governs biological evidence and would be interested in participating in a UK or European provenance demonstrator, The BioChain would welcome the conversation.
Get in touch
Regulatory developments, technical notes and platform news — sent occasionally, straight to your inbox.