
This is the complete text of The BioChain’s EU policy white paper, From Data Governance to Evidence Governance: Building Verifiable Evidence Chains for European Life Sciences — published in full here for anyone who wants to read, cite or share it without a download gate. A formatted PDF is also available to download, and a shorter introduction to the argument is available in our companion article, From Data Governance to Evidence Governance: Our EU Policy White Paper.
This paper proposes that European health, research and regulatory policy should distinguish more explicitly between data governance and evidence governance. The distinction is not intended to suggest that existing European frameworks neglect provenance, auditability, data quality or reproducibility. On the contrary, those concerns are increasingly visible across the European Health Data Space (EHDS), the European medicines regulatory network, the European Open Science Cloud (EOSC), the Data Act, the Data Governance Act and the Artificial Intelligence Act. The argument advanced here is narrower and more constructive: as data are reused across distributed infrastructures, transformed through increasingly complex computational workflows and converted into scientific or regulatory conclusions, Europe would benefit from treating the lineage of the resulting evidence as an object of governance in its own right.
The paper therefore introduces the concept of a verifiable evidence chain: a persistent, machine-readable and independently verifiable record that links source data, transformations, analytical processes, derived artefacts, interpretations and consequential decisions. Such a chain need not expose sensitive source data and need not depend upon a single technology. It is instead a set of governance and technical properties that can be implemented across existing infrastructures. The purpose of the proposal is to support trust, reproducibility and accountability without weakening privacy, legitimate confidentiality or the federated character of European data environments.
This is a policy paper rather than legal advice. References to European Union legislation and institutional initiatives are current to 14 September 2026 and have been checked against primary or authoritative European Union sources; all online sources were last accessed in September 2026. Legislative files still in progress, in particular the data track of the Digital Omnibus and the proposed European Biotech Act, are identified throughout as proposals rather than as law in force.
Europe is building the most ambitious federated infrastructure for health and research data anywhere in the world. It rests on a deliberate choice: rather than pooling sensitive data into central repositories, Europe is arranging for analysis to travel to the data and for results to travel back. That choice is right, and it is what makes the question raised in this paper unavoidable. Once data have been transformed into an analysis, an analysis into an evidence artefact, and an artefact into a regulatory or public-health decision, the object that matters to the decision-maker is no longer the dataset. It is the claim derived from it. Governing access to the dataset does not, by itself, keep that claim connected to the source artefacts, transformations, software versions, reference resources and analytical choices upon which it depended. The purpose of this paper is to name that second problem, to show where in the European architecture it arises, and to propose a proportionate and technology-neutral way of addressing it before the relevant technical specifications are fixed.
The timing is set by an unusually active European acquis. The EHDS creates a common framework for the primary and secondary use of electronic health data; it applies in general from 26 March 2027, which is also the deadline for its key implementing acts, and its staged implementation brings most secondary-use rules into application from 26 March 2029 and extends them to the remaining categories, including human genetic, epigenomic and genomic data, from 26 March 2031. The Data Act has applied since 12 September 2025; the Data Governance Act provides a framework intended to increase trust in data sharing, although the Commission has proposed repealing it and consolidating its operative rules into the Data Act; and the Artificial Intelligence Act, as amended by the Digital Omnibus on AI in July 2026, adds a risk-based governance structure for artificial intelligence. In December 2025 the Commission proposed a European Biotech Act which treats data and artificial intelligence as enablers of biotechnology and provides for data quality accelerators intended to deliver interoperable, well-annotated datasets. In parallel, EMA and the European medicines regulatory network are expanding the use of real-world evidence through DARWIN EU, while EOSC is developing a federated environment for reusable research data, tools and services. These developments share an important objective: enabling data to move, or to be analysed across distributed environments, without sacrificing trust. (European Commission, 2025a; European Commission, 2025c; European Commission, 2025d; European Commission, 2025e; European Commission, 2026a; EMA, 2026a; European Commission, 2026b).
Yet the governance of data access and the governance of evidence are not identical problems. A secure processing environment can determine who may access data, under what conditions, and whether raw personal data may leave the environment. A data-quality framework can determine whether a data source is sufficiently reliable and fit for a particular analysis. An interoperable common data model can make distributed analysis possible. None of those functions, taken alone, necessarily ensures that a later scientific or regulatory claim remains persistently linked to the precise source artefacts, transformations, software versions, analytical decisions and derived outputs upon which it depended. In modern life sciences, the object that ultimately matters to a regulator, policymaker or clinician is frequently not the source dataset itself but the evidence produced from it.
This paper calls that additional concern evidence governance. Evidence governance is defined here as the policies, standards and technical mechanisms through which the origin, transformation, analysis, derivation, interpretation and use of evidence can be reliably established across its lifecycle. It does not replace data governance, privacy law, cybersecurity, quality management or research integrity. Rather, it connects them at the point where data become consequential claims.
The proposal is particularly relevant to Europe because its emerging infrastructure is intentionally distributed. DARWIN EU, for example, works across multiple real-world data partners. EMA describes a process in which healthcare information is collected, accessed through participating data sources, standardised to the OMOP common data model, analysed, interpreted and then used to support regulatory decisions. EOSC is likewise being developed as a federation of data, tools, services, repositories and infrastructures. The EHDS will depend upon national health data access bodies and secure processing environments connected through HealthData@EU. The institutional architecture is therefore not a single database under a single custodian; it is a network. In such networks, auditability confined to an individual institution is valuable but incomplete. Evidence may cross an institutional boundary even where underlying personal data do not. (EMA, 2026a; European Commission, 2026b).
The central recommendation is that the European Commission, relevant agencies and Member States should consider a common minimum model for persistent evidence provenance. This should identify the evidence artefact, its parent artefacts, the relevant process or actor, material methods and software or model versions, timestamps, integrity references, applicable governance context and links to derived artefacts. The model should be machine-readable, interoperable and capable of supporting verification without requiring unnecessary disclosure of sensitive underlying data. It should also be proportionate: a routine descriptive statistic does not require the same evidential record as a genomic attribution supporting an outbreak intervention or an AI-assisted analysis supporting a regulatory conclusion.
The practical route proposed is not a new centralised European platform. It is a staged programme of standards and pilots. The Commission and relevant agencies could define a minimum evidence record, test it against genomics, real-world evidence and cross-border outbreak investigation, and assess how it interacts with EHDS secure processing environments, DARWIN EU, EOSC, existing research standards and regulatory quality systems. The objective would be to ensure that evidence lineage survives changes of system, institution and jurisdiction. The timing matters more than the headline dates for secondary use suggest. The specifications that will determine what European secure processing environments record, and what accompanies an approved output when it leaves one, are being settled now, ahead of March 2027, not in 2029.
| Stage | What evidence governance adds |
|---|---|
| Source data Data governance: privacy, lawful access, quality |
— |
| Transformation Mapping, cleaning, standardisation, versioning |
Persistent identity, lineage, integrity, material version, governance context, verification, dependency |
| Analysis Code, workflow, AI or model, reference resources |
Persistent identity, lineage, integrity, material version, governance context, verification, dependency |
| Evidence artefact Result: dataset, report, signal, genomic conclusion |
Persistent identity, lineage, integrity, material version, governance context, verification, dependency |
| Interpretation | Persistent identity, lineage, integrity, material version, governance context, verification, dependency |
| Decision | Persistent identity, lineage, integrity, material version, governance context, verification, dependency |
Data governance determines whether the source data may be used. Evidence governance is concerned with what must persist from the first transformation onwards, so that the resulting claim can still be understood and verified after it has left the environment in which it was produced.
| Recommendation | Purpose |
|---|---|
| 1. Recognise evidence governance as a distinct policy concern. | European policy should distinguish the governance of access to data from the governance of lineage, integrity and verifiability of evidence derived from those data. |
| 2. Define a European minimum evidence record. | A machine-readable core should describe material provenance relationships across source, transformation, analysis and derived evidence. |
| 3. Preserve lineage across organisational boundaries. | Evidence provenance should remain interpretable when analysis moves between data holders, secure environments, research infrastructures and regulatory bodies. |
| 4. Make versioning a first-class evidential property. | Material software, reference data, models and analytical specifications should be identifiable where they can alter the scientific meaning of a result. |
| 5. Build privacy-preserving verification. | The ability to verify lineage and integrity should not imply broad disclosure of personal, confidential or otherwise protected source data. |
| 6. Extend provenance principles to AI-assisted evidence. | Where AI materially contributes to a scientific or regulatory conclusion, its relevant role, version and validation context should form part of the evidence record. |
| 7. Pilot before mandating. | Initial pilots should focus on genomics, distributed real-world evidence and cross-border outbreak investigation, where provenance challenges are both visible and consequential. |
| 8. Keep the framework technology-neutral. | European policy should specify properties and interoperability requirements rather than prescribe a particular ledger, database or vendor architecture. |
This paper is a contribution to work already under way rather than a proposal for a new programme. Three things would advance it, and each attaches to a file that is open now.
First, that the provenance of derived outputs be addressed explicitly in the technical specifications accompanying the EHDS implementing acts due by 26 March 2027, so that an approved output leaving a secure processing environment can carry a verifiable reference to the permit, data sources and processing context under which it was produced.
Second, that the data provisions of the proposed European Biotech Act, and in particular the biotechnology data quality accelerators, be underpinned by a defined minimum evidence record. Without one, each accelerator will have to decide for itself what a verified provenance record contains, and the resulting datasets will not be comparable across projects or across Member States.
Third, a technical exchange. The BioChain would welcome the opportunity to present the minimum evidence record set out in section 3, together with the working demonstrator described in section 13, to officials and technical staff working on EHDS implementation, on the Biotech Act data provisions, or on EOSC federation interoperability, and to contribute to any pilot that tests evidential continuity across European infrastructures.
The policy context for evidence governance is unusually favourable because Europe has already undertaken much of the harder work of defining the legal and institutional conditions under which data may be accessed and reused. The GDPR remains the foundational framework for the protection of personal data. The Data Governance Act is intended to strengthen trust in data-sharing arrangements, while the Data Act establishes harmonised rules on fair access to and use of data and has applied since 12 September 2025. The EHDS now adds a sector-specific framework for electronic health data, and the Artificial Intelligence Act establishes a horizontal risk-based framework for AI. These instruments do not form a single code, and they pursue different objectives, but together they illustrate the direction of European policy: data should be usable, interoperable and economically or socially productive, yet subject to safeguards that sustain fundamental rights, security and trust. That architecture is also still in motion; two developments since late 2025 are described at the end of this section. (European Parliament and Council, 2016; European Parliament and Council, 2022; European Parliament and Council, 2023; European Parliament and Council, 2024).
The EHDS is especially important because it moves health-data governance beyond bilateral access arrangements and isolated national infrastructures. The Commission describes three principal objectives: empowering individuals in relation to their electronic health data for care; enabling the secure and trustworthy reuse of health data for research, innovation, policy-making and regulatory activities; and fostering a single market for electronic health-record systems. The implementation is deliberately staged. The Regulation entered into force on 26 March 2025 and applies in general from 26 March 2027. That date also carries the deadline for the Commission’s key implementing acts and the start of application for several Chapter IV provisions directly relevant to secondary use, including the templates for data access applications, permits and requests, the dataset description and data quality and utility label provisions, and the requirements for secure processing environments. From 26 March 2029, the first group of priority categories for primary use and the secondary-use rules for most categories begin to apply. From 26 March 2031, the second group of priority categories and the remaining secondary-use categories, including human genetic, epigenomic and genomic data, follow. A further milestone on 26 March 2035 allows third countries and international organisations to apply to participate in HealthData@EU. (European Parliament and Council, 2025; European Commission, 2025a).
This staged timetable creates a policy window, but a narrower one than the headline dates for secondary use suggest. The technical specifications that will determine what a secure processing environment records, and what accompanies an approved output when it leaves, are being settled ahead of March 2027. Decisions taken in the implementing acts and the accompanying technical work will be expensive to revisit once national health data access bodies have built to them. The EHDS architecture already contains significant safeguards for the source data: access within a secure processing environment is to be in anonymised form where the purpose of the processing can be achieved that way, and in pseudonymised form only where it cannot, and health data users may download only non-personal data from the environment. Those rules are principally concerned with protecting the source data and the people to whom they relate. The separate question considered in this paper is what persistent information should accompany an analytical result after an authorised user has transformed those data into a new evidence artefact, and after that result has left the environment in which the safeguards operated. (European Parliament and Council, 2025; European Commission, 2025b).
A similar logic appears in European research policy. EOSC is intended to provide researchers and innovators with an open and trusted multidisciplinary environment in which data, tools and services can be published, found and reused. The Commission describes EOSC as a federated system of systems and explicitly links it to reproducibility and trust in research. It also emphasises high-value, machine-actionable digital objects produced across the research lifecycle, including software and publications. This is close to the evidential perspective advocated here: the scientific object of interest is not merely a dataset but an interconnected set of digital artefacts and processes. (European Commission, 2026b).
Two things have happened since 2025 that a list of instruments does not capture. The Digital Omnibus, proposed by the Commission on 19 November 2025, split into two tracks. The artificial intelligence track was agreed in May 2026 and adopted as Regulation (EU) 2026/1744, published on 24 July 2026 and in force from 27 July 2026; it defers the application of the AI Act’s high-risk obligations to 2 December 2027 for stand-alone systems listed in Annex III and to 2 August 2028 for systems embedded as safety components in products already covered by sectoral legislation. The data track, which proposes to repeal the Data Governance Act and the Open Data Directive and consolidate their operative rules into the Data Act, remains under negotiation at the time of writing. Any proposal addressed to the European data acquis should therefore be expressed in terms of durable properties rather than of particular instruments, since the instruments are being rearranged. (European Parliament and Council, 2026; European Commission, 2025d).
The second development is the closest existing point of attachment for the argument made here. On 16 December 2025, under the Strategy for European Life Sciences published in July 2025, the Commission proposed a European Biotech Act. The proposal treats data and artificial intelligence as enablers of biotechnology innovation and provides for biotechnology data quality accelerators, described as intended to deliver high-quality, interoperable and well-annotated datasets to support the training, testing and validation of AI systems used in biotechnology, alongside trusted AI testing environments. Commentary on the proposal has characterised those datasets as provenance-verified. That characterisation is apt, but there does not appear to be a generally applicable European definition specifying what constitutes verified provenance for such datasets. What counts as verified provenance, what must be recorded for a dataset to qualify, and whether such a record remains meaningful once the dataset has been analysed and the results have moved elsewhere are precisely the questions this paper addresses. The file is in the ordinary legislative procedure and is therefore open to exactly this kind of technical input. (European Commission, 2025e; European Commission, 2025f).
Data governance and evidence governance overlap, but they begin from different questions. Data governance asks, among other things, who controls data, who may access them, for what purpose, under what legal basis, with what safeguards, and according to which quality and interoperability requirements. Evidence governance begins later in the lifecycle. It asks what exactly was produced from authorised data use, how that result came into being, which upstream artefacts materially affected it, whether those relationships have remained intact, and whether an authorised independent party can establish the lineage without simply trusting the organisation that generated the conclusion.
The distinction matters because modern scientific claims are frequently several computational steps removed from the original observation. A clinical event recorded in an electronic health record may be mapped into a common data model, filtered according to a phenotype definition, incorporated into a cohort, transformed by statistical code and aggregated with results from other databases before appearing in a regulatory report. A genomic sample may pass through extraction, sequencing, alignment, quality control, variant calling, annotation, phylogenetic analysis and epidemiological interpretation before it supports an outbreak hypothesis. An AI-assisted workflow may introduce additional layers of model versioning, retrieval sources, configuration, validation and human review. In each case, access to the original data is only the beginning of the evidential story.
Evidence governance is therefore proposed here as a connective discipline. It should not duplicate the responsibilities of data protection authorities, cybersecurity teams, quality-management systems or research-integrity processes. It should instead specify what information must persist between them if a consequential evidential claim is to remain traceable. This persistence is especially important in federated systems because evidence may be transported as reports, derived datasets, model outputs or summary statistics after the originating data have remained securely in place.
The concept also helps to clarify the relationship between audit and verification. An audit trail is often system-specific: a database records that user A changed record B at a particular time. That information is valuable, but it may cease to be visible once an artefact is exported or combined with outputs from another system. A verifiable evidence chain is concerned with continuity across such boundaries. It should permit an authorised verifier to establish that artefact C derived from artefacts A and B through process P, that the relevant identities and versions can be established, and that the recorded relationship has not been silently altered. This is a different design problem from simply retaining local logs.
It is reasonable to ask whether evidence governance is simply provenance under another name. It is not, although provenance is its largest component. A workable evidence record has to carry at least seven things: provenance, meaning where an artefact came from; integrity, meaning whether the artefact or the recorded relationship has since changed; governance context, meaning the permit, protocol or authority under which the processing occurred; materiality, meaning which transformations and versions could actually alter scientific meaning and therefore warrant recording at all; portability, meaning whether those relationships survive the move to another institution, system or jurisdiction; verifiability, meaning whether an authorised party other than the originator can inspect the claim; and dependency, meaning which downstream conclusions rest upon a given upstream artefact.
The last of these is the least familiar and, in practice, the most useful. Provenance answers the question of where a result came from. Dependency answers the question a regulator or a public-health authority actually asks when something upstream turns out to be wrong: what else is affected? A reference dataset is superseded, a phenotype definition is corrected, a coding error is found in an analytical package. Establishing which conclusions depended upon the affected artefact is a traversal problem, and it is only tractable if the relationships were recorded in a form that can be traversed.
Data governance asks whether data may be used. Evidence governance asks whether the resulting claim can be traced, tested and trusted.
A verifiable evidence chain can be understood as a sequence of linked evidential states rather than as a particular database architecture. At its simplest, the chain begins with a source observation or dataset, records acquisition and transformation, links those transformations to an analytical process, identifies the derived evidence, and then preserves the relationship between that evidence and any interpretation or decision that materially depends upon it. The chain is “verifiable” when those relationships can be checked by an appropriately authorised party using information that has not been reduced to narrative assertion alone.
The framework should be deliberately technology-neutral. Cryptographic hashes may be useful for proving that a digital artefact has not changed. Digital signatures may be useful for establishing authorship or organisational attestation. Persistent identifiers may be useful for locating or referring to objects. Workflow standards may express computational steps. Existing laboratory or clinical quality systems may provide authoritative event records. Distributed ledgers may be appropriate in some circumstances, but they should not be treated as synonymous with evidence governance. The policy objective is to define the properties that must survive, not to privilege one mechanism by which they are implemented.
Those properties can be expressed through a small set of questions. What is the artefact? From which prior artefacts did it derive? Who or what process created it? When was it created? Under whose organisational responsibility? Which method, workflow, software, reference dataset or model materially affected the result? What integrity evidence exists? Which governance context authorised the processing? Which later artefacts or decisions rely upon it? A well-designed evidence record need not answer every question with equal detail in every context. Proportionality is essential. But it should provide a common grammar through which evidential dependencies can be represented and inspected.
The problem identified in this paper is not the absence of provenance standards. Provenance has been the subject of sustained technical work for two decades, and much of that work is mature. The W3C PROV family provides a general model for describing entities, activities and agents and the relationships between them. Research-object approaches, including RO-Crate, package data, software and context together with machine-readable metadata. Workflow description standards express computational steps in portable form. Persistent identifier infrastructures, including DOIs, ORCID and the identifier services operated by European research infrastructures, make objects and actors referenceable over time. The FAIR principles and the associated FAIR Digital Object work address findability, accessibility, interoperability and reuse at the level of the digital object. EOSC maintains its own body of work on metadata and provenance within the federation, and life-science infrastructures maintain domain-specific conventions for sample, sequence and analysis identifiers.
What is missing is not syntax. It is a common policy expectation that the relevant provenance survives the transition from data to consequential evidence, and survives it across organisational and jurisdictional boundaries. A standard can describe a lineage; it cannot by itself require that the lineage be created, that it be retained when an artefact is exported from a secure processing environment, that it remain interpretable to the receiving institution, or that it be verifiable by someone other than its author. Those are governance questions, and they are the questions this paper addresses. The recommendation that follows is therefore that European policy should specify the properties which must persist and map them onto the standards that already exist, not that it should commission a new vocabulary.
| Field | Purpose |
|---|---|
| Persistent identifier | Uniquely identifies the evidence artefact or event. |
| Parent artefact(s) | Identifies the immediate inputs from which the artefact was derived. |
| Actor or process | Identifies the responsible person, organisation, instrument, workflow or service. |
| Time | Records a sufficiently precise creation or transformation time. |
| Method or workflow | Links the artefact to the procedure by which it was produced. |
| Material version information | Records relevant software, model, reference-data or protocol versions where these may affect meaning. |
| Integrity reference | Provides a means to detect unauthorised or unrecorded alteration. |
| Governance context | Links processing to the relevant permit, study, protocol, authority or policy context where appropriate. |
| Derived artefacts | Supports traversal forward to results, reports, submissions or decisions. |
The proposed distinction is easiest to see by examining data quality. EMA, the Heads of Medicines Agencies and the Joint Action Towards the European Health Data Space co-produced a Data Quality Framework for EU medicines regulation, and EMA made available its real-world data chapter, adopted by the Committee for Medicinal Products for Human Use, in March 2026. The framework is important precisely because regulatory conclusions are sensitive to the quality and fitness for purpose of their underlying data. Evidence governance does not compete with that work. It asks what must persist so that a later reviewer can establish which quality-assessed data source, which extraction, which transformation and which analysis actually supported a particular conclusion. (EMA, 2026b).
Quality is a property assessed in relation to a purpose; provenance is the history necessary to understand how an artefact reached its present form. A high-quality source may still yield unreliable evidence if it is transformed incorrectly, analysed with the wrong specification or combined with a different version of a terminology or reference set. Conversely, a provenance record cannot rescue fundamentally poor data. The two functions are complementary. A mature evidence-governance regime would allow a quality assessment to remain attached to the exact data or derived artefact to which it applied, rather than becoming an unstructured statement separated from the object being evaluated.
The same logic applies to reproducibility. Reproducibility is strengthened when the data, code, environment, versions and decisions necessary to repeat or interrogate an analysis are identifiable. EOSC explicitly connects its federated research infrastructure with reproducibility and trust. Evidence governance provides a way of carrying that aspiration into regulated and privacy-sensitive environments where source data cannot simply be published alongside every analysis. It therefore sits between open-science ideals and the legitimate constraints of health-data governance. (European Commission, 2026b).
DARWIN EU provides an unusually clear case through which to examine the argument because EMA publicly describes both its purpose and its analytical sequence. Established in 2022 by EMA and the European medicines regulatory network, DARWIN EU works with real-world healthcare databases and supports regulatory decision-making on the use, safety and benefits of medicines. EMA states that data partners standardise data to the Observational Medical Outcomes Partnership common data model, that experts analyse the data and create study reports, and that EMA committees and EU regulators interpret those reports as part of regulatory decision-making. A document dated 16 June 2026 lists forty data partners as onboarded. EMA’s fourth report on regulator-led studies, covering February 2025 to February 2026, records 108 research topics assessed and 88 studies completed or ongoing. (EMA, 2026a; EMA, 2026c).
This is precisely the kind of distributed infrastructure in which evidence governance becomes important. The source healthcare data do not need to be assembled into a single European super-database for regulatory value to be created. Instead, standardisation and distributed analysis allow evidence to be generated from multiple environments. The regulatory object that moves is often a study result or report rather than the underlying identifiable record. Accordingly, the provenance problem changes form: the important chain may link local source characteristics, data-model transformations, cohort definitions, analytical code, study protocol versions, partner-specific results, aggregation procedures and the final regulatory interpretation.
DARWIN EU documents these matters, and does so within a sophisticated scientific and regulatory environment. The question raised here is not about the completeness of that documentation but about its portability. Lineage held across study protocols, partner systems, analytical repositories and committee papers is intelligible to the people and systems that created it; it is not currently expressed in a form that travels intact to a different infrastructure, a different agency or a different decade. The question for European policy is whether a persistent, interoperable minimum should be defined so that it does. If a result is revisited five years later, perhaps because a phenotype definition has changed, a coding error has been discovered, or a reference dataset has been superseded, the system should make it possible to identify which conclusions depended upon the affected artefact. That is a graph problem as much as it is a document-retention problem.
The same approach could strengthen regulatory learning. If a study output is linked to its analytical lineage, subsequent assessments can distinguish between a change in the underlying population, a change in source data, a change in method and a change in interpretation. Such distinctions are central to good regulatory science, yet they are difficult to reconstruct if the evidential history is scattered across local logs, PDF reports, code repositories and institutional records.
Genomics is likely to be one of the most demanding tests of European evidence governance. The EHDS implementation timetable is particularly significant because human genetic, epigenomic and genomic data sit among the remaining secondary-use categories for which the rules begin to apply from 26 March 2031. That is often read as a distant horizon. It is better understood as a short one: the implementing acts and technical specifications that will define how secure processing environments describe datasets and outputs are due by March 2027, and genomic data will later arrive into a framework those decisions have already shaped. The interval is an opportunity to embed provenance expectations before cross-border genomic secondary use reaches maturity, not a reason to defer the question. (European Parliament and Council, 2025; European Commission, 2025a).
A genomic conclusion is rarely a direct reading of a biological specimen. Between sample and statement lie multiple transformations: sample collection and accession, extraction, library preparation, sequencing, base calling, quality control, alignment or assembly, variant calling, filtering, annotation, comparison with reference resources, and often statistical or phylogenetic interpretation. Each stage can introduce assumptions, thresholds and versions that alter the result. A statement such as “this variant was present” or “these cases belong to the same genomic cluster” therefore depends upon a structured history.
The public-health implications are substantial. ECDC has long promoted the use of whole-genome sequencing in surveillance and cross-border outbreak investigation, and in 2026 highlighted the role of genomic data in tracking antimicrobial resistance across Europe. As sequencing becomes routine, the evidential question will no longer be whether Europe possesses genomic capacity, but whether the provenance of genomic conclusions remains sufficiently clear when data and interpretations cross laboratories, national surveillance systems, research infrastructures and European agencies. (ECDC, 2019; ECDC, 2026).
A European minimum evidence record could be especially valuable here because genomics is naturally version-sensitive. Reference genomes change; annotation databases are updated; taxonomies and lineage definitions evolve; analytical software is revised; thresholds for quality or cluster inclusion may be modified. A result may be technically correct relative to the resources available on the day it was generated yet appear inconsistent with a later reanalysis. Persistent provenance converts that apparent contradiction into an intelligible history. It allows a reviewer to say not merely that two reports differ, but why.
Outbreak investigation demonstrates why evidence chains must extend beyond a single domain. A foodborne or zoonotic event may involve farm records, veterinary observations, environmental sampling, laboratory diagnostics, human clinical data, genomic sequencing, food-supply information and epidemiological interviews. No single institution necessarily controls the complete evidential picture. The public-health conclusion is constructed by linking partial observations held under different legal, professional and organisational regimes.
This is the practical meaning of One Health evidence governance. The aim is not to collapse human, animal and environmental datasets into one unrestricted repository. It is to preserve the evidential relationships that justify a conclusion while respecting the distinct controls that apply to each source. A regulator or outbreak team may need to know that a human isolate and a food-chain isolate were linked by a particular genomic analysis without being entitled to broad access to every associated clinical or commercial record. Evidence governance should support that separation.
Such a model would be valuable in cross-border events. A sample collected in one Member State may be sequenced in another, compared against a European reference dataset and interpreted alongside epidemiological evidence from several jurisdictions. If the relevant lineage is represented only in institutional documents, reconstruction can become slow precisely when time matters. A machine-readable chain would not automate scientific judgement, but it could reduce the friction involved in establishing which artefacts underlie a hypothesis and whether those artefacts are still current.
Artificial intelligence makes the distinction between data and evidence even more important because the analytical process itself may become less transparent to the end user. The EU Artificial Intelligence Act is a risk-based legal framework intended to support trustworthy AI, and the Commission notes that high-risk AI systems, including certain AI-based medical software, are subject to requirements relating to matters such as data quality, information, oversight, robustness and risk mitigation. In healthcare, AI is also governed alongside sectoral frameworks such as medical-device law. The timetable for those high-risk obligations has since moved. Regulation (EU) 2026/1744, the Digital Omnibus on AI, defers their application to 2 December 2027 for stand-alone systems listed in Annex III and to 2 August 2028 for AI embedded as a safety component in products already covered by sectoral legislation, which is the route most clinical AI will follow. The deferral does not reduce the underlying requirements. It extends the period during which the supporting expectations, including those concerning data, documentation and traceability, are still being settled, and therefore the period in which provenance conventions can still be shaped. (European Commission, 2026a; European Commission, 2024; European Parliament and Council, 2026).
Evidence governance should not attempt to reproduce the AI Act. It should address a different question: when an AI system materially contributes to a scientific or regulatory conclusion, what should be preserved about that contribution so the resulting evidence remains intelligible? Depending upon the application, this might include the model identity and version, validated intended use, relevant configuration, retrieval or reference sources, input artefacts, output artefacts, material post-processing and human review. In some settings, exact prompts or full model internals may be neither necessary nor legally or commercially available. A proportionate evidential standard should therefore focus on what is needed to understand and verify the role played by AI in producing the consequential artefact.
This distinction will become increasingly important as AI is used not merely to automate clerical tasks but to classify images, prioritise variants, identify safety signals, construct cohorts, generate code, summarise evidence and support scientific interpretation. If an AI-assisted conclusion enters a regulatory submission or public-health decision, an evidential chain should reveal that AI played a material role rather than allowing the output to appear as though it arose from an unspecified human analytical process. Transparency at this level is compatible with innovation: it records dependency without requiring that every model be made public.
The European Union has a structural reason to care about cross-border evidential continuity: its scientific and regulatory strength derives partly from networks. HealthData@EU, DARWIN EU, EOSC, European Reference Networks, multinational research consortia and surveillance systems all depend upon cooperation between institutions that retain their own responsibilities. The challenge is therefore not simply to make each institution auditable. It is to preserve meaning when an artefact leaves the institution in which it was created.
Consider a stylised example. A Finnish cohort is defined within an authorised environment; a Belgian group develops an analytical workflow; a German infrastructure executes a standardised analysis against local data; a French consortium combines summary results; and an EU-level body uses the combined evidence. Every institution may have exemplary internal logging, yet the final report may still lack a single navigable representation of the dependencies that connect its conclusions to the upstream artefacts. In that circumstance there is no failure of local governance. There is a gap between local governance systems.
A European evidence-governance framework should therefore be designed around portable assertions and interoperable identifiers. The principle is that provenance should follow the evidence. When an artefact crosses an organisational boundary, the receiving environment should be able to preserve the relevant upstream relationships, add its own transformations and pass the extended chain onwards. This resembles the way digital certificates or persistent scholarly identifiers remain meaningful outside the system that first issued them. The exact mechanism requires technical standardisation, but the policy objective can be stated now.
The EHDS is the natural policy setting in which to test the distinction because secondary use converts protected health data into research, policy and regulatory outputs at European scale. The EHDS is expressly intended to enable secure and trustworthy reuse for research, innovation, policymaking and regulatory activities. Health data access bodies, data permits, dataset catalogues, secure processing environments and HealthData@EU together establish the conditions under which authorised secondary use can occur. (European Commission, 2025a; EUR-Lex, 2025).
The policy opportunity lies after authorised processing. A researcher may leave a secure environment with an approved aggregate output, statistical result or other non-disclosive artefact rather than with the underlying personal data. That is correct from a privacy perspective, but the exported object may then enter a publication, regulatory submission, model-development workflow or further research project. The provenance requirements that applied inside the secure environment need not automatically remain attached to that new object. Evidence governance would create a bridge by ensuring that an approved output can carry a reference to its permitted analytical lineage without exposing the underlying records.
This distinction could also improve accountability without centralisation. A health data access body need not become responsible for validating every downstream scientific interpretation. Its role could remain focused on lawful access and compliance. The evidence chain would simply allow later parties to establish that an output came from a particular authorised process and to traverse, subject to permission, the relevant provenance relationships. Responsibility would remain distributed, but the history would not be fragmented.
A useful European standard should be minimal enough to travel between domains yet expressive enough to support meaningful verification. It should not attempt to encode every laboratory observation, statistical parameter or regulatory judgement in a single universal schema. Instead, it should define a small core of mandatory provenance relationships and permit domain-specific extensions. This mirrors successful approaches elsewhere in digital standards: interoperability is achieved by agreeing a stable core and allowing richer specialist profiles above it.
The minimum evidence record proposed in this paper contains nine conceptual elements: a persistent identifier; parent artefacts; responsible actor or process; time; method or workflow; material version information; integrity reference; governance context; and derived artefacts. Some records would also require status, quality assertions, signatures, access restrictions or links to validation evidence. The critical point is that the core relationship between input, transformation and output is explicit and machine-readable.
The standard should also distinguish evidence from metadata about evidence. A hash may prove that a file has not changed, but it does not explain the scientific relationship between two files. A timestamp may prove that an event was recorded at a particular time, but not whether the correct reference genome was used. A workflow description may show intended processing but not prove which inputs actually entered a particular run. Evidence governance requires these elements to work together rather than treating any single technical control as sufficient.
Material evidence artefacts should be identifiable over time. Persistence does not require that the underlying object remain publicly accessible; indeed, in health research it often cannot. It requires that references remain resolvable within the authorised governance environment and that the identifier not be silently reassigned to a different object.
Provenance should be captured as part of routine processing rather than reconstructed retrospectively for audit. Retrospective reconstruction is expensive, error-prone and particularly difficult when staff, software and organisational structures have changed. Systems that already know which input produced which output should preserve that relationship at the moment of creation.
Where a version can materially alter scientific meaning, it should be treated as evidence. This includes not only analytical software but also reference datasets, terminology releases, ontologies, genomic references, model versions and protocols. Versioning should be selective and materiality-based; indiscriminate capture can make provenance unusably noisy.
Verification should not be conflated with disclosure. A reviewer may need to establish that an artefact derives from a permitted source and validated process without being entitled to view individual-level health data. The architecture should support different layers of visibility, with provenance assertions capable of being verified independently from the protected content to which they refer.
A provenance standard that works only within one vendor product would fail the central European use case. Records must survive export, federation and institutional change. Alignment with the existing research and semantic-web provenance standards described in section 3 — W3C PROV, research-object and RO-Crate approaches, workflow description standards, persistent identifier infrastructures and FAIR Digital Object work — should therefore be the starting point, and the creation of new syntax should require justification rather than be assumed. European policy should seek semantic compatibility and transportability rather than platform uniformity.
The entity that generated an evidential assertion should not be the only entity capable of verifying it. Independent verification need not mean public verification; it may occur only between authorised organisations. The principle is that trust should be supported by inspectable evidence rather than by the reputation of the originating system alone.
The BioChain is developing a demonstrator using ancient DNA and other public scientific datasets to test whether persistent evidence provenance can be represented end to end. The choice of material is deliberate, and its limits are acknowledged at the outset: ancient DNA is neither clinical data nor regulated data. What it is, is public, versioned, identifier-bearing, computationally transformed through long pipelines, and subject to reinterpretation as reference resources improve. The provenance problems are therefore structurally the same as those in the three pilot streams proposed below, and they can be worked on in the open rather than behind a data permit. A demonstrator built on protected health data could not be shown to anyone, which is the wrong property for a proposal that asks to be tested. Ancient DNA therefore functions here as a safe experimental substrate for evidence-infrastructure design, not as a proxy for the governance requirements of clinical data.
The demonstrator is not presented as evidence that one technical architecture should become a European standard. Its role is narrower: to test whether source, ingestion, transformation, analysis and verification can be represented as a navigable provenance chain without replacing the systems that hold the original scientific records. A planned extension to ancient pathogen data provides a bridge to modern genomic surveillance, because it introduces organism identification, sequence analysis and epidemiological interpretation while retaining the advantages of non-clinical public datasets.
This approach reflects a broader principle of responsible infrastructure design. Governance concepts should be testable. A policy proposal for evidence lineage is more credible if it can be demonstrated on real, messy, versioned scientific material rather than only described in diagrams. At the same time, demonstrations should remain subordinate to policy goals: the purpose is to learn what information must persist, where interoperability fails and what verification actually requires.
The appropriate next step is experimentation rather than immediate regulation. A European Verifiable Evidence Pilot could be conducted across three streams chosen because they expose different aspects of the problem: genomic analysis, distributed real-world evidence and cross-border outbreak investigation. Each pilot should use existing infrastructures and workflows, adding only the minimum provenance layer necessary to test whether evidential continuity can be improved.
The genomic stream should test persistent lineage from source specimen or public dataset through sequence data, analytical transformation, reference resources and derived interpretation. Particular attention should be paid to versioning and to the ability to identify which downstream conclusions are affected when an upstream reference or analytical step changes.
The real-world-evidence stream should model a DARWIN-like distributed study in which local data are standardised, analysed and returned as results rather than pooled as identifiable records. The pilot should test whether a common evidence record can represent cohort specifications, common data-model versions, analytical packages, local results, aggregation and final interpretation without revealing protected patient-level data.
The outbreak stream should connect laboratory, genomic and epidemiological evidence across at least two Member States and one European-level analytical or coordinating body. The purpose would be to test whether a shared provenance model can accelerate reconstruction of an evidence chain during an evolving event while preserving domain-specific access controls.
Evaluation should examine scientific usefulness, implementation burden, interoperability, privacy, security, legal fit, long-term maintainability and human factors. A pilot that creates perfect lineage but imposes excessive manual documentation would not be successful. The objective is to capture provenance from systems that already generate it and to make those relationships portable.
The Commission and relevant European agencies should use the term evidence governance, or an equivalent concept, to distinguish the governance of derived scientific and regulatory artefacts from the governance of access to source data. Formal recognition would create a shared vocabulary across health data, regulatory science, research infrastructure and AI policy. It would also reduce the risk that provenance is treated as a purely technical afterthought.
A small cross-institutional working group should map existing standards and define the minimum relationships necessary for European use cases. The objective should be compatibility with existing provenance, research-object and workflow standards rather than invention for its own sake. Domain extensions for genomics, real-world evidence and public health could then be developed above the core.
As EHDS implementing acts and technical specifications mature ahead of March 2027, the Commission and Member States should consider how authorised analytical outputs retain links to the permits, data sources and processing contexts from which they derive. This need not expand the remit of health data access bodies; it can instead ensure that outputs remain anchored to the governance conditions under which they were produced. The same requirement should be made explicit in the data provisions of the proposed European Biotech Act, where the biotechnology data quality accelerators will otherwise each have to define for themselves what a verified dataset provenance record contains. Settling it once, consistently across the two files, would avoid divergent sectoral and national interpretations that are far harder to reconcile later. (European Commission, 2025a; European Commission, 2025e).
European guidance should state that versions of software, models, reference datasets and analytical specifications form part of the evidential record where they can materially alter a result. This is especially important in genomics, AI-assisted analysis and rapidly evolving classification systems.
Evidence provenance should be capable of being shared at a different permission level from source data. Technical work should explore how integrity, lineage and governance assertions can be verified without revealing protected health information, trade secrets or other confidential content.
Where an AI system materially contributes to evidence used in a regulated or public-interest decision, the evidence record should identify that role at a level sufficient for scientific and regulatory review. This recommendation should complement, not duplicate, obligations under the AI Act and sectoral legislation.
EHDS, EOSC, regulatory networks and public-health infrastructures should be encouraged to exchange provenance in interoperable form. The strategic objective is that evidence history should not disappear when an artefact moves from one infrastructure to another.
A mandatory standard imposed too early could freeze immature approaches or create administrative burden without scientific value. Europe should therefore pilot, measure and refine. Mandatory requirements should follow only where pilots demonstrate a clear benefit and where the proportionality of record-keeping can be defended.
It does not require a central European database containing every research artefact. The European landscape is federated and should remain capable of accommodating national, institutional and domain-specific infrastructures. Evidence governance can operate through linked assertions and persistent identifiers without centralising protected content.
It does not require publication of sensitive data. The framework should explicitly permit a distinction between verification of provenance and access to content. In many health and commercial contexts, the verifier may be able to establish the integrity and authorised lineage of an artefact without receiving the underlying individual-level data.
It does not require blockchain. Distributed ledger technologies may be useful where multiple organisations require a shared tamper-evident record without a single trusted operator, but they are one implementation option among several. A policy standard should specify integrity, portability and verification properties, not a fashionable architecture.
It does not automate scientific judgement. A perfect provenance chain can document a poor scientific decision with great precision. Evidence governance improves the capacity to inspect how a conclusion was reached; it does not determine whether the interpretation was scientifically correct. Expert judgement, validation and peer review remain indispensable.
Evidence governance also has a strategic dimension. Europe increasingly treats trusted digital infrastructure, research-data access and technological sovereignty as related concerns. EOSC policy now explicitly discusses permanent accessibility, verifiability, security, resilience and interoperability of European research data and services. These concerns are not isolationist; they are conditions for dependable collaboration. If Europe is to rely on distributed digital evidence, it needs confidence that the provenance of that evidence remains intelligible even when infrastructures, service providers or geopolitical conditions change. (European Commission, 2026b).
A persistent evidence layer can contribute to institutional memory. Scientific infrastructures routinely outlive individual grants, software versions and teams. Regulatory decisions may be revisited years after the original analysis. Public-health conclusions may be re-examined after new evidence emerges. In such circumstances, evidence lineage becomes part of the durability of the European knowledge base. It helps distinguish between loss of access to an old system and loss of the scientific history itself.
There is also an economic argument, and it is stronger than it first appears. Poorly portable provenance creates a recurring reconstruction cost. Each time an evidence artefact crosses an organisational boundary, somebody has to re-establish where it came from, which versions of software and reference resources were used, under whose authority it was produced and what has happened to it since. That work is done once for the regulatory submission, again for the audit, again for the partner collaboration, again for the due-diligence exercise, and again when a question is reopened years later. It is largely manual, it depends on individuals who may have moved on, and it produces an answer of variable quality each time.
The cost is therefore not a one-off compliance overhead. It is a charge on every subsequent reuse of evidence, which is precisely the activity that European policy is trying to increase. A reusable evidence record converts provenance from a repeated manual exercise into infrastructure. The returns are in reduced duplication, faster assessment and re-assessment, lower friction in cross-border and cross-sector collaboration, greater confidence in externally generated evidence, and better resilience when a system, a supplier or an institution changes. These are competitiveness and administrative-burden questions as much as scientific ones, and they bear directly on the simplification agenda that motivated the Digital Omnibus. The benefits would not arise from recording more information indiscriminately, but from recording the right relationships once and allowing them to be reused across legitimate governance contexts.
Europe has spent the past decade constructing increasingly sophisticated legal and technical foundations for trustworthy data use. The EHDS, Data Governance Act, Data Act, AI Act, DARWIN EU and EOSC demonstrate a common direction, and the proposed European Biotech Act extends it into biotechnology: data should be usable across organisational boundaries while remaining secure, governed and worthy of trust. The next challenge is to ensure that the evidence produced from those data retains an equally coherent history.
Evidence governance offers a way to describe that challenge without displacing existing frameworks. Its concern is the continuity of scientific meaning. It asks whether an authorised reviewer can determine which source artefacts, transformations, methods, versions and interpretations produced a consequential claim; whether those relationships survive institutional and technical boundaries; and whether integrity can be verified without unnecessary disclosure of protected content.
The case for action is strongest where Europe is already building federated infrastructures. In DARWIN EU, distributed healthcare data become regulatory evidence. In the EHDS, protected health data will increasingly support cross-border secondary use. In EOSC, research data, tools and services are being federated to improve reuse and reproducibility. In genomics and outbreak investigation, conclusions increasingly depend upon chains of computational transformation. In AI-assisted science, the analytical agent itself becomes another provenance-bearing component. None of these developments requires a new central database. They require continuity.
The practical proposal is therefore modest: define a minimum evidence record, make it interoperable, preserve material version information, separate verification from disclosure, and test the approach in carefully chosen European pilots. If those pilots demonstrate value, evidence-governance principles can then be incorporated progressively into technical standards, funding conditions, regulatory guidance and data-space implementation. The opportunity to do so cheaply is time-limited. EHDS implementing acts are due by March 2027 and the Biotech Act is in negotiation now; provenance conventions adopted at that stage cost far less than provenance reconstructed afterwards.
Data governance asks whether data may be used. Evidence governance asks whether the claim produced from those data can still be understood and verified after the data have passed through complex systems, organisations and jurisdictions. Europe increasingly needs both.
The policies, standards and technical mechanisms through which the origin, transformation, analysis, derivation, interpretation and use of evidence can be reliably established across its lifecycle — distinct from data governance, which asks only whether the underlying data may be used.
Because the EHDS implementing acts and the technical specifications that will determine what a secure processing environment records, and what travels with an approved output when it leaves one, are due by 26 March 2027. Those decisions will be expensive to revisit once national health data access bodies have built to them.
No. The proposal is explicitly designed around the federated, distributed character of European data infrastructure — DARWIN EU, EOSC and the EHDS all work by moving analysis to the data, not the reverse. Evidence governance is meant to work through linked, portable assertions, not centralisation.
No. The paper treats verification and disclosure as separable. A hash can prove a file is unchanged without revealing it; an identifier can refer to a protected object without publishing it; a permit reference can show an authorised process existed without granting access to a single patient record.
A proposed small, stable core of nine conceptual elements — persistent identifier, parent artefacts, responsible actor or process, time, method or workflow, material version information, integrity reference, governance context, and derived artefacts — that should travel with a consequential scientific or regulatory result.
We are building a demonstrator using ancient DNA and other public scientific datasets to test whether persistent evidence provenance can be represented end to end, on real and messy scientific material, without the governance complications of clinical data. It is a test of what breaks in practice, not a proposal that our own architecture become a standard.
The Minimum Evidence Record, v0.1 — the nine-field technical specification this paper’s central recommendation is built on — is published separately as The Minimum Evidence Record, v0.1. For the argument applied to European cloud and AI sovereignty policy, see Europe Knows Where Its Servers Are. It Doesn’t Know Where Its Evidence Came From.
All online sources were last accessed in September 2026. Where a reference is to a legislative proposal rather than to an adopted instrument, this is stated.
The following propositions were checked against authoritative EU sources during preparation of this revision. The EHDS is Regulation (EU) 2025/327 of 11 February 2025, published in the Official Journal on 5 March 2025 and in force from 26 March 2025; it applies in general from 26 March 2027, with most secondary-use rules from 26 March 2029 and the remaining categories, including human genetic, epigenomic and genomic data, from 26 March 2031. The Data Act has applied since 12 September 2025. The Commission’s Digital Omnibus proposal of 19 November 2025 would repeal the Data Governance Act and consolidate its operative rules into the Data Act, and remains under negotiation. The Digital Omnibus on AI, Regulation (EU) 2026/1744 of 8 July 2026, was published on 24 July 2026 and entered into force on 27 July 2026, deferring the AI Act’s high-risk obligations to 2 December 2027 and 2 August 2028 respectively. The European Biotech Act was proposed on 16 December 2025 and is in the ordinary legislative procedure. DARWIN EU was established in 2022, and an EMA document dated 16 June 2026 lists forty onboarded data partners. EMA made the real-world data chapter of its Data Quality Framework available in March 2026. ECDC published the results of its multi-country CCRE genomic survey on 13 May 2026. EOSC is described by the European Commission as a federated environment intended to support reuse, reproducibility and trust. Where this paper proposes “evidence governance”, a “minimum evidence record” or a “verifiable evidence chain”, these are policy proposals of The BioChain rather than descriptions of existing EU legal duties.
If your organisation generates, manages, analyses or governs biological evidence and would be interested in participating in a UK or European provenance demonstrator, The BioChain would welcome the conversation.
Get in touch
Regulatory developments, technical notes and platform news — sent occasionally, straight to your inbox.