Ten frameworks shaping data, AI and evidence governance across the European Union, from the EHDS and the AI Act to the Data Union Strategy’s emerging provenance-standardisation work.
The European Health Data Space Regulation establishes a common European framework for the primary and secondary use of electronic health data.
The Regulation entered into force on 26 March 2025 and applies in general from 26 March 2027. That date also carries the deadline for the Commission’s key implementing acts and the start of application for several Chapter IV provisions directly relevant to secondary use, including the templates for data access applications, permits and requests, the dataset description and data quality and utility label provisions, and the requirements for secure processing environments.
Most secondary-use provisions begin applying from 26 March 2029. The remaining categories, including human genetic, epigenomic and genomic data, follow from 26 March 2031. A further milestone on 26 March 2035 allows third countries and international organisations to apply to participate in HealthData@EU.
The commonly quoted 2029 and 2031 dates describe when secondary use arrives. 2027 is when the technical decisions that shape it are made.
The EHDS creates an important distinction between access to source data and the evidence subsequently derived from those data.
Within a secure processing environment, access is to be in anonymised form where the purpose of the processing can be achieved that way and in pseudonymised form only where it cannot, and health data users may download only non-personal data. Those safeguards protect the source data effectively.
The open question concerns what happens next. As authorised users generate statistical outputs, research findings, models and regulatory evidence, what information about source, permit, processing context, versions and transformations should remain attached to those outputs after they leave the environment?
For The BioChain, this is one of the clearest examples of the difference between data governance and evidence governance — and the implementing work now under way is where it will be settled.
The Digital Omnibus is the Commission’s simplification package for the digital acquis. It has proceeded on two tracks.
The artificial intelligence track was adopted as Regulation (EU) 2026/1744 of 8 July 2026, published on 24 July 2026 and in force from 27 July 2026. It defers several AI Act application dates (see the AI Act entry below).
The data track remains under negotiation. As proposed, it would repeal the Data Governance Act and the Open Data Directive and consolidate their operative rules into the Data Act, alongside targeted amendments to the GDPR, the ePrivacy Directive, NIS2 and DORA.
Because the final text is not settled, organisations should treat the current instrument-by-instrument map of European data law as provisional.
The practical lesson of the Omnibus is that instruments are being rearranged faster than compliance architectures can be rebuilt around them.
That argues for designing evidence and provenance arrangements around durable properties — persistent identity, lineage, material version information, integrity and controlled verification — rather than around the structure of any particular regulation.
Arrangements built to the shape of a specific instrument tend to require rework each time the instrument moves. Arrangements built to the underlying question generally do not.
The AI Act establishes the European Union’s risk-based regulatory framework for artificial intelligence. Requirements vary according to the nature and risk of the system. High-risk systems are subject to obligations including risk management, appropriate data quality, documentation, logging, human oversight, robustness and traceability.
The application timetable was amended in July 2026 by the Digital Omnibus on AI. High-risk obligations now apply from 2 December 2027 for stand-alone systems listed in Annex III, and from 2 August 2028 for AI embedded as a safety component in products already covered by sectoral legislation. Most clinical and medical-device AI follows the second route.
Transparency obligations under Article 50 began applying on 2 August 2026. These include disclosure of interactions with AI and machine-readable marking of certain AI-generated or manipulated content. Application of the Article 50(2) marking requirement to systems already placed on the market was deferred to 2 December 2026.
The deferrals do not reduce the underlying requirements. They extend the period in which supporting expectations, including those concerning data and documentation, are still being settled.
The AI Act makes traceability and documentation increasingly important properties of AI systems.
For evidence governance, the question goes further: when an AI system materially contributes to a scientific, clinical or regulatory result, can the organisation later identify which model, version, input, reference source and human review process contributed to that result?
That is particularly important where an AI-derived output subsequently enters a conventional scientific report or regulatory submission and may otherwise lose its machine provenance at the point of transfer.
The Data Act establishes harmonised rules concerning access to and use of data, with particular focus on data generated by connected products and related services. It forms a central part of the European strategy for a more usable and competitive data economy.
Under the Digital Omnibus proposal, the Data Act would additionally absorb the operative provisions of the Data Governance Act and the Open Data Directive. That proposal is not yet adopted.
Greater availability and portability of data increase the importance of preserving context.
An organisation receiving data from another system may know that it has lawful access to those data while still lacking information about how they were generated, changed or previously interpreted.
Interoperability therefore needs to include not simply the ability to move data, but the ability to retain enough provenance to understand what those data represent.
The Data Governance Act established mechanisms intended to increase trust in the reuse and sharing of data, including arrangements around data intermediation services and data altruism.
The Digital Omnibus proposal of 19 November 2025 would repeal the Data Governance Act and move its operative rules into the Data Act. Until that proposal is adopted, the DGA continues to apply in its own right.
Organisations with compliance arrangements referencing the DGA specifically should monitor the negotiation, as citations and obligations may need to be re-pointed.
Trusted data sharing depends not only upon access controls but also upon confidence in the data received.
As data are exchanged through intermediary and federated arrangements, persistent identifiers and portable provenance help receiving organisations understand the history and authority of the material they are using.
That requirement is unaffected by which instrument houses the rules.
The Data Union Strategy is intended to increase access to high-quality data for AI, simplify the European data-rule landscape and strengthen Europe’s position in international data flows. The Commission explicitly links it with data sovereignty and with the ability of European businesses, particularly SMEs, to participate in the AI economy.
The Strategy commits the Commission to launching a standardisation request for a European data quality standard covering completeness, consistency, provenance, semantic clarity and governance. A related initiative addresses the standardisation of annotation and labelling practices, with the stated aim of ensuring trust in data’s origin and conditions of use.
This is, to our knowledge, the most direct commitment to provenance standardisation currently on the European agenda.
High-quality data are not defined solely by accuracy. For many scientific and regulated uses, quality depends upon knowing where data came from, which transformations occurred, what standards or versions were used, and whether that history remains reliable.
The inclusion of provenance within the proposed data quality standard makes this an active standardisation question rather than a theoretical one.
Our view is that the same standardisation work should address what persists when data become evidence — not only the quality of a dataset before analysis, but the traceability of the result afterwards. We have published a proposed minimum evidence record as a contribution to that discussion.
The proposed European Biotech Act follows the Commission’s Strategy for European Life Sciences of July 2025. It aims to strengthen the Union’s biotechnology and biomanufacturing sectors, and treats data and artificial intelligence as enablers of biotechnology innovation.
Among its provisions are biotechnology data quality accelerators, intended to deliver high-quality, interoperable and well-annotated datasets to support the training, testing and validation of AI systems used in biotechnology, alongside trusted AI testing environments.
The Biotech Act comprises two instruments adopted by the Commission on the same day: the proposed Regulation, COM(2025) 1022, and an accompanying proposed Directive, COM(2025) 1031. The proposal is at an early stage and its final content may change substantially.
Commentary on the proposal has described the datasets these accelerators would produce as provenance-verified. There does not appear to be a generally applicable European definition specifying what constitutes verified provenance for such a dataset.
What must be recorded for a dataset to qualify, and whether that record remains meaningful once the dataset has been analysed and the results have moved elsewhere, are open questions.
For organisations in life sciences, this is the file most likely to determine what dataset provenance is expected to mean in practice.
The Commission presented its European Technological Sovereignty Package on 3 June 2026. It comprises the Cloud and AI Development Act, a revision of the Chips Act, an EU Open Source Strategy and a strategic roadmap for digitalisation and AI in the energy sector.
CADA proposes to expand EU cloud and AI computing capacity and introduces a framework of assurance levels for cloud services used in sensitive public-sector contexts. As described in the proposal, these range from processing and storage within the Union at the lowest level, through independence from third countries and transparency over the software supply chain, to EU ownership and control and full supply-chain control with no third-country interference at the highest.
Organisations should consult Article 16 and Annex II of the proposal for the operative criteria rather than relying on summaries.
The assurance-level framework reflects an important principle: knowing where a system is located establishes very little about what that system depends upon.
The same principle applies to the evidence those systems produce. A result can be generated on European infrastructure, from lawfully accessed European data, and still depend on an opaque reference resource, an unversioned analytical component or a model whose configuration is no longer recorded.
We describe this as evidential sovereignty: the capacity of an institution to understand and independently verify the lineage and integrity of the evidence on which its consequential decisions depend. It is a question of verification, not of self-sufficiency.
The GDPR remains the principal European framework governing the processing of personal data. It establishes requirements relating to lawful processing, purpose limitation, data minimisation, transparency, security, accountability and individual rights.
Scientific and health research may operate under particular provisions and safeguards, but organisations remain responsible for demonstrating that processing is lawful and appropriately governed.
The data track of the Digital Omnibus proposes targeted amendments. The framework itself is not proposed for replacement.
Provenance does not require unrestricted visibility of the underlying data.
A well-designed evidence system should be capable of preserving information about the origin, processing context and integrity of an evidence artefact while maintaining appropriate restrictions on access to personal information.
For evidence governance, verification and disclosure should therefore be separable. A hash establishes that a file is unchanged without revealing it; an identifier refers to a protected object without publishing it; a permit reference shows that an authorised process existed without granting access to any record.
EOSC is developing a federated environment — a system of systems — through which European research data, tools and services can be found, accessed and reused across institutional and national boundaries.
EOSC policy work addresses metadata, provenance, permanent accessibility, verifiability and interoperability within the federation.
EOSC is the clearest existing example of the problem this page is concerned with: research outputs that are intended to be reusable across infrastructures, by parties who were not involved in producing them.
Reuse at that scale depends on the receiving party being able to establish what an artefact is, what it derived from and which versions mattered. EOSC interoperability work is therefore a natural venue for provenance expectations that travel.
26 March 2027, not 2029 or 2031. That date carries the deadline for the Commission’s key implementing acts and brings the secure-processing-environment provisions into application — the decisions that shape what happens later are made then, not when secondary use itself begins.
Yes. The Digital Omnibus has proposed repealing it and folding its rules into the Data Act, but that proposal is not yet adopted. Until it is, the DGA continues to apply in its own right.
It commits the Commission to launching a standardisation request for a European data quality standard that explicitly covers provenance, alongside completeness, consistency, semantic clarity and governance. That is the most direct European commitment to provenance standardisation we are aware of.
Our term for the capacity of an institution to understand and independently verify the lineage and integrity of the evidence its consequential decisions depend on — a question of verification, not self-sufficiency. It extends the Cloud and AI Development Act’s assurance-level logic from infrastructure to the evidence that infrastructure produces.