
Europe knows where its servers are. It doesn’t know where its evidence came from.
The EU has learned that location tells you nothing about dependency. That lesson applies to scientific evidence too — and the window to act on it closes in March 2027.
In June 2026, the Commission published the Cloud and AI Development Act. Its most argued-over feature is a ladder: four assurance levels that a cloud provider must climb to serve sensitive public-sector workloads. Level 1 asks only that processing and storage happen inside the Union. By Level 4, nothing foreign can reach the workload — not by ownership, not by staffing, not through the supply chain, not by operation of a third country’s law.
Most of the commentary has been about what this means for American hyperscalers. The more interesting thing is the admission built into the design. The Commission has accepted that knowing where a system runs tells you almost nothing about what that system depends on. Sovereignty, on this reading, is not a property of geography. It is a property of visibility: can you see the chain of things your critical capability rests on, and can you assess them?
That is a significant intellectual move, and it is correct. It is also incomplete, because Europe has applied it to infrastructure and not yet to the thing infrastructure produces.
Consider a European research environment a few years from now, built exactly as policy intends. The data sit in the Union. The compute is European. The permits are in order and the secure processing environment works as designed. A researcher produces a risk estimate that goes into a regulatory submission.
Now ask the CADA question of that estimate rather than of the servers. Which source extracts did it come from? Which cohort definition, which build of which statistical package, which version of which reference resource? Was a model involved, and if so, which one, configured how? If any of those change next year, will anyone be able to say which conclusions were affected?
The infrastructure can score Level 4 while the answer to every one of those questions is “it’s in someone’s notebook, probably, if they still work here.” The dependency is no longer on a foreign cloud provider. It is on undocumented institutional knowledge. That is still a dependency, and it is not one that sovereignty policy currently sees.
This is not a complaint about European infrastructure. It is a consequence of the thing European infrastructure does well.
Take DARWIN EU, the network EMA uses to generate real-world evidence for medicines regulation. As of 2026 it reports data reaching roughly 250 million patients across Europe, drawn from around forty onboarded data partners, with 108 research topics assessed in the most recent reporting year. Its design principle is that the data stay where they are. Local partners map their records to a common model, run the analysis locally, and return results. Nothing identifiable is pooled into a European super-database, and nothing needs to be.
This is a genuine achievement. It is also precisely why provenance becomes hard. The thing that moves between institutions is a result, not a record. Each partner keeps excellent internal logs. The study protocol exists. The code exists. The mappings exist. They simply exist in different places, in different systems, under different custodians, expressed in ways that made sense locally. Nothing is missing. Nothing is portable either.
The European Health Data Space will sharpen this considerably. Its privacy architecture is deliberately strict: secondary use happens inside a secure processing environment, access is anonymised where the purpose allows and pseudonymised only where it does not, and users may download only non-personal data. That is right. But it means the moment of greatest evidential risk is the moment the system is working best. An approved aggregate leaves the environment. It enters a publication, a submission, a model, another study. The governance context that produced it — the permit, the source categories, the analytical workflow — has no reason to travel with it, and currently doesn’t.
Nobody has failed here. Each institution is well governed. What is missing sits between them.
The failure mode is rarely fraud, and it is rarely error. It is forgetting.
Staff leave. Repositories migrate. Software is retired and suppliers change. Annotation databases update; reference genomes are revised; classification thresholds move. A conclusion that was entirely intelligible when it was produced becomes progressively opaque, not because anyone did anything wrong, but because the environment that made it legible no longer exists.
Regulators discover this when a decision is reopened five years later. Public health authorities discover it mid-outbreak, when a sample sequenced in one Member State and interpreted against a European reference resource has to be reconciled with epidemiology from three others, and the reconstruction takes days it does not have.
There is a commercial version of the same cost that anyone in life sciences will recognise. Organisations spend real money recreating provenance for audits, submissions, partner collaborations and due diligence: working out which file produced which output, which version was used, who approved what. The same questions get answered from scratch every time, because the relationships were never recorded in a form anyone else could reuse. That is not a one-off compliance overhead. It is a charge levied on every reuse of evidence, which is the exact activity European policy is trying to increase.
It would be easy to read this as a pitch for a new European platform. It is the opposite.
The technical work is largely done. W3C PROV has described entities, activities and agents for over a decade. RO-Crate and research-object approaches package data, code and context together. Persistent identifier infrastructures make objects and actors referenceable over time. FAIR Digital Object work and EOSC interoperability activity address much of the rest. Nobody needs to invent a vocabulary.
What is missing is not syntax. It is an expectation. A standard can describe a lineage; it cannot require that the lineage be created, that it survive export from a secure environment, that the receiving institution be able to read it, or that anyone other than its author be able to verify it. Those are governance questions, and they are the ones nobody has answered.
The answer need not be elaborate. A small, interoperable record — what the artefact is, what it came from, who or what made it, when, by what method, which versions mattered, what integrity evidence exists, under what authority, and what depends on it downstream — would carry most of the weight. It exposes nothing. A hash proves a file is unchanged without revealing it. An identifier refers to a protected object without publishing it. A permit reference shows an authorised process existed without granting access to a single patient record.
Europe does not need a new initiative to start. The Data Union Strategy, published by the Commission in November 2025, already commits it to a standardisation request for a European data quality standard covering completeness, consistency, semantic clarity, governance — and provenance. That request is the natural home for a minimum evidence record, and it is open now.
The second venue is the EHDS. Its implementing acts are due by March 2027, and the same date brings the Chapter IV provisions on secure processing environments into application. Whatever those specifications say about what an environment records, and what accompanies an approved output when it leaves, will be built to by twenty-seven national health data access bodies. Revisiting it afterwards will be expensive. The commonly cited 2029 and 2031 dates are when secondary use arrives. 2027 is when the decisions are made.
None of this requires scientific autarky, and none of it requires centralisation; European science depends on global data, software and collaboration, and should continue to. It does not require publishing sensitive information, because verification and disclosure are separable. It certainly does not require a blockchain. And it will not make bad science good — a perfectly preserved chain can document a poor decision with great precision.
What it would do is let a European institution answer a question about its own conclusions: where did this come from, what does it depend on, and can someone other than its author check? Europe has decided that question is worth asking about its cloud infrastructure. It is at least as worth asking about the evidence that infrastructure exists to produce.
A framework of assurance levels a cloud provider must meet to serve sensitive EU public-sector workloads, from Level 1 (data processed and stored inside the Union) up to Level 4 (no foreign ownership, staffing, supply chain, or third-country legal reach into the workload).
No. It explicitly does not require scientific autarky or centralisation — European science depends on global data, software and collaboration. The proposal works through portable, interoperable provenance records, not a central repository.
It applies the same underlying argument — that Europe governs data access but not evidence lineage — to the specific lens of the Cloud and AI Development Act’s sovereignty framework, rather than to the EHDS and Biotech Act directly. See the full argument in From Data Governance to Evidence Governance.
Two venues are already open: the Data Union Strategy’s standardisation request for a European data quality standard (which already names provenance), and the EHDS implementing acts due by March 2027, which will determine what a secure processing environment records and what travels with an approved output when it leaves one.
This piece develops an argument set out in full in our policy white paper From Data Governance to Evidence Governance, which proposes a European minimum evidence record and a staged pilot programme — download the full paper here.
If your organisation generates, manages, analyses or governs biological evidence and would be interested in participating in a UK or European provenance demonstrator, The BioChain would welcome the conversation.
Get in touch
Regulatory developments, technical notes and platform news — sent occasionally, straight to your inbox.