
On 3 September 2026, the Allen Institute, University of Washington and Fred Hutch Cancer Center announced AI BioDesign, a new collaborative research accelerator supported by nearly $95 million from the Fund for Science and Technology, the philanthropic vehicle established from the estate of Microsoft co-founder Paul Allen.
Its ambition is considerable. The initiative intends to create open AI models, datasets, assays and tools capable of exploring biological possibilities beyond those produced by natural evolution. The underlying idea is not simply to use artificial intelligence to analyse biology — it is to place AI inside an experimental feedback loop. Models propose biological designs. Researchers construct and test them. Experimental measurements return to the models. Those results influence what the system designs and tests next. The Allen Institute describes this as a continuous design-build-measure-learn cycle.
That distinction matters. As AI moves from interpreting biological observations towards proposing biological objects that researchers subsequently manufacture and test, the provenance problem changes with it. The question is no longer simply “where did this dataset come from?” Increasingly it becomes: how did this biological design come into existence?
Consider a conventional genomic analysis. A biological specimen is collected. DNA is extracted. A library is prepared. A sequencing instrument produces reads. A computational pipeline processes them. Researchers interpret the results. Reproducibility requires good records throughout that process, but the broad direction of travel remains relatively straightforward: specimen → experiment → data → analysis → result.
An AI-driven biological design loop can be considerably more complicated: source data → training dataset → preprocessing → model → model version → design objective → generated design → computational selection → physical construct → experiment → measurement → new dataset → model update → next design. And the process can repeat many times.
The eventual biological object may therefore have both a physical experimental lineage and a computational lineage. Those lineages need to meet. If a protein is proposed by an AI system, synthesised, experimentally tested, modified by a scientist, tested again and ultimately developed further, knowing its nucleotide or amino-acid sequence tells us what the final object is. It does not necessarily tell us how it got there.
One useful test of research infrastructure is deceptively simple: could another researcher reconstruct this result six months later? For AI-designed biology, answering that properly could require knowing which dataset trained the relevant model; which version of that dataset was used; which model architecture and checkpoint generated the candidate; what software and parameters were involved; what objective or constraints were supplied; which candidates were rejected and which progressed; whether a human altered the selected design; which physical construct was actually manufactured; which experimental protocol tested it; and which resulting measurements subsequently entered another training cycle.
That is a substantial evidence chain. At small scale, a laboratory can attempt to capture much of this through electronic notebooks, repositories, model registries, laboratory information systems and disciplined documentation. AI BioDesign, however, demonstrates why scale changes the problem. The initiative explicitly proposes multiplex experiments capable of testing large numbers of designs and intends to generate reusable models, datasets, assays, reagents and benchmarks for the wider scientific community. Contemporary reporting on the project describes experiments involving millions of DNA sequences. At that scale, provenance cannot realistically depend upon somebody remembering to document the important relationships afterwards. It has to become infrastructure.
There is another important feature of the AI BioDesign announcement: openness. The collaboration intends to produce resources that other scientists can reuse. That is enormously valuable. But an open model without reliable provenance is not necessarily a reproducible model, just as an open dataset without adequate metadata is not necessarily a reusable dataset.
This distinction already exists within the FAIR principles for scientific data stewardship. FAIR asks that research assets be Findable, Accessible, Interoperable and Reusable, and the principles explicitly state that reusable data should be associated with detailed provenance. There are also mature approaches for representing provenance: the W3C PROV family of standards models the entities, activities and agents involved in producing data or other objects; Research Object Crate (RO-Crate) provides a machine-readable method for packaging research data together with contextual information such as software, workflows, people and equipment; and Workflow Run RO-Crate extends that thinking to the provenance of computational workflow executions.
These approaches demonstrate something important. The problem is not a lack of ways to describe provenance. The harder problem is maintaining provenance continuously as research objects cross systems, organisations and computational and physical environments.
This is precisely the class of problem that interests us at The BioChain. The BioChain is not an AI biological-design platform, and we are not involved in AI BioDesign, nor are we suggesting that projects such as AI BioDesign lack appropriate provenance systems internally. The more useful question is architectural: what would a verifiable evidence layer for this kind of research need to preserve?
In our view, it should not replace model registries, laboratory information systems, object stores, sequencing repositories, electronic laboratory notebooks or workflow engines — those systems have jobs they already perform well. Instead, an evidence layer should preserve the relationships between the important objects those systems create. For example: Dataset D17 was used to train Model M4.2, which, using Design Objective O81, generated Candidate C14592, which was selected and subsequently represented by Construct X224, which was tested under Assay A19, producing Dataset D18, which subsequently contributed to Model M4.3. Each object can remain in the system where it belongs. What matters is that the relationship between them can be independently demonstrated later.
There is an important difference between recording provenance and being able to verify it. A database field can say that Dataset D18 came from Assay A19. That is provenance. A stronger evidence architecture can additionally establish that a particular version of D18 existed at a particular point, that its relationship to A19 was recorded at that time, that the relevant organisation or system made the assertion, and that the referenced object has not subsequently changed without that change being detectable. That is closer to verifiable provenance.
The BioChain approach is to create cryptographically committed, tamper-evident records around important events and relationships without requiring every underlying research object — particularly sensitive genomic or commercial data — to be placed into one central database or onto a public blockchain. That distinction matters enormously in biological research: the evidence can be shared or verified without necessarily sharing the biological data itself.
Software researchers have long understood the importance of versions. In AI-driven biology, however, versioning becomes part of the scientific evidence. Imagine two apparently identical biological designs: one was generated by model version 4.1, and the other was generated six months later by version 5.0, trained partly on experimental results derived from the first generation of designs. The sequences might be similar. The provenance is not. The second object sits downstream of evidence that did not exist when the first was generated.
That lineage could matter for reproducibility, attribution, intellectual property, regulatory review and eventually safety. It may also matter for understanding why an AI system generated a particular candidate in the first place. The provenance record therefore needs to encompass more than the final sequence — it needs to describe the history of the design.
Open, collaborative biological design creates another question: who contributed what? A successful biological construct could ultimately incorporate contributions from an original public dataset, researchers who generated experimental observations, developers of a foundation model, an organisation that fine-tuned it, a scientist who specified the design objective, an automated selection system, a researcher who manually modified the generated candidate and a laboratory that experimentally validated it.
That does not mean all of those contributors necessarily possess an intellectual-property claim. It means that the evidence needed to establish contribution and lineage becomes increasingly distributed. A trustworthy provenance architecture cannot decide ownership. It can, however, provide much better evidence with which ownership, attribution and responsibility can subsequently be determined.
AI-enabled biological design also carries legitimate biosecurity questions. Work by the Nuclear Threat Initiative and others has examined potential guardrails for AI biological-design tools, including technical safeguards and managed-access approaches. Provenance does not replace those controls. But if biological design increasingly involves automated systems, an audit trail capable of showing which model generated a design, under what authorised workflow, what happened to that design and whether it subsequently entered a physical laboratory becomes part of a broader responsible-research infrastructure. The same evidence architecture that improves reproducibility can therefore support accountability.
That is an important principle for The BioChain: provenance should not be something constructed only when an auditor, regulator or investigator asks for it. It should be a natural by-product of doing the work.
Perhaps the most interesting thing about AI BioDesign is not AI itself. It is the increasingly narrow gap between computational biology and physical biology. A model proposes an object. A laboratory makes it. An instrument measures what happened. Those measurements become data. The data change the model. The model proposes another object. The cycle begins again.
At that point, treating computational provenance, experimental provenance and biological chain-of-custody as separate disciplines becomes increasingly artificial. They are different sections of the same evidence chain. AI BioDesign is an ambitious example of what biological research may increasingly look like: collaborative, automated, iterative, data-intensive and distributed across computational and physical systems.
If that future arrives at the scale its proponents expect, biological infrastructure will need to answer more than “what did we discover?” It will need to answer: which data, model, person, system, biological material and experiment produced this — and can we still prove that years later? That is the problem The BioChain is being built to explore, and the same problem we’ve examined from the sequencing-cost side in our piece on what happens when genomic surveillance becomes cheap enough to deploy everywhere, and from the pure-AI-analysis side in our piece on reproducing what an AI model did to a genome.
AI BioDesign is a collaborative accelerator involving the Allen Institute, University of Washington and Fred Hutch Cancer Center. It intends to combine AI with large-scale experimental biology to design and test novel biological molecules and systems, sharing the resulting models, datasets and tools openly.
Because the final biological object can depend on datasets, model versions, parameters, generated candidates, human decisions and physical experiments. Reproducing the final result therefore requires preserving both the computational lineage and the experimental lineage, not just the final sequence.
Not by itself. Open access can make a model or dataset available, but reproducibility also depends on knowing exactly which versions, workflows, inputs and relationships produced a result — an open model without reliable provenance isn’t necessarily a reproducible one.
No, that’s not the architecture we’re proposing. The underlying research objects can stay in their existing repositories and systems. The BioChain is concerned with preserving verifiable claims about their identity, provenance and relationships, not hosting the data itself.
They allow later verification that a referenced object or provenance assertion corresponds to the object that existed when the record was made, without necessarily exposing the underlying sensitive data — a database field alone can’t prove it hasn’t been altered since.
No. AI BioDesign is used here as a real-world example of an emerging research architecture that illustrates why verifiable scientific provenance is becoming increasingly important — not a project we’re part of.
Regulatory developments, technical notes and platform news — sent occasionally, straight to your inbox.