
There is an appealing assumption at the centre of modern biological artificial intelligence: that more data produces better models. Sometimes it does. But there is a qualification that deserves considerably more attention — more inconsistent data can produce more confidently inconsistent models. On 1 September 2026, researchers writing in Nature Biotechnology argued that reference materials should be routinely used as common calibrators in multiomics studies, co-profiled alongside experimental samples so that results can be expressed relative to a shared standard rather than treated as isolated absolute quantities. Their stated objective is reproducibility. The implications extend considerably further, because as biological datasets increasingly become training material for artificial intelligence, the quality of the measurement process becomes part of the integrity of the model itself. AI does not begin with an algorithm. It begins with a measurement.
Consider a seemingly straightforward biological quantity. A researcher measures the abundance of a molecule in a sample; another laboratory measures the same molecule. It would be convenient to assume that the resulting numbers describe precisely the same biological reality. They may not. Different laboratories can use different instruments, different reagents, different sample-preparation protocols, different software, different reference databases, different analytical pipelines, different thresholds and different quality-control procedures. Measurements can also drift over time: an instrument used in January is not necessarily behaving identically in October, a software update may alter processing, a reference database may change, and a new reagent lot may introduce another source of variation. These effects are familiar to experimental scientists. AI gives them a new significance.
Suppose an AI system is trained using multiomics data generated across a hundred laboratories. The model encounters variation, and some of it represents genuine biology — disease, population differences, environmental exposure. But some of it may simply represent the fact that Laboratory A measured something differently from Laboratory B, and the machine does not inherently know which is which. If technical variation is sufficiently systematic, an AI system may learn it anyway, becoming extremely good at detecting which laboratory processed a sample rather than the biology researchers hoped it would discover. This is one manifestation of a much larger problem: artificial intelligence can learn the history of our measurement systems alongside the biology those systems were intended to measure.
The argument advanced in Nature Biotechnology is therefore important. Reference materials can act as common calibrators: instead of treating every measurement as an isolated absolute quantity, researchers can co-profile a shared reference material alongside study samples, and express results relative to that reference — sample divided by reference, rather than simply sample. That provides an anchor across instruments, laboratories and experiments, much as agreeing on a common ruler before comparing measurements made in different countries. Without a common ruler, numbers may appear precise while remaining difficult to compare.
This leads to the next question. Which reference material, which batch, when was it produced and how was it characterised? Which laboratory used it, on which instrument, under which protocol, which software version and which quality-control thresholds? Which transformation converted the raw observation into the value ultimately supplied to the AI model? Reference materials improve comparability, but the reference itself becomes part of the evidence chain, and that chain must be preserved just as carefully as the sample it was meant to calibrate.
Consider what happens when multiomics data becomes AI training material. The lineage might resemble: biological sample, sample preparation, reference material, instrument, raw measurement, signal processing, quality control, normalisation, feature extraction, dataset, dataset harmonisation, training corpus, model, prediction, experimental validation. At every stage, information can change — some transformations physical, others computational; some performed by humans, others automatically; some reversible, others not. By the time a prediction emerges from the end of that chain, its relationship with the original biological material may be extremely difficult to reconstruct. Yet scientifically, that relationship still matters.
AI has encouraged a computational understanding of reproducibility. We preserve code, record software environments, version models, save random seeds and containerise analytical pipelines. These are good practices, but biological AI introduces another requirement: the computational workflow can be perfectly reproducible while the biological evidence entering it is not. Run the same model against the same dataset and it will faithfully return the same result — that is computational reproducibility working exactly as intended. It says nothing, however, about whether the measurements inside that dataset were themselves reproducible, or whether a systematic laboratory effect has simply been reproduced alongside the biology. A model can be perfectly deterministic and still be confidently wrong, because determinism guarantees that an analysis can be repeated, not that the evidence it was repeated on was sound in the first place.
It is worth being precise about what reference materials do and do not fix. They address measurement-level reproducibility: whether two laboratories measuring the same biological quantity arrive at comparable numbers. They do not, on their own, address evidence-level provenance: whether anyone downstream can reconstruct which reference batch was used, which instrument generated a given measurement, or which version of a pipeline produced the value that ultimately reached a model. A biomedical AI system can therefore fail in two distinct ways that are easy to conflate. It can fail because the underlying measurements were never made comparable in the first place — the problem reference materials are designed to solve. Or it can fail because the measurements were comparable, calibrated and correct at the time, but nobody preserved enough information to know that, years later, when the model’s behaviour needs to be investigated. Reference materials solve the first problem. They create, rather than remove, a need to solve the second.
This is not a narrow methodological point confined to multiomics. We have made a version of this argument before in the context of the European debate over biological data infrastructure for AI-driven medicine: building the infrastructure to generate biological data is only half the problem; the evidence flowing through that infrastructure also has to remain traceable as it moves between laboratories, repositories, models and, eventually, regulators. Reference-material calibration is exactly the kind of detail that determines whether that traceability is possible in practice. A provenance-aware infrastructure needs to record not only that a sample was measured, but which reference material accompanied it, so that a value expressed as sample-over-reference can still be interpreted correctly when the underlying reference standard is itself revised, replaced or discontinued.
At The BioChain, this is precisely the kind of detail our evidence layer is designed to hold onto. We are not proposing another multiomics repository, calibration standard or analytical pipeline — those are scientific instruments that should remain in the hands of the laboratories and standards bodies that build and maintain them. What we record is the relationship between them: that dataset D was measured relative to reference material R, using instrument I, on date T, processed by pipeline P version V, and later incorporated into training corpus C. If R is later recharacterised, or P is updated, or a laboratory’s instrument is later found to have drifted, that relationship lets researchers work out exactly which downstream models and predictions are affected, rather than treating the entire training corpus as equally suspect or equally trustworthy by default.
For how this connects to the wider transatlantic and UK infrastructure picture, see Europe Is Building the Infrastructure for AI Biology. It Should Build the Evidence Chain With It..
Britain is now running directly into this problem at national scale — see The Missing Layer in Britain’s AI-Ready Health Data Infrastructure.
The broader lesson is that AI-ready biology needs more than volume. It needs measurements that are calibrated against something held in common, and it needs the history of that calibration to survive every transformation between the instrument and the model. A dataset can be enormous and still be untrustworthy if nobody can say which reference standard it was measured against, or whether that standard changed halfway through collection. Reference materials give biology a common ruler. Provenance is what lets us prove, years later, that the ruler used at the start of a study is the same one being referred to at the end of it.
That reference materials should be used as common calibrators in multiomics studies, co-profiled alongside experimental samples so that results can be expressed relative to a shared standard rather than as isolated absolute measurements, improving reproducibility across laboratories, instruments and time.
Not necessarily. If technical variation is systematic — for example, if one laboratory consistently measures a quantity differently from another — an AI model can learn that pattern as if it were biological signal, rather than the pattern cancelling out.
It is related but distinct. Reference materials address whether measurements are comparable in the first place. Provenance addresses whether anyone can later reconstruct which reference, instrument, batch and pipeline version produced a given value. A dataset can be well-calibrated and still poorly documented, or vice versa.
No. It guarantees an analysis can be repeated exactly, not that the underlying biological measurements were themselves reliable or comparable. A perfectly reproducible computational workflow can still faithfully reproduce a measurement artefact.
The BioChain is not a calibration standard or a multiomics repository. It is an evidence layer that can record the relationship between a dataset, the reference material and instrument it was measured against, and the pipeline that processed it, so that relationship remains reconstructable even after any one of those things changes.
If your organisation generates, manages, analyses or governs biological evidence and would be interested in participating in a UK or European provenance demonstrator, The BioChain would welcome the conversation.
Get in touch
Regulatory developments, technical notes and platform news — sent occasionally, straight to your inbox.