October 5, 2026
Open access made much of the scientific record openly available. There is now an enormous opportunity to make what science has learned computable and systemic: representing methods, measurements and findings in a common, connected form across disciplines.
Global R&D expenditure is now approaching $4 trillion a year. Yet when science reports what it has learned, its basic unit is still the paper: a document written for a human and usually filed under a single discipline. This creates two problems that need to be solved together:
The first is a formalization problem. The core objects of scientific reasoning: what was done, what was measured, and what was found are locked in prose and figures. Machines can’t reliably read and compute over them, and neither can the AI systems increasingly reshaping how research is done.
The second is a systems problem. Knowledge is organized by field, journal or department, but the world is not. Today’s complex challenges that matter the most, from pandemics to climate change to chronic disease, are systemic and cut across these boundaries. The knowledge needed to address them is scattered across evidence bases that rarely meet.
What science needs is a representation that is both computable and systemic. A finding in cardiology and a finding in environmental epidemiology should be expressed in the same terms, and should meet whenever they share a variable. Here, we set out an agenda for building that representation, and the parts of it that already exist.
Thirty years ago, the Bermuda Principles established a radical norm for the Human Genome Project: publicly funded sequence data should be rapidly released into the public domain. Since then, open science and open access have moved from radical propositions into the mainstream. About 68% of research published in 2025 indexed by OpenAlex is open access, up from 31% a decade earlier.
Researchers have spent nearly two decades making pieces of the scientific record machine-readable. The FAIR Principles established machine actionability as a goal for research data; Nanopublications have explored persistent identity and provenance for individual assertions; initiatives like OMOP and CDISC have demonstrated the value of common models at scale; and efforts such as the Open Research Knowledge Graph show how contributions buried in papers can become structured knowledge.
What’s missing is a representation of the scientific record at the scale of the scientific record itself, and one that crosses and connects domains. Earlier approaches have asked researchers to do the formalizing, with little formalized science to connect to, so the returns on effort have been slow.
AI has upended the economics of formalizing science. Autoformalization is possible: we can now extract, structure and connect the existing scientific corpus at scale, at a cost that was previously impractical. As scientists increasingly use AI in their workflows, new science can be born computable and connected. Formalization can stop becoming a burden on researchers and become something they get value from immediately.
AI systems already reconstruct pieces of this structure every time they read the literature: identifying entities, interpreting measurements, extracting claims and connecting evidence across papers. But today, much of that work is repeatedly done transiently and non-deterministically inside individual models and products.
If we do nothing, computational representations of science will increasingly be recreated independently across organizations, often inside proprietary systems that are difficult to inspect, reuse, or improve.
The alternative is to treat this representation as shared scientific infrastructure, always open, persistent, interoperable, and connected back to the evidence from which it was derived.
The early evidence for this way of working is promising. Earlier this year, Nasri et al. (2026) showed that giving AI agents a deterministic interface to viral genomic data dramatically improved their accuracy at retrieving the right datasets. Their work describes “deterministic data access as critical infrastructure for reliable agentic science.”
In materials science, Lee et al. (2026) found that structuring experimental literature increased LLM retrieval accuracy from 21% to 71% using the same RAG pipeline, and combined with improved query reformulation, accuracy reached 86%.
And in our own work — our differential diagnosis algorithm, which pairs AI with our clinical graph, achieves industry-leading performance.
What would a deterministic interface onto the whole scientific record, connected across scales and fields, as one system, make possible?
The necessary infrastructure to accomplish this requires representing three related layers of scientific reasoning, each dependent on the one beneath it:
Methods describe what was done: study design, protocols, populations, interventions, and analysis. Without them, it’s hard to tell if similar results are comparable. Some domains have rich schemas but there is no general layer describing how most studies were conducted.
Measurements describe what was observed: the variables, instruments or assays, units, timing and context. Ontologies and common data models cover pieces, but they are fragmented by domain and mostly describe types of measurement, not the measurements individual studies actually report.
Findings describe what was learned: which variables relate, in which direction, by how much, with what uncertainty in whom. They are what scientists, funders, and AI systems most want to reason over, and are the largest gap.
Measurements are also the crux of solving the systems problem: the same variable, such as blood pressure or particulate exposure, appears across thousands of studies in dozens of fields under different names. If you can reliably resolve those variables, findings from separate literatures can connect into one system.
System was founded with a purpose to relate everything, and we have spent eight years building towards this infrastructure. Important parts of it are open today.
The System Graph makes findings first-class objects. Instead of treating a result as a sentence buried in a paper, we represent it as a structured relationship between variables, together with its direction, magnitude, statistical context, population, and provenance.
Today, the graph contains over 14 million findings extracted from the scientific literature and other trusted sources. Because every finding uses the same structure, whatever its field, the graph is connected by design: formalizing a finding is also the act of placing it in the system.
That changes what researchers can do. Instead of reading papers one by one, they can query findings directly, compare evidence across studies, surface replications and contradictions, and assemble the system of evidence around a problem. Users can build algorithms on top of the graph. AI systems get grounded, structured data to reason over. Evidence stays linked to its source, so errors can be found and fixed rather than disappearing into model weights. The System Public Graph, opens this through an API and bulk downloads for research use.
The Open Variable Index, coming soon, will be a versioned, citable, and openly downloadable index of the scientific variables catalogued in the System Public Graph. It will expose both grounded and ungrounded variables, how they are used, and where possible how they map to external ontologies. It is a map of measurement space showing where existing standards and ontologies give good coverage, and where the gaps are.
Building shared scientific infrastructure that becomes widely adopted is an organizational problem as much as a technical one. It needs a formalization model that already works at scale, a way of sustaining itself, and a credible commitment to staying open. Few organizations hold all these. Here is where we stand:
This work is far from finished. The System Graph grows every day as new science is indexed, yet still only represents part of the overall record and a subset of topics it could cover. The automated extraction used to build the Graph is reliable but not perfect. Variable extraction and grounding also works well, but a large proportion of variables we extract remain ungrounded in existing taxonomies and vocabularies.
These are all reasons to build this infrastructure openly, not reasons to wait. Because every object stays connected to its source, all of these issues can be seen, measured, and corrected by anyone.
Crucially, this is infrastructure that should be built once and reused, rather than repeatedly reconstructed inside separate projects and products. Over the next year, we want to work with three communities:
Researchers: Use the Public Graph and Open Variable Index on real scientific problems and tell us where they work and where they don’t. The best test of this infrastructure is whether it helps researchers do things better than they can do today.
Builders: Build on it. Create agents, evidence tools, visualisations, meta-analyses, and entirely new research workflows on top of it. Suggest which datasets and ontologies we should connect to. We want the value created on top of this layer to be much larger than anything System could build alone.
Funders: System’s open-science effort has received incredible support from the Wellcome Trust; now, we call on more organizations to partner with us to help expand this shared scientific infrastructure to become even more durable and useful. Because commercial products on the same infrastructure will in time fund the pipeline that maintains it, support here buys expansion and openness rather than basic operation. This agenda needs:
These are high-leverage investments. As it stands, the world already spends $4 trillion a year generating new knowledge. What it does not yet have is a way for that knowledge to accumulate as one connected, computable system rather than millions of separate objects.
Thirty years ago, a generation of researchers and funders established the principle that scientific knowledge should be open. Openness gave us a corpus. The next step is to make it a computable system, and to build that system in the open.