Interveil Labs
The Journal · Essay

Forensic Provenance for Synthetic Worlds

AI is moving beyond language into models of physical reality. The governing question becomes: is this simulated world faithful enough to trust? That's a provenance problem — and it's the one we were built to solve.

For a few years the story of AI has been a story about language — models that read and write, that summarize and converse. That chapter isn't over, but a new one has opened beside it. Researchers at Stanford's Institute for Human-Centered AI, led by Daniel Zhang with Fei-Fei Li and colleagues, recently argued that we are entering the era of world models and spatial intelligence: systems that build working representations of physical environments and predict how those environments change in response to action.[1] A language model completes a sentence. A world model completes a situation — it imagines what happens next if the robot turns left, if the water rises, if the part is placed just so.

This is a genuine shift, and it moves the hardest question in AI from the page into the room. When a chatbot is confidently wrong, you get a bad paragraph. When a world model is confidently wrong — when its simulated physics diverges from the real kind — you get a robot that mishandles a patient, a factory that trains on a fiction, a decision made against a reality that was never really there. The HAI brief names the governing question exactly: whether a simulated environment "matches physical reality closely enough to train or test another system or guide a real-world decision."[1]

Read that question again, because it should sound familiar. It is not a new question at all. It is our question, wearing new clothes.

The same problem, one dimension up

Everything Interveil has built rests on a single discipline borrowed from forensic audio: restore the signal, never falsify the record. Pull the meaning out of the noise — but if something wasn't there, don't invent it, and if the source is uncertain, say so. We have applied that to text (a claim that is fluent is not therefore true) and to reading (enrichment restores context; it never fabricates it). A world model asks for the very same discipline in a higher dimension. A generated environment that looks physically plausible is the spatial twin of a sentence that sounds authoritative. Plausibility is not provenance. Fidelity is not fluency.

So the instrument the age now needs is not another engine for producing synthetic worlds — the largest labs are already spending enormously on that. It is an instrument for trusting them: a way to ask of any simulated environment, where did this come from, and how faithful is it, and can you prove it? Call it forensic provenance for synthetic worlds.

What it would mean

Concretely, it means treating a synthetic environment the way a careful lab treats any piece of evidence. Its lineage is visible and cut-able: you can trace which real data it was built from and which parts are inferred, the way a layered material shows its own grain when you cut it. Its fidelity is graded, not asserted — this region is validated against measurement, that region is extrapolation, this corner is frank guesswork — each carrying its status the way we grade every empirical claim we publish. Its integrity is anchored, so that a world used to train a safety-critical system can be shown to be the world that was actually validated, unaltered since. And through all of it, the human stays the author of the decision the simulation informs — the model proposes a world; a person, seeing its provenance and its grade, ratifies or refuses.

None of that requires out-building the frontier labs. It requires standing one layer above their output and refusing to let a synthetic world become authoritative until it is sourced, graded, and anchored. It is the assayer's role, not the miner's — and it is exactly the layer a small, rigorous house can own while the giants spend the fortunes on capability.

Why it matters now

The HAI authors flag three things that make this urgent, and each one sharpens the case. There are no adequate benchmarks yet for evaluating world models in safety-critical use, and closing that measurement gap needs public investment — which is to say the discipline of grading and publishing the ledger is not a nicety here, it is the missing infrastructure.[1] The data these models need — action-labeled interaction data, robot trajectories, fleet logs — cannot be scraped from the open web, which means it will be captured at scale by whoever can afford the fleets, concentrating control unless provenance and access are designed in early.[1] And the dual-use stakes are real: capable autonomous systems, democratized, shift advantage in ways that make verifiable trust a safety issue, not a branding one.[1]

The wagerWe will not win the race to build synthetic worlds, and we don't intend to enter it. We intend to be the ones who can look at any synthetic world and say, with proof: this much of it is real, this much is inferred, and here is the record.

Between veil and vision, extended to worlds

Our emblem line is "Between Veil and Vision." The veil is the membrane between a mind and everything mediated; a simulated world is simply that membrane made three-dimensional and set in motion. The same commitment carries over without a seam: keep the membrane honest — permeable to context, provenance, and meaning; resistant to the quiet substitution of a convincing fiction for the real. As the frontier turns from words to worlds, the principles don't need to be reinvented. They need to be extended — and someone has to hold the line where the map starts to move.

We think that's worth doing in the open, early, before the synthetic worlds are everywhere and the habit of trusting them without proof has already set. The question was never whether we could simulate reality convincingly. It is whether we will be able to tell, afterward, which parts were true.

How to read these claims

The empirical claims here come from the Stanford HAI policy brief and are attributed as such; our extension of them into a "forensic provenance" discipline is stated as Interveil's position, not as established fact.

  • AttestedThe shift toward world models / spatial intelligence; the benchmark and measurement-science gap; the non-scrapability of action-labeled data and its concentration risk (Stanford HAI, 2025).
  • PositionThat the right response is a verification-and-provenance layer above the model layer ("forensic provenance for synthetic worlds") — argued here, offered for scrutiny, not asserted as settled.

Interveil Labs makes human-agency-preserving AI for media and mind. This essay argues a position and attributes the empirical claims it builds on. Drafted with AI tools, edited and stood behind by D. Hardwick. We make the veil honest.

Sources

  1. [1] Stanford HAI (D. Zhang, Fei-Fei Li, et al.), "The World Model and Spatial Intelligence Era: Governing AI Beyond Language." hai.stanford.edu
  2. Provenance & content authenticity standards: C2PA. c2pa.org · NIST AI Risk Management Framework. nist.gov
Interveil disclosure norm: an argued position paper; empirical claims are attributed and graded; the extension is labeled as our position. Drafted with AI tools, edited and stood behind by D. Hardwick. The record speaks.