For years the conversation was about the wrong thing. Headsets or glasses. VR or AR. Consumer adoption or enterprise. Which platform wins. These debates are understandable — device cycles are visible, headsets ship or don't ship, funding moves. But they live at the surface. They focus on XR as an interface layer while ignoring the stack beneath. And that stack is where the real shift is happening.
While the market debated devices, the XR, simulation and digital twin industry built something different: real-time 3D engines, physics models, sensor fusion, low-latency pipelines, spatial anchoring, digital twins that connect to live data. Without that stack, autonomous systems don't "understand" space — they detect signals and execute commands. Physical AI doesn't revive XR. It exposes XR as infrastructure.
This page is the architecture of that shift — from visualization to enabler to infrastructure to merged reality — for the companies already building the spatial substrate that Physical AI is going to need, whether or not that was the market they originally set out to serve.
The same underlying technology — real-time 3D, spatial computing, digital twins — can occupy four distinct architectural positions. Where it sits isn't a reflection of how capable the technology is. It's a strategic choice about what role it plays — and what else depends on it.
Industrial training, flight and military simulation, engineering design reviews, remote operations support, soft skills development. These are not prototypes — they are embedded in operational workflows. Pilots training in immersive simulators. Engineers validating systems before physical deployment. Teams rehearsing complex procedures in controlled environments. The industry may not be in a hype phase. But infrastructure rarely looks exciting when it stabilizes. It looks boring. Reliable. Measurable. That is often the moment when real value compounds. The model is built to be looked at, then set aside. Nothing downstream depends on it surviving past the session.
The model stops being rebuilt from scratch and starts connecting to something live — telemetry, IoT feeds, a PLM or ERP system. It updates instead of getting redrawn. In robotics, this appears as world models. In industrial environments, as digital twins. In simulation and training platforms, as virtual worlds connected to live data — the same kind of environment already used to train autonomous vehicles and now increasingly used to train robots. Different vocabulary. Same architectural function: a shared spatial reference layer. Real progress — and also where most platforms plateau: a better dashboard, but not yet something else actually depends on.
Leading research now points in the same direction: MIT Technology Review describes manufacturers moving from asset-level twins toward system-level, closed-loop twins that connect design, simulation and operations into a continuous feedback cycle. Capgemini finds 76% of executives say moving from pilots to scaled deployment is a significant challenge — not because the technology doesn't work, but because the architecture was never designed to carry the operational load. The spatial model is no longer a representation. It becomes the reference model machines and AI use to act, the interface humans use to understand and intervene, the coordination layer across systems. The industrial metaverse made this step before most of the industry noticed: it wasn't a detour on the way to Physical AI — it was the precondition for it.
XR stops being the interface to a model you inhabit alone and becomes the interface to a shared reality occupied simultaneously by humans, AI agents and autonomous systems. At that point, the spatial model can no longer be a private representation — it has to become what I call an operational twin: the live, shared layer through which every actor coordinates, that arbitrates when systems disagree, and that governs who acts on what and when. Not a place you visit. A layer the physical world acquires. The full architecture of this is what the section below develops.
Not because XR doesn't work. That is the easy explanation. The harder truth is structural.
The buyer treats it as an innovation experiment. The supplier delivers a bespoke project. The demo impresses. But scaling a demo to an operational deployment requires something completely different: repeatable architecture, integration with real workflows and legacy systems, deployment playbooks, enterprise-grade support, and customer success at scale.
I went through this transition several times in different companies and verticals. It is one of the most difficult scaling thresholds a deep-tech company faces — the challenge of standardising the portfolio while retaining the ability to configure and extend the platform for specific client needs.
The XR suppliers that break through make a deliberate transition: from project-based delivery to product-oriented operating models. That transition rarely happens by itself.
The right question is not "Can we build an impressive XR experience?"
It is: "Can this become a repeatable capability inside the way the organisation actually works?"
That is where the real scaling work begins — and it is also, precisely, where the architecture decision underneath the product becomes impossible to defer.
This pattern has a name. Architecture built for pilots, not platforms — it's one of five recurring failure modes in Physical AI and XR deployments, each with a specific diagnostic signal.
See the five failure patterns →Companies that position themselves as content providers will compete in device cycles. Companies that position themselves as spatial infrastructure providers will shape the stack. That is not a go-to-market choice. It is an architecture decision — and it has to be made before the next integration locks in the wrong foundation.
The training module, the visualisation, the immersive review. Each one is valuable and each one is self-contained. When the session ends, the model can be discarded without anything else breaking. Replacing the vendor next year costs roughly what it cost to engage them the first time. The relationship re-procures rather than compounds.
A planning system queries the model. A maintenance workflow updates it. An autonomous system navigates by it. It has to stay current because something else's correctness depends on it staying current. Replacing it means replacing what has been built on top of it — which is exactly why it is worth more, and why it shapes the stack rather than competing inside it.
This is the same logic as compound innovation, applied one level down — to a single product rather than a whole market. Each capability that makes the spatial model something other systems rely on compounds the next: better data in, more systems depending on it, more reasons it cannot be removed. Capabilities that stay self-contained — however impressive the rendering — stay additive. They make the demo better. They don't make the business harder to leave. PwC projects a €430 billion global Physical AI market by 2030. Several platforms already sitting in XR and digital twins are closer to that value than their current positioning suggests.
"In my own work across industrial simulation and XR deployments, the limiting factor has rarely been immersion quality. It has almost always been the integrity and synchronisation of the underlying spatial model. When that layer is coherent, XR scales. When it is fragmented, initiatives stall — regardless of how impressive the interface looks."
Choosing infrastructure over content is the first decision. The second is what that infrastructure eventually has to coordinate — not just one platform's twin, but every operator, authority and system sharing the same physical space once enough of them are running.
See where this leads →This is the architecture decision, one company at a time: how the runway has to evolve as it matures past this threshold — the general pattern behind the content-vs-infrastructure choice above.
See the architecture page →Building the interoperability architecture underneath the XR stack. Creating the infrastructure category on top of it. Then scaling enterprise adoption from inside it. That sequence is the authority this page is built on.
Co-founded and helped scale a platform that made exactly this transition on purpose. RealityOS started where most XR platforms start — visualization and collaboration — and was deliberately re-architected into industrial-metaverse infrastructure: an operational layer connecting AI, simulation, XR, digital twins and IIoT, not a viewer bolted onto each of them separately. ROOM3D and HOLOH weren't features added to an existing product — they were new categories, holographic and spatial communications, built because the underlying infrastructure could support categories that didn't exist before. Scaled from zero to 150 people and over €10M in funding inside that transition.
The interoperability and federation problem underneath the XR stack isn't new — I built a platform to solve it a decade before Physical AI had a name. Simware solved the connectivity problem: how to federate simulators and real systems to train as you fight, how to evolve monolithic simulations into interoperable simulation services running cloud-to-edge, and how to expand the use of simulation beyond training through the concept of the Internet of Simulations. Adopted by more than fifty customers across Europe, the US and China — the first commercial implementation of this architecture, co-published in peer-reviewed form at IEEE SoSE 2017. The spatial infrastructure layer now emerging in Physical AI is solving the same interoperability problem at a larger scale, with higher stakes.
Saw the same transition from the adoption side — and from a platform that makes it concrete. Threedy is turning digital twins into active engineering collaboration infrastructure: digital twins integrated with agents and knowledge graphs to convert static 3D visualisation into an intelligent layer for engineering collaboration, where the model doesn't just represent the asset but reasons about it. OEMs like BMW, Porsche, Daimler Truck and Stadler Rail weren't just adopting a viewer — they were embedding a system their engineering workflows would depend on continuously. The work was making that dependence reliable at scale: multi-project management, third-party integrations, delivery discipline that held as adoption grew past the first deployment. That is the actual difference between a successful pilot and an infrastructure relationship — and it is where most platform vendors get stuck, not in convincing the customer the technology works, but in becoming something the customer cannot easily remove.
Imagine an engineer connecting to a wind farm — not through a video call but by stepping holographically into the environment through a robotic avatar. Wearing a lightweight XR headset, they see, hear and interact with colleagues and machines in real time, assisted by an AI co-pilot predicting maintenance needs before failures happen. At headquarters, another team operates through a digital twin: a living virtual replica mirroring every action in real time. Remote and onsite teams collaborate as if co-located, interacting seamlessly with intelligent systems and cobots. That is what happens when AI, XR, edge computing and digital twins fuse into one intelligent ecosystem.
Once a spatial model is infrastructure rather than a tool, XR's job changes. It is no longer the interface to a private walkthrough — it becomes the interface into a shared replica: a virtual version of a real space that more than one person, and more than one system, can occupy and act through at the same time. Smart spaces. Smart systems. Autonomous agents, increasingly. All present, all acting, inside the same merged reality your platform already knows how to render.
When that happens, something uncomfortable follows. The shared spatial model is no longer a representation — it becomes infrastructure. And whoever controls the digital twin influences real-world outcomes. In conversations with cities, governments and infrastructure operators, one concern keeps coming up: no one wants to outsource their reality to a platform. This is not about content or data. It is about the model of reality itself — who defines what is true, who can read and write it, how it is governed when systems disagree.
The moment multiple systems act on the same spatial model, you have quietly inherited a problem XR was never built to answer on its own. If a robot's world model and a technician's headset describe the same room differently, whose version governs what happens next? That is not a rendering problem or a latency problem. It is a governance problem — and it is exactly the problem the operational twin is designed to solve: a shared runtime coordination layer, live and authoritative, through which every system, agent and human in the same environment coordinates. Not a digital twin in the marketing sense. The layer that governs distributed autonomy in shared physical space.
The full architecture of the operational twin — the fighter pilot, the corridor problem, the world model distinction. Built out in full on its own page.
Go deeper: The Operational TwinIf the technology is already closer to infrastructure than your pricing or your pitch deck admits, that's an architecture and positioning conversation — not a roadmap item to revisit next quarter.
Repositioning what you've built is an architecture decision. Whether the organisation behind it can actually sell, deliver and support that shift — without unwinding the per-seat, per-project model that got you this far — is what the Scaling System Maturity Framework assesses, across Technology, Organisation and Trust simultaneously.
If what you've built already has the bones of an infrastructure business — and most of what's on this page suggests it does — the question isn't whether to make the shift. It's how to do it without breaking what already works for your current customers.