Digital Twins for Operational Resilience
How digital twins spot system drift before thresholds trip.
My first experience with digital twins came in 2018 when I founded Valor Cycles, a project with the goal of winning a stage of the Tour de France with an American manufactured bike. The frames were made from carbon tubes, joined at lugs we 3D printed out of chopped carbon fiber. The commercial goal was to be able to create bespoke bikes that fit each customer’s unique size, riding style, and compliance preferences. What I didn’t do was experiment by printing and building bikes. Before a single tube was cut or a lug printed, I simulated each geometry and tube with Autodesk Fusion 360 to solve as much as possible before making a build. I modeled frame flex, weight distribution, and what road riders refer to as “compliance,” which relates to how a bike feels on the road.
In 2018, “digital twin” barely existed as a term outside a few manufacturing journals. It was mostly theory: model the physical object, compare it against what the object did once it existed, see where the two disagreed. Most of what I learned came from what the model didn’t tell me versus what it did, and this fed the learning loop that improved the model.
I shut down Valor in 2019 after deciding that the personal liability outweighed the revenue potential, but it taught me a lot about experimentation in the context of mechanical engineering.
Since then, digital twins have become a technology that I’ve watched emerge from engineering obscurity into a necessity for experimentation in the age of AI, and even an operational fail-safe. This shifts the use case from a simulation you run before you build something to a live system that tells you your operational system is about to have a bad day before it happens.
My hypothesis for why it took this long: digital twins were intimidating. The cost to build, deploy, and interpret one was high, and there were too few experienced engineers who knew how to do it well, so most companies skipped it. What changed is pressure. Companies need to show ROI on their AI investment, and the data center capacity enterprises built out for AI now carries its own pressure to prove utilization and results. That combination is making digital twins commonplace, for experimentation and for operational monitoring. Cheap sensors help too.
The gap before the alarm
Systems rarely fail the way people picture it, with a threshold crossed and an alarm firing at the exact moment things go wrong. The example that always comes to mind for me is the scene in the 2019 HBO series “Chernobyl” where all of the monitors crossed into red and the “all is lost” moment is achieved. A system spends hours, sometimes weeks, behaving slightly differently than it should while staying inside every alarm limit anyone ever set. Again, thinking back to Chernobyl, the engineers began expressing concern when the gauges had deviated from normal, but the manager told them it was a non-issue because they were still within tolerance. By the time the sensors were out of tolerance, the meltdown was well underway and recovery was not possible. Even in 2026, this is still the way many operations perform, but they don’t have to.
Sitting at the beach as I type with my family, which includes six ER physicians, I am reminded that this concept is not new in the ED. A patient can hold normal blood pressure and a normal heart rate while trending toward a crash, and the clinicians who catch it early are reading the trend rather than any single number. A digital twin gives machines and their operations that same instinct. The platform asks whether a value is moving somewhere its own history says it shouldn’t: it’s monitoring for deviation, and how known incoming variables such as a weather event may impact the system’s known baseline.
We all have been raised on alarms: fire alarms, weather alarms, alarm clocks. All of them share the same blind spot. Alarms are built around known limits: max temperature, max pressure, max latency, fire, weather, incoming threat. A twin runs on a model of expected behavior, drawn from equipment specs and historical operating data, and fed by the same real-time data the alarms already see. When the system and the model start to disagree, the twin flags the gap. An alarm-based stack throws that information away by design, and the twin exists to keep it.
A twin can’t forecast an earthquake or a market crash, but it can tell you what to expect if one of these events were to happen. It can also help predict the impact of an inbound hurricane and help plan how to modify operations or output based on expected impact. It can accurately tell you, earlier than any threshold, that a system has stopped matching its own history. In operations, that’s usually the information that matters.
Data centers
The standard data center playbook alarms on inlet temperature, PDU load, and chiller loop delta (a big thanks to my friends at Siemens for the deep-dive into the engineering marvel that the modern data center represents). Each of those is a threshold, and a threshold only reports that a limit has been crossed.
A twin of the same facility models expected thermal behavior for a given load and airflow configuration. When a rack’s fans start running ten or fifteen percent harder to hold the same exit temperature, nothing alarms, because the rack is still in spec. The twin sees the gap, and the gap shows up before the temperature spikes. It usually means a filter is loading up, a CRAC unit is losing capacity, or airflow got redirected by another anomalous change. These events caught in the first week of drift are fixed via a work order. Caught after the threshold trips, and the facility enters an emergency change window at two in the morning with an angry customer demanding a status update from your war room. Having been in my share of those war rooms, I can promise you that they’re never as much fun for the client as they are for the advising consultant who thrives on adrenaline and pressure. Ahem. Not to out myself.
Siemens has been building toward this with Insights Hub, the industrial IoT layer of its Xcelerator platform, which connects a facility’s real-time operational data, IT and OT both, into a model continuously compared against what’s happening on the floor. Siemens has said its campus and facility customers see utility savings of up to 25 percent per square foot running this kind of connected twin, mostly by catching small drifts before they compound. The company pushed further this year with Digital Twin Composer, unveiled at CES 2026, which pulls live data from manufacturing execution systems, quality systems, and industrial IoT sensors to keep a facility’s twin current in close to real time. PepsiCo is running an early version across several US plants and warehouses, and started by using the twin to establish a performance baseline before changing anything at all.
Breweries
A brewery runs on a version of the same physics, just slower and considerably better smelling. A fermentation tank follows a temperature curve: a rise, a hold, a controlled step down, timed to how the yeast behaves at each stage. The glycol system exists to hold that curve.
When a chiller starts drawing more load than the batch’s current stage should require, the system is working harder to hit the same number, and the number itself hasn’t moved yet. A twin of the cellar, built off the temperature, gravity, and CO2 off-gas data the brewhouse already collects, can flag that days before the batch drifts into a stuck fermentation or an off-flavor. Most breweries, even good ones, still catch it from a hydrometer reading and a shift lead’s gut, which works right up until the shift lead is on vacation.
Distribution networks
The same pattern stretches across geography. A lane between a distribution center and a regional hub has an expected transit time and an expected dock dwell time, built from months of runs.
When dwell at a dock creeps up ten minutes at a time, or a reefer trailer’s compressor cycles more often to hold the same set point, no threshold trips. The truck arrives, the product is in spec, and the lane is quietly diverging from its own history. That divergence usually points to a dock about to bottleneck, a compressor a few weeks from failing mid-route, or a driver rotation breaking a schedule nobody re-baselined. A twin of the network surfaces the lane while the fix is still a routing adjustment. Without one, the first confirmation is a load of temperature-sensitive product missing its window, followed by a claim, a customer call, and an argument about whose fault it was.
Bridging the gap to action
A twin that only feeds a dashboard adds very little to its operator. It still depends on someone looking at the right screen at the right time, which is the very dependency the twin exists to remove. The organizations getting value from these systems wire the gap directly into an action with an owner and a clock: a drifting rack opens a work order, a hot fermenter pages the cellar lead, a slowing lane triggers a routing review before the next load is scheduled onto it. It’s like the original Toyota Production System, only running at machine speed. The modeling stopped being the hard part a while ago. The discipline to act before the threshold is where most organizations stall.
There’s a person on the other end of that discipline, or the lack of it. Somebody takes the call at three in the morning to explain what already broke, makes an emergency change half-informed over the phone, and spends the next day writing an incident report for a failure the data had been describing for weeks.
The models I built for Valor in Fusion 360 helped me answer a basic question: will this frame ride the way I designed it? The operational version answers a better one: will this system keep behaving the way it did yesterday? If your monitoring can only report that a limit has been crossed, you’re learning about failures at the same time as your customers. Pick the system whose failure hurts most, model what normal looks like, and give the gap an owner. You likely already have the data, but most operations are throwing it away.
This is the work my team and I are doing at The Select Group. We help enterprises put digital twins to work in both of the modes this post describes: building twins that let teams rapidly test ideas against their strategic imperatives before committing capital, the same way I tested frame geometries before printing a single lug, and standing them up inside live operations as a site reliability layer that catches the drift before the threshold trips. If your organization is sitting on operational data it can’t yet act on, that’s exactly the problem we like. Reach out and let’s talk about it.


