Every software engineer has a debugger. You set a breakpoint, step through execution, and watch state change one instruction at a time. When something goes wrong, you don't reason about it from a dashboard, you replay it.

Manufacturing has no debugger. A factory generates enormous telemetry: machine states, work order transactions, quality results, and operator sign-offs. All of it lands as rows in tables. But the thing that actually produced the defect, like a fixture seated wrong, a part staged on the wrong side of the bench, a torque tool shared across three stations that's unavailable for eleven minutes, happens in the gap between the log lines. The transaction says the step completed. It doesn't say how. That gap is what Factory Playback closes.

What Factory Playback is, and the problem it actually solves

Factory Playback synchronizes video from cameras on the production floor with the operational event stream captured by Tulip onto a single timeline. Click a failed leak test in the record and the video jumps to the second it happened. Ask "show me every time a technician re-worked the harness routing on line 2 last week" and you get the clips, not a report.

The framing matters for anyone building vision agents: this is not a video analytics problem, it's a grounding problem. Point a vision language model (VLM) at raw factory footage and you get generic captions like "a person is working at a table." The model has no idea of operational context like that the table is Station 7, that the person is executing step 4 of work order 88213, that the torque spec is 42 Nm, or that the last three units through this station failed the final test.

Operational visibility is inherently hard in these environments for three structural reasons:

  • Large, dynamic physical environments. Multiple stations, wide work zones, and material, tooling and fixtures that move constantly. There is no fixed geometry to anchor to.

  • Engineer-to-order production. The product at a given station may be unique or highly customized. A model trained on "what a correct unit looks like" degrades immediately.

  • The work lives between the events. Systems capture transactions. The physical context that explains them is missing.

This is exactly where NVIDIA and Tulip’s technology stacks meet. NVIDIA models and compute bring accelerated perception and reasoning capability; Tulip's event stream brings the operational context that tells those models which moments matter and what they mean in the language of the factory.

Why Tulip uses NVIDIA

NVIDIA has spent years building the model and microservice layer for physical AI: NVIDIA Cosmos for world reasoning and generation, the NVIDIA Metropolis Blueprint for video search and summarization (VSS) , NIM for deployable inference, and Omniverse for simulation. It's a deliberately general foundation, built to be adapted to a domain rather than presuppose one. Adapting it to industrial operations calls for one input that is genuinely scarce: labeled operational reality at scale. Synthetic data generation, which Cosmos is exceptionally good at, multiplies that seed, but something has to seed it, and the seed has to be a faithful record of how skilled people actually perform complex physical work. That is not a dataset anyone can crowdsource. It only exists where the work is already being executed through software.

Tulip has it as a byproduct of the day job. The platform runs in 1,000+ factories across 45 countries, with roughly 19 billion events a year flowing through frontline apps, and 41,000+ active applications spanning assembly, inspection, kitting, and batch execution. Every one of those events is a human-authored, timestamped assertion about what was supposed to happen and what did.

There's a nice symmetry here that we've started calling kaizen as Reinforcement Learning from Human Feedback (RLHF). Continuous improvement is already a feedback loop where humans observe work, judge it, and correct it. Formalized and timestamped, that loop is a preference signal. It’s a ground truth for what good physical work looks like, generated continuously by the people doing it.

So the division of labor is clean: NVIDIA builds the brain. Tulip is the nervous system and the runtime that connects it to real work.

How the stack fits together

Factory Playback is built on the NVIDIA Metropolis VSS blueprint uses NVIDIA Cosmos 3 as the reasoning VLM, with Tulip supplying the operational index. The pipeline runs on-premise, so camera feeds never leave the building, which is a hard requirement in GxP, ITAR and works-council environments.

Ingest and detection. Station cameras stream into the VSS ingestion path on an on-prem GPU server, currently NVIDIA RTX PRO 6000 class hardware at the edge. Video is chunked, and a NIM-deployed object detection model (YOLOX in the current configuration) produces per-frame detections with tracking IDs, so objects persist across chunk boundaries rather than being re-discovered every few seconds.

Reasoning. Chunks go to NVIDIA Cosmos 3 for dense, timestamped captioning and spatio-temporal reasoning. Cosmos 3's mixture-of-transformers architecture pairs a reasoning transformer with an expert generation transformer, so the model reasons about object interactions, motion and spatial-temporal relationships as a first-class operation rather than inferring them from frame-level captions. For Playback, the payload we care about is timestamp precision and relational understanding: not "a component is clamped in a fixture," but "the fixture sits unclamped from 00:03:47 to 00:04:12 while the unit waits." That interval is the actual finding.

Retrieval. Captions, CV metadata and Tulip events are embedded and indexed through NeMo text embedding and reranking NIMs, which is what makes the timeline queryable in natural language instead of by camera and clock.

The Tulip half. Tulip's event stream does three jobs no vision pipeline can do for itself.

  • It supplies the schema: station, work order, operator, app step, machine state, quality result – a vocabulary for what a moment is, not just what it looks like.

  • It supplies the triggers: a failed test or an Andon pull tells the pipeline which seconds of footage are worth reasoning over, which is the difference between a tractable inference budget and boiling the ocean on 24/7 multi-camera video.

  • And it supplies the clock: one authoritative timeline that video, telemetry and human action are all aligned against.

That last one sounds mundane and is the whole trick. Without a shared clock, you have two archives. With it, you have a record. Downstream, the same aligned sequences feed Omniverse real operational behavior as the grounding layer for digital twins, so simulation is visualized and validated against what the floor actually does rather than against a process engineer's assumption of it.


What it's worth

The clearest way to see the payoff is to look at what an investigation costs today. When an issue comes back from the field, the team has to manually rebuild the story of what happened out of whatever survived: machine logs, exported reports, and the recollections of whoever was on shift. That work can take hours or days and usually pulls in several people, and what it produces is still a reconstruction rather than a record.

What makes it expensive is not a shortage of expertise. Teams generally know the starting point and the end result; what they don't have is a reliable picture of what happened in between, so most of the effort goes into manufacturing evidence after the fact. When that evidence already exists as a replayable timeline, the investigation starts roughly where it used to finish.

Four categories of value follow from that, and they compound as coverage expands across a plant:

  • Throughput. Cycle-time variation between builds becomes visible and measurable rather than anecdotal. Low single-digit percentage gains in throughput are meaningful in high-value discrete production, where each point translates directly into units out the door.

  • Quality and rework. Correlating quality events with the video moment that produced them turns defect analysis from statistical to causal. Teams stop debating which of four plausible mechanisms caused a failure mode and watch the one that did.

  • Anomaly detection. Machine behavior that degrades gradually — and therefore never trips a threshold — shows up when you can replay the same operation across weeks and see the drift.

  • Safety and productivity. Unnecessary movement, waiting time and awkward station layouts are obvious on a replayable timeline and nearly invisible in a transaction log. The same evidence supports line balancing and workstation redesign.

In modeling exercises at heavy discrete-manufacturing sites, these categories have scoped to seven-figure annual opportunity at a single plant. Those are modeled figures, not realized results, and they depend heavily on production value density, but the direction is consistent across the deployments we've scoped.

There's a second-order return that matters more over time. Aligned video and events accumulate into a structured record of how the process actually runs, including the practical know-how that today lives in the heads of a plant's most experienced people and walks out the door when they retire. Captured as process data, that expertise becomes teachable to the next person hired, and it becomes the grounding layer for whatever industrial AI gets built on top. It accrues as a byproduct of running the plant.

The broader point for anyone building on Cosmos: the models are no longer the bottleneck. Grounding is. The teams that get real value out of physical AI in industrial settings will be the ones who can hand the model a structured, human-authored account of what the work was supposed to be and then let it reason about the gap between that and what happened. Factories have never been able to debug reality. Now they can step through it.


See the architecture, live

Everything above is the short version. If you want the long one, Rony Kubat, Tulip's co-founder and CISO, is hosting a LinkedIn Live session where he'll walk through the Factory Playback architecture end to end: how station video moves through the VSS ingestion path on the edge GPU, how Cosmos 3 reasoning gets scoped by Tulip's event triggers instead of running on every frame, and how operational events get indexed into one queryable timeline, and where the hard engineering problems actually were. He'll also take questions, so bring the ones this post didn't answer.

Register for the LinkedIn Live →
Wednesday, September 30 at 2pm EST/11 am PST

Get hands-on at Operations Calling

Join manufacturers already getting return on AI at Operations Calling. Build a solution from an SOP, work with connected machine and video data, and move it between sites.

Bottom CTA Global