Session

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

Large language models are transforming how we build software, but physical AI systems expose a hard limit: you cannot recompute reality.
When robots, sensors, and production systems interact with the real world, failures are causal and time-based, not semantic. You cannot go back to record sunlight you missed, human behavior you never captured, or robot telemetry lost to fragile connectivity. Backfills rewrite history. Late data arrives with new stories. Model performance drifts without obvious errors.
In this talk, I will show how these constraints fundamentally change how we design ML data platforms for robotics, agriculture, manufacturing, and other physical-world domains. Using real production workflows, I will walk through how engineers correlate time-aligned telemetry, inference metadata, and operational events to debug subtle drift and root causes that dashboards and LLM-based tooling often miss.
We will cover concrete architectural patterns for capturing irreversible data reliably at the edge, building reproducible ML pipelines when recomputation is impossible, managing late and out-of-order data without rewriting history, and unifying analytics, training, and debugging on a shared data backbone.
I will close with where LLMs do fit powerfully in this workflow, and why physical AI still needs fast, reliable analytics as its foundation.

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

San Jose, California, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top