An Phan

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

San Jose, California, United States

Actions

An Phan is a Senior Data Infrastructure Engineer working at the intersection of data, infrastructure, robotics, decision-making systems, and machine learning.

His work focuses on building the reliability layer beneath AI systems, where telemetry, reproducibility, and historical truth matter more than novelty. He designs architectures for sensor-rich and autonomous systems operating across cloud and edge environments, where signals are incomplete, environments change, and failures often cannot be replayed once they are lost.

Over the years, he has worked across industrial AI, manufacturing, maritime systems, and robotics, where many production failures emerged not from the model itself, but from fragmented ownership, missing telemetry, and systems that could no longer explain reality over time.

Badges

Area of Expertise

  • Information & Communications Technology
  • Real Estate & Architecture

Topics

  • Software Engineering
  • Data Engineering
  • MLOps & AI Infrastructure
  • AI Engineering
  • AI & ML Solutions
  • AI Ethics
  • Physical AI
  • Scalable Data Infrastructure
  • Data Platform
  • Generative & Agentic AI

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

Large language models are transforming how we build software, but physical AI systems expose a hard limit: you cannot recompute reality.
When robots, sensors, and production systems interact with the real world, failures are causal and time-based, not semantic. You cannot go back to record sunlight you missed, human behavior you never captured, or robot telemetry lost to fragile connectivity. Backfills rewrite history. Late data arrives with new stories. Model performance drifts without obvious errors.
In this talk, I will show how these constraints fundamentally change how we design ML data platforms for robotics, agriculture, manufacturing, and other physical-world domains. Using real production workflows, I will walk through how engineers correlate time-aligned telemetry, inference metadata, and operational events to debug subtle drift and root causes that dashboards and LLM-based tooling often miss.
We will cover concrete architectural patterns for capturing irreversible data reliably at the edge, building reproducible ML pipelines when recomputation is impossible, managing late and out-of-order data without rewriting history, and unifying analytics, training, and debugging on a shared data backbone.
I will close with where LLMs do fit powerfully in this workflow, and why physical AI still needs fast, reliable analytics as its foundation.

From Prompts to Physics: Designing APIs for Real-World AI Systems

Most APIs were designed for software systems where requests, responses, and failures happen entirely in the digital world. Physical AI systems introduce a different challenge. Robots lose connectivity. Sensors drift. Actions have delayed consequences. The state of the world changes while requests are still being processed.

Drawing from production deployments across agriculture robotics, manufacturing, and industrial AI systems, this session explores how API design changes when software must interact with physical reality. We'll discuss event-driven architectures, asynchronous workflows, telemetry feedback loops, observability patterns, and strategies for handling unreliable edge environments.

Attendees will learn practical patterns for building APIs that remain reliable when machines, sensors, and real-world operations become part of the system.

Designing a Cloud-Edge Data Backbone for Physical AI Systems

Physical AI systems in robotics, industrial automation, and other real-world environments rely on continuous telemetry from sensors, machines, and human-in-the-loop actions. Unlike cloud-native software systems, these signals represent irreversible real-world events that cannot be reconstructed if they are not captured when they occur. However, many production data pipelines still assume that data can be recomputed or backfilled, which leads to irreproducible training datasets and blind spots in debugging and drift analysis.

This session presents a production cloud-edge data architecture that treats telemetry as immutable historical truth and unifies raw sensor data, inference metadata, and operational events into a time-aligned, append-only record. The architecture separates capture correctness from downstream compute, preserves late-arriving data, and enables reproducible reconstruction of training datasets. Practical workflows for model drift debugging and historical dataset reproduction are discussed, along with trade-offs in storage cost and operational complexity.

Why Most AI Platforms Break in Production

Many organizations successfully build AI proofs of concept, only to see those systems struggle once deployed in production. Models continue to run. Pipelines execute as expected. Dashboards appear healthy. Yet performance drifts, trust erodes, and teams can no longer explain what changed or why.

The root cause is rarely the model itself. More often, the problem lies in the platform and organizational architecture around data. Backfills rewrite historical datasets. Features are recomputed without clear provenance. Inference runs without traceability. Teams move quickly in isolation, and no one owns end-to-end reproducibility.

This talk reframes AI reliability as a systems and organizational design problem rather than a modeling problem. Drawing from real-world AI systems and machine telemetry pipelines, it shows how data lineage, telemetry capture, observability, and reproducibility must be treated as core platform capabilities rather than afterthoughts.

Attendees will learn how organizational structure, team boundaries, and architectural decisions interact to either preserve or undermine trust in AI systems over time. The session concludes with practical guidance on structuring teams and platforms so AI systems can scale without accumulating silent data and technical debt.

WeAreDevelopers World Congress 2026 - North America Sessionize Event Upcoming

September 2026 San Jose, California, United States

API World + CloudX + AI TechWorld 2026 Sessionize Event Upcoming

September 2026 Santa Clara, California, United States

IEEE Cloud Summit 2026 Sessionize Event

June 2026 Washington, District of Columbia, United States

AI DevSummit + DeveloperWeek Management 2026 Sessionize Event

May 2026 South San Francisco, California, United States

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

San Jose, California, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top