Session

Private, Offline AI in the Browser with WebGPU and Transformers.js

Modern browsers include a GPU, a neural-network runtime, and enough compute to run useful AI models locally. That creates a different architecture for features that need privacy, offline operation, or low-latency interaction: no inference request has to leave the device.

This session runs text generation, image classification, and semantic search directly in a browser tab using Transformers.js, ONNX Runtime, and WebGPU. Along the way, we will examine the limits that determine whether a model belongs in the browser: download size, quantization, warm-up time, memory pressure, browser support, and the point where a server is still the better answer.

Attendees will leave able to select browser-appropriate models, choose between WebGPU and WASM fallbacks, design an offline-capable inference flow, and explain the privacy and operational tradeoffs of client-side AI.


Audience: Web developers and architects evaluating client-side AI.
Format: 45–60-minute technical talk with live browser demonstrations.
Demo: Local text generation, image classification, and semantic search with no inference backend.
Technology reference: https://github.com/huggingface/transformers.js
Evidence: New reusable session; currently in evaluation for Live! 360 Tech Con Orlando 2026.
Materials: Talk-specific repository, performance notes, slides, and recording are not yet published.
Vendor scope: Vendor-neutral web standards and open-source runtimes.

Ron Dagdag

Microsoft MVP / Research Engineering Manager @ Thomson Reuters

Fort Worth, Texas, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top