Call for Papers

Call for Papers is closed. Submissions are no longer possible. Sorry.
in 3 months

PyTorch Conference North America 2026 - POSTER SESSIONS

event starts

20 Oct 2026

event ends

21 Oct 2026

location

San Jose Convention Center San Jose, California, United States


Join us in San Jose, CA, October 20-21 for PyTorch Conference North America 2026. This two-day event hosted by the PyTorch Foundation gathers top-tier AI pioneers, researchers, and developers to explore the future of open source AI and the impact of PyTorch Foundation projects like PyTorch, vLLM, DeepSpeed, and Ray.

PyTorch Conference features in-depth technical talks, hands-on workshops, and candid conversations spanning the full AI stack, from bare metal infrastructure to applications and agent-based systems. The program features keynote sessions from leading voices in AI and practical deep dives on training, inference, GenAI, and focused tracks on responsible AI & compliance, security & privacy, frameworks & compilers, and more.
finished 7 days ago
Call for Papers
Call opens at 12:00 AM

31 Mar 2026

Call closes at 11:59 PM

26 Jul 2026

Call closes in Pacific Daylight Time (UTC-07:00) timezone.
Closing time in your timezone () is .

CFP GUIDE

Please review our CFP Guide to answer many common questions before submitting. 

PLEASE NOTE - Only an abstract is required with your Poster submission. If selected to present a poster, you’ll receive additional details about printing requirements and event logistics by email.

DATES TO REMEMBER:

  • Poster Session CFP Close: Sunday, July 26  at 11:59 pm PDT (UTC +-7)
  • Poster Session Notifications: Monday, August 17, 2026 
  • Poster Sessions Announced: Tuesday, August 18, 2026
  • Event Dates: Tuesday, October 20 - Wednesday, October 21, 2026

Reminder: This is a community event — so no product and/or vendor sales pitches.

CODE OF CONDUCT

By submitting, you agree to The Linux Foundation's Code of Conduct.

COMMITMENT TO INCLUSIVITY

Please review The Linux Foundation's Inclusive Speaker Orientation and Inclusive Language Initiative.

Please be mindful that The Linux Foundation does not allow any talk with 3 or more speakers to only have men participating in an effort to increase equity & inclusion.

PRIVACY POLICY

At The Linux Foundation, we are committed to safeguarding your privacy. Sensitive speaker information will only be accessible to event organizers and program committee members who adhere to the highest confidentiality standards. Rest assured that your information will never be sold or shared beyond these parties.

Speaker personal information including name, company, job title, biography and photo will appear on the public schedule. For information on our privacy practices and commitment to protecting your privacy, please review our Privacy Policy.

You can update or delete your account information at any time through your Sessionize profile account settings. Please contact support@sessionize.com directly for any questions or issues. 

QUESTIONS?

Question about submitting a proposal? Contact us at cfp@linuxfoundation.org.  


all submitted sessions

publicly listed on this page
292 submissions
Submitted sessions
Angel Li
  • Resolving Numerical Differences in Attention for RL
Krutarth Bhatt
  • Agentic Triton Kernel Development: Using AI Agents to Accelerate GPU Kernel Optimization
Krutarth Bhatt, Bhavya Shah
  • Scaling LLM Inference with Megakernels
Evgeny Leksikov
  • Network fault tolerance for AI workloads
Salvador Escobedo
  • Fractional-Order Memory in Single-Cell Habituation
Monica Song
  • Neptune Serving Engine: General Purpose Inference
Benjamin Ummenhofer, Sameer Sheorey
  • KernelFoundry: Hardware-Aware Evolutionary GPU Kernel Optimization
Jinjie Liu, Chunlei Men
  • Trident: Compiling Triton Kernel Launch Paths from Python to LLVM via MLIR-Based JIT Compilation
Joshua Lochner
  • Agentic Kernel Optimization at the Scale of the Web
Mankeerat Singh Sidhu
  • Koi: A Self-Learning Planner for LLM Inference on Heterogeneous Multi-Cloud GPUs
show all submissions
Pratinav Seth, Zera Lyngkhoi
  • AlignTune: One Workflow for LLM Post-Training
Klaus Zimmermann
  • Rebuilding the Build: PyTorch's Move to a Standards-Based Packaging Backend
Oh Jiseong, Hoon Choi
  • Building On-Device AI Agents on Exynos with ExecuTorch
Shabista Haider
  • The Foundation Effect: How Umbrella Governance Reshapes vLLM and DeepSpeed
Krishna Patel, Rahul Vishwakarma
  • Fenceline: a self-hosted prompt firewall for locally served open-weights models
Kushagra Rastogi
  • Understanding Distributed Checkpointing
Parshant Sharma
  • Don't Tune It Yourself: Autotuning GPU Kernels with Helion
Harshita Varma, Nikita Verma
  • Safe Checkpoint Provenance using Sigstore + Safetensors
  • Profiling CUDA Memory Fragmentation in Dynamic Shape PyTorch Workloads
Nikita Verma, Harshita Varma
  • Characterizing Graph Breaks in torch.compile Across Production Vision Models
  • Understanding KV Cache Memory Behavior in vLLM for Long Context
  • Static Security Analysis of PyTorch Checkpoints for Safe Model Deployment
Tanvi Gupta
  • Firetap: Model Internal Telemetry for Large-Scale ML Training
Derrick Johnson
  • Physical AI on the Edge: Adapting and Deploying VLMs with ExecuTorch on Qualcomm HTP
  • Spatial Reasoning for Edge Robotics Using Parameter-Efficient Vision-Language Models on Ventuno Q
  • From Simulation to Edge AI: A PyTorch Sim2Real Deployment Workflow for the Arduino Ventuno Q
Harsh Shah
  • From PyTorch to Edge: Accelerating GenAI Inference with ExecuTorch on Qualcomm NPUs
Felix Baum
  • On-Device LLM Inference on Android with ExecuTorch and Qualcomm QNN
Frank Lin
  • Memory Management in CUDA Graphs
Sesh Seshagiri
  • From Hugging Face to Edge: Enabling Native PyTorch Model Onboarding with Optimum–ExecuTorch
Shiksha Patel, Rishi Teja Madduri
  • MoRI + vLLM: Wide Expert Parallelism and RDMA KV-Cache Transfer for Disaggregated MoE Serving on AMD
Shashank Hegde, Mohamed Husain Noor Mohamed
  • Performance and optimization of distributed MoE on Arm-CPU clusters
Priya Joseph
  • ExecuTorch and Monarch powered Snowflake CoCo applications
  • Smart routing for packing many model lifecycle functions such as training, RL, inference,serving
Manad Desai
  • Red-Teaming Self-Hosted Agents: What Each Injection Defence Stops, Breaks, and Costs
Niyati Parameswaran, Vitalii Dziuba
  • TorchTPU Agent Suite Automating PyTorch Model Migration and Benchmarking on Cloud TPU
Joy Song
  • Bringing vime RL Post-Training to ROCm: End-to-End Training and Rollout on AMD Instinct
Shrey Modi, Krishna Patel
  • Seal Forward, Diagnose Backward: Signed Evidence Around a PyTorch Inference Step
Youpeng Zhao
  • Reliable and High-Availability LLM Serving with In-place Failover and Elastic Recovery
Krish Malik
  • Accelerating PyTorch LLM Inference by Reducing Tensor Layout Overhead
Radim Urban
  • TCCL: Low-latency Communication Backend for Apple Silicon
Chunxing Yin
  • Composable Kernel HSTU Attention Kernel for Generative Recommenders
Ishan Aryendu
  • Retrieval-Augmented Autotuning: Navigating the Quality–Readiness Frontier with Historical Evidence
Rishi Sinha
  • Integrating RCCL GPU-Initiated Networking into TorchTitan’s MoE Communication
  • GraphTrainer on AMD Instinct: Bringing a New Distributed-Training Paradigm to Eager Parity
Nathan Wichmann, Naveen Ravi
  • Extending OpenSHMEM for AI and MoE Communication Patterns
Geonsik Moon, Kaoutar El Maghraoui
  • Closing the Source-to-Kernel Provenance Gap for Out-of-Tree PyTorch Accelerators
Manpreet Sokhi
  • From Bottlenecks to Autonomy: Profiling RLHF Pipelines with PyTorch TorchTItan
Abigail Fernandes
  • Scaling PyTorch Validation with Neuron as the success story
Rudraksh Karpe, Shivay Lamba
  • From FlashAttention to DeepGEMM: IO-Aware Kernel Design for Inference workloads
  • When Acceptance Rates Are Not Enough: An Inference Systems Perspective on Speculative Decoding
Shreya Srivastava, Lucas Hendren
  • Integrating Out-of-Tree Accelerators with PyTorch Profiler
Ian Ballantyne
  • "Are We There Yet?" (AWTY): A Framework for Assessing Small Model Viability on Local Hardware
Suvrakamal Das, Shivay Lamba
  • Five Bits on Metal: Adding Q5_K Quantization to ExecuTorch's MLX Backend
Elfie Guo Guo, Syed Ahmed
  • TorchTitan Performance Optimizations
Mansi Agarwal, Parshant Sharma
  • Beyond Registration: What It Actually Takes to Be a First-Class PyTorch Backend
Mansi Agarwal
  • The Graph Break You Didn't Expect: Making torch.compile and FSDP2 Work Together
Ratnakar Malla, Anurag Sharma
  • Building Reproducible ML Infrastructure for Digital Pathology with PyTorch
Xinyu Kang
  • One Policy, Two Engines: Closing Training–Rollout Drift in DeepSeek-V4 Flash RL
Kyle Huang
  • From MP4 to Tensor: Optimizing Multi-Camera Video Data Loading for PyTorch Training
Sherlock Huang
  • TorchTitan GraphTrainer: LLM Distributed Training with PyTorch’s Compiler Toolkit
Riya Punia, Mansi Agarwal
  • Scaling PyTorch Validation for a Multi-Accelerator Future
Naveen Ravi, Nathan Wichmann
  • Connectionless Scale-Out EP: GPU-Initiated MoE Dispatch/Combine for vLLM
Justin Turney, Erik Postma
  • Getting Cozy in the Scratchpad: Compile Time Memory Planning for the Spyre Dataflow Accelerator
Yufeng Lyu
  • Zero-Code Model Portability: Running HuggingFace Transformers & Diffusers on Any Chip via Torch-FL
  • One PrivateUse1, Many Chips: A Virtual Device Layer for PyTorch
Parichay Das
  • From MCP to On-Device AI: Building Multimodal Agents with SafeTensors and ExecuTorch
  • Trust Before Inference: Verifiable Multimodal Model Deployment with SafeTensors and ExecuTorch
  • From Model Integrity to Runtime Integrity: Trusted MCP Agents on ARM Edge Devices
Luka Govedič
  • vLLM IR: Graph Compiler Intermediate Representation for Fast LLM Inference
Georgia Phillips
  • Upgrading CPU Preprocessing Modules from TorchScript to PT2: IR Design, Developer Experience, and Pr
Hang Xiao
  • Triton-Based Operator and Compiler Optimizations for LLM Inference Across AI Accelerators
Yiqun Huang
  • Accelerating Modern Attention Kernels with Triton-TLE
Riya Punia, Arkadip Maitra
  • Testing the Edge: Scalable Validation for ExecuTorch
Apoorv Gupta, Yang Jin
  • Neuron - PyTorch Native Backend for Trainium
Sam Jin
  • Neuron Native PyTorch Backend for Trainium
Masahiro Hiramori
  • Portable Semantics, Specialized Kernels: PyTorch FlexAttention in ONNX
Jalil Alva, Sebastian Villalobos Alva
  • Evaluating Local INT4 SLM Judges on Apple Silicon
Wei Chen US, Ville Kallioniemi
  • Production FP8 Training for Multimodal Transformers in PyTorch
Huanran Wang
  • FlexBudget: Interchangeable KV Importance Signals under One Compiled FlexAttention Budget
  • The Agent Roofline: When GPU Offloading Helps (and Hurts) in LLM Agent Runtimes
Subrata Goswami
  • Scaling TorchTitan on Intel XPU Architecture : Very Large MoE LLM Training at 10K Scale.
  • AutoBalance: A Self-Improving Agent Architecture for Imbalance Mitigation in Distributed Execution
  • TorchTitan on Intel XPU: From Desktop to Supercomputer - A Guide to Large-Scale LLM Training
Witold Czubala
  • From Client Communications to Competing Risks: A PyTorch Transformer for Client Attrition Model
Nate Mauer
  • Fidelity Over Capability: Distilling a Paid LLM Judge from Its Public Verdict Archive
  • Reward Hacking in Long-Horizon LLM Agents: A Health-Gated Reward for an OpenEnv Benchmark
Akshat Kaul
  • The Grader Looked Fine Until the Model Learned to Hack It
Aishwarya Ramasethu
  • Serving Interpretability Probes on Open Models
Vishal Jaimin Vakil, Dr Harsh Kapadia
  • AugVein: Vein Detection through Augmented Reality with PyTorch
Debashish Chakraborty
  • Phalanx: Adaptive Kernels for Sparse Retrieval Encoding
Venkatesh Bellale
  • When Does Offloading FSDP's Reduce-Scatter to a DPU Actually Pay Off?
Michal Shalev
  • Hardware-Accelerated Strided KV-Cache Transfers for Asymmetric LLM Inference
Shubham Jain, Ritik Raj
  • MAKCI: A Multi-Abstraction AI Kernel Cost Instrumentor for Performance Modeling in AI Compiler
Manan Chawda, Nikunj Doshi
  • Semantic Tensor Contracts: Detecting and Localizing Train–Serve Corruption in PyTorch
Jerome Mitchell
  • Evaluating TorchComms Adoption in vLLM on XPU
Aditya Srikanth
  • Tuning the Compiler, Not Just the Kernel: Bespoke -O3 for your PyTorch Helion GPU kernels
Sean McGovern, Arkadip Maitra
  • When Should I Disaggregate? Portable Diagnostics for LLM Inference on vLLM and llm-d
Mathew Odden, Alessandro Sangiorgi
  • A Remote Autotuning Service for Helion: Async, Parallel Tuning on Production Hardware
Manpreet Singh
  • 85% of the Time Was Preprocessing: A Vendor-Portable GPU Kernel for Real-Time ICU Monitoring
  • The Same Kernel Bug Costs 10x More on AMD: Vendor Asymmetry in Protein-Binding Code
  • From 129 Seconds to 79: Fused GPU Kernels for Whole-Slide Pathology Foundation Models
Abdulsalam Bande
  • ERRC: Entropy-Reinvested Residual Correction for Tensor-Parallel LLM Communication
Aneesh Shetty
  • Serving Video Diffusion on Trainium by fusing Communication, Attention, and Quantization
KC Hema Prasanna, Pradipta Ghosh
  • Layout-Aware Compilation Cache Validation for Custom AI Accelerators in PyTorch
Shruthip Venkatesh, Pradipta Ghosh
  • No Hardware? No Problem : A Compiled-Artifact Execution Layer for Out-of-Tree PyTorch Backends
Tao Huang
  • Batch Cost Lookahead: Coordination-Free Workload Regularization for Packed Variable-Length Training
Abhinav Srivastav, Abhijeet Pendyala
  • Step-Aligned Telemetry for Diagnosing Distributed Training Slowdowns
Santosh Bahir
  • Pytorch Legate Interoperability through DLPack
Deval Shah
  • Inference Performance Analysis using TraceLens
Manu Nicholas Jacob
  • What actually governs edge inference on a Raspberry Pi 5
Emilio Belmar Quijada
  • A 3D U-Net Surrogate for CFD Prediction in Industrial Enclosures
Vedant Malik
  • TraceWeave: Eliminating Weight Reloads for Faster LLM Inference
Raj Thakur, Finn Thompson
  • Bringing High-Performance Inference Serving to Trainium with an Out-of-Tree vLLM Plugin
Julian Ng-Thow-Hing
  • ExecuTorch WebGPU: PyTorch's First GPU-Accelerated Inference in the Browser
Rob Timpe
  • Improving the torch.compile user experience
Chun Tao, Souvik Kundu
  • Telemetry-Complexity Aware Adaptive Routing for Agentic Workloads on Heterogeneous Hardware Systems
Adarsh Singh
  • Neuron Native PyTorch Runtime - An Asynchronous Execution Runtime
Yuchen Jiang, Ashraf Mahgoub
  • Efficient EPD Architecture to Accelerate VLM Performance on AWS Trainium
Amit R
  • Agents as Distributed Processes: A Native PyTorch Design Pattern for Multi-Agent Orchestration
  • Gradient Ghosts: Real-Time Numerical Fault Detection for PyTorch-Native Training
Maria Lyubimtseva
  • AI Edge Quantizer: Streamlining PyTorch Model Optimization for LiteRT On-Device Inference
Christine Cheng
  • Race-Free TorchInductor Artifact Reuse for Elastic Training
Alan Zhu
  • GenieX: A Dual-Runtime PyTorch Inference Backend for NPU/GPU/CPU Deployment on Snapdragon
Shuang Yu, Huiying Li
  • Day-0 RL at Scale for Emerging LLM and VLM Architectures
Yuhe Zhang
  • From Domain Documents to Custom Embeddings: A PyTorch Recipe for Better Retrieval
Tzu-Hsin Yang, Yidi Wu
  • Regional-AOTI: A PyTorch-Native Partial Compilation Framework For High-Performance Inference
Jiaqi Zeng
  • Multi-Teacher On-Policy Distillation for SOTA Agentic AI models
Elham Harirpoush, Fadi Arafeh
  • Accelerating Whisper Inference on Arm CPUs through PyTorch Ecosystem Optimizations
Jinsun Yoo
  • Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
Yanwen Xu
  • Day-1 Cooperative Matrix Support in ExecuTorch: Accelerating On-Device Inference on Mobile GPUs
Matthew Lowdon
  • Beyond export: building Raspberry Pi camera apps from PyTorch vision models with ExecuTorch
Bhushan Sonawane
  • From Cloud to Device: A Guide for Deploying AI Models on Embedded Systems Using Quantization
Lan Luo
  • Bridging the Edge Deployment Gap: High-Performance PyTorch with Torch-TensorRT and ExecuTorch
Chun Tao, Todd Malsbary
  • Hypothesis-to-Skill Loop for Agentic SYCL Kernel Optimization
Frost Mitchell
  • Empirically-guided Tensor Placement in AutoParallel
Dominic Zhang, Sara Mchugh-Grant
  • Hospital Bed Assignment: A PyTorch-Forecast and LLM Decision Support Architecture
Anubhav Jana, Mansi Agarwal
  • Turning Device Backend Bring-Up into Upstream-Tested Confidence: A PrivateUse1 Journey with Spyre
Filippo Simini, Guoqiong Song
  • Reinforcement Learning for LLMs with Intel XPU and PyTorch on ALCF Aurora
Nurlan Nazaraliyev
  • Automatic Graph-Based Activation Checkpointing and Offloading for TorchTitan Frontier Model Training
Apurba Bose US, Naren Dasan
  • Beyond the Python Runtime: AOT-Compiled Expert-Parallel MoE for Edge Deployment in TensorRT
Volodymyr Kysenko
  • The Power of Graph-level Optimizations for LLM Inference on Edge CPUs: Efficient SDPA with YNNPACK
SoonEe Ong
  • RAFT for API Code Generation Under Distractor Noise
Yannick Schnider, Sophie du Couédic de Kergoualer
  • Beyond the Execution Layer: Integrating Custom Accelerators with vLLM
Srujan Teja Thomdapu
  • Efficient Segmented LoRA Inferencing on Intel GPUs
Pratinav Seth, Vinay Kumar Sankarapu
  • DLBacktrace: Tracing How AI Models Reach Their Decisions
Pratinav Seth, Soham Bhattacharjee
  • CuratorKIT: Building Better LLM Training Data from Raw Documents
Pratinav Seth, Aditya Tanna
  • TabTune: One Interface for the Tabular Foundation Model Lifecycle
Pratinav Seth
  • SafeTune: Auditing and Repairing Safety Drift After Fine-Tuning
Pratinav Seth, Hem Gosalia
  • CircuitKIT: From Circuit Discovery to Model Intervention
Kevin Zhao
  • Helion Beyond the GPU: A Qualcomm Hexagon Backend for PyTorch
Xiaoyan Liu
  • Cross-Language Compilation in Triton: From CUDA Source to Distributed GPU Kernels
Sreyashi Chatterjee
  • Fail Fast or Restart? Automatically Triaging Distributed PyTorch Training Failures
Siddharth Narayanan
  • Efficient Data Representations for Large-Scale PyTorch Recommendation Training
  • Scaling PyTorch Recommendation Models with Request-Level Training
Jagadish Krishnamoorthy
  • Grouped GEMM on AMD ROCm in PyTorch
Shubh Pachchigar, Vishal Jaimin Vakil
  • An Empirical Study Of torch.compile across Inference Engines
Jiahao Tan
  • veRL: Extreme Optimization Practices for Ultra-Large-Scale MoE Models
Hao Chen
  • vLLM Ascend: Production-Grade Inference on Huawei Ascend NPUs
Tai-Hsiang Peng, Chao-Lin Lee
  • Support Helion Flow to MLIR Linalg for RISC-V Vector Extension
Takuya Nakaike, Mori Ohara
  • Beyond SPMD: Compiling Triton Kernels for Dataflow AI Architectures
Tao Chang
  • Multi-Chip KV-Cache Transfer for Disaggregated vLLM Serving with FlagCX
  • FlagCX: Portable Device-Side Communication Primitives for Triton Kernels on Any Accelerator
Gilliean Lee
  • Solving Tail-Latency Inflation in Production Vision-Language Serving with Virtualized GPUs
Gilliean Lee, Chun Tao
  • Resource-Efficient Secure and Privacy-Preserving LLM Inference via GPU Virtualization
Ziqi Hong, Fengchun Hua
  • Triton-Ascend: Simplifying NPU Kernel Programming and Enabling the Open-Source Kernel Ecosystem
Kushal Mittal
  • Session-Aware Routing for Local Agentic Workflows: Mixture-of-Models on Intel Workstation GPU
  • Scaling Satellite Object Detection: A PyTorch Geospatial Pipeline on Intel XPU
Phalani Paladugu, Mohamed Husain Noor Mohamed
  • Scaling LLM Inference Beyond Accelerators with CPU Clusters
Boram Yoon
  • StageFrontier: Always-On Stage Accounting for Distributed PyTorch Training
Shrinath Thube, Vaibhav Tupe
  • Which Quantization Scheme Should You Actually Use? A Decision Tree for INT8, FP8, GPTQ, and AWQ
  • From Training to Your Pocket: Mapping the PyTorch Ecosystem End-to-End
  • What Privacy Actually Costs: An Accuracy-vs-Epsilon Curve for LoRA Fine-Tuning
  • Sandboxing Agentic Tool Calls: An Isolation Architecture for Untrusted Execution in Production
Shikhar Mathur
  • Your Agent Passed the Benchmark. Then Production Broke It.
Mengmei Ye
  • Offloading KV Cache to Secondary Memory: A Three-Tier Hierarchy Emulator using vLLM
  • Resilient LLM Serving with Shared Secondary Memory: Fast KV-Cache Failover Across Deployment
Mengmei Ye, Maroon Ayoub
  • Where the Blocks Live: Tier-Aware KV-Cache Routing for vLLM at Scale
Ofer Achler
  • Using HugePages with PyTorch for Faster LLM Inference
Chun-Lin Huang, Wei-Shen Huang
  • EvoTuner: Evolutionary Search with a Learned Cost Model for Triton Kernel Autotuning
Jui-Ting Chen, Jenq-Kuen Lee
  • TorchInductor for SGLang Inference on RISC-V CPUs
Stanislaw Wozniak, Francesco Fusco
  • Enabling APC for hybrid models in vLLM
Eylon Toledano
  • Transparent Pipelined Compression for PyTorch Inference via a NIXL Utility
Venly wu
  • FlagRelease: A Self-Verifying Agent Platform for Model Migration Across Heterogeneous Accelerators
Aravind Neelakantan
  • Fault-Tolerant PyTorch FSDP at Scale: Benchmarking NVRx
Guang Liu
  • KernelGen: A Self-Evolving AI Kernel Engineer for PyTorch Across Heterogeneous Accelerators
Chunlei Men
  • TLE: Compiling End-to-End LLM Inference into Triton Megakernels with E-Graphs
You Zhou
  • SGLang-Plugin-FL: A Unified Multi-Chip Backend for SGLang Inference
Andrew Espira
  • What Should SRE Watch? Calibration-Gated Decisions and the Missing Golden Signals for GPU Clusters
Jeffrey Wan
  • New Activation Memory APIs in PyTorch
Shangdi Yu
  • Regional Inductor: Selective Inductor Compilation within torch.compile
Guilherme Leobas
  • More Than Meets the Eyes: Mirroring CPython's Object Protocol in Dynamo
Goutham Annem
  • Warm Standby, Cold Start: Multi-Region Resilience for PyTorch Inference on Kubernetes
Hyungyo Kim, Apoorve Mohan
  • Train More with Less: Fewer GPUs, Less Data Movement, Better Cluster Utilization
Kushal Mittal, Chun Tao
  • Choosing BF16, FP8, or TurboQuant: A PyTorch KV Cache Guide for Memory‑Constrained GPUs
Shangdi Yu, Boyuan Feng
  • Private Pool Memory Visualization for CUDA Graphs
Pranav Prashant Thombre US
  • Efficient, Large-Scale LoRA and Fine-Tuning for Diffusion Models in PyTorch
  • Beyond Autoregression: Fine-Tuning Diffusion LLMs at Scale in PyTorch
Soham Ghadge
  • Incremental Neighbor Sampling for GNN Training on Continuously Growing Graphs
Sarath Nandu Ramachandran Nair, Fayçal Benmlih
  • Power-Aware Profiling: Bringing Arm Telemetry into PyTorch for Efficient Cloud–Edge Inference
Kazuaki Ishizaki
  • Real Shapes, Real Bugs: Model-Derived Operator Tests for Quick Out-of-Tree Accelerator Bring-Up
Baris Demir
  • Bridging PyTorch Export to Edge GPU Backends: Enabling High-Performance grid_sample with ExecuTorch,
Zhaopeng Qiu, Jingqi Zhang CN
  • FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning
Aditya Tewari
  • High-Performance CPU Inference in PyTorch: TorchInductor, oneDNN, and vLLM Ecosystem Integration
Krzysztof Laskowski, Adam Banas
  • From FX Graph to Silicon: Building a Full-Stack PyTorch Backend for Custom AI Accelerators
Yuchu Fang, Jiahao Chen
  • Production Stability for vLLM: Elastic Scaling and Fault Tolerance Mechanisms on Ascend Platform
Wei Liu
  • FlagQuantum: A PyTorch-Native Runtime for Differentiable and Distributed Quantum Computing
Quentin Anthony, Dhabaleswar K (DK) Panda
  • OSU HPC-AI: Distributed Pre-training, Reinforcement Learning, and Inference with MPI Communication
Syed Ahmed
  • Graphing the Ungraphable: Making MoE Training CUDA-Graphable in PyTorch
Schwinn Saereesitthipitak, Vikram Sharma Mailthody
  • Dynamo Snapshot: Autoscaling and Failure Recovery in Seconds Using GMS
Weifeng Yao, Xiaojun Zhang
  • Heterogeneous CPU + GPU EPD Disaggregation to Boost VLM Serving
Weifeng Yao
  • Better Together: Get More Concurrent Users with Heterogeneous Multi-Agent Workload optimization
Sagar Surendran
  • Beyond the Learner: The CPU’s Role in Scaling PyTorch Multi-Agent Reinforcement Learning
  • Beyond the Learner: The CPU’s Role in Scaling PyTorch Multi-Agent Reinforcement Learning
Sagar Surendran, Na Li
  • Beyond the Learner: The CPU’s Role in Scaling PyTorch Multi-Agent Reinforcement Learning
Abhilash Majumder
  • Inductor-TV: Formal Methods for the Pytorch Compiler
Josiah Davis, Kshitij Sisodia
  • Optimizing Streaming ASR on Arm Edge Devices with ExecuTorch
Samuel Nordmann
  • Elastic Endpoint P2P Transfers for Inference with NIXL
Elena Zhelezina
  • From PT2E to Arm/TOSA: Dynamic W8A8 Linear Quantization in ExecuTorch
  • X-Raying a Delegate: A Debugging Toolkit for Delegated Inference in ExecuTorch
  • Two GPU Delegates, One Vulkan Context: Mixed VGF–Vulkan Execution in ExecuTorch
Thomas Ortner, Daniel Galvez
  • Shape-Stable Dynamic Control Flow in PyTorch CUDA Graphs
Zhaoqiong Zheng
  • Intel Client GPU Expands PyTorch Use Cases: Enabling Robotics AI on iGPU and dGPU
Jianyi Zhang, Eikan Wang
  • Debuggable and Profileable Intel GPU Software Stack for AI Computing with PyTorch
Kevin Fu
  • AOTI with CUDA Graph
Qi Bao, Ugur Kaynar
  • Tensor KV Cache: A PyTorch-Native Approach to KV Cache Offload
Pengzhan Zhao
  • Inside TokenSpeed-Kernel: Portable, High-Performance LLM Inference Kernels Across GPUs
Alessandro Sangiorgi
  • Accelerating Helion Autotuning with Warm-Start and Shared Caches
Yuchuan Gou
  • VLAs for Autonomous Driving: Inference Optimization and Production Lessons
Bruce Chang
  • Resilient PyTorch Distributed with NCCL Shrink and Revoke
Xingguo Li
  • Bringing SmolLM2-135M to Resource-Constrained Arm Edge Devices with ExecuTorch
Zongyu li
  • ASystem: An Open-Source Reinforcement Learning System for Foundation Model Post-Training
Dhritiman Das, Vishal Shah
  • Retrieval as Inference: Building a Production-Scale Torch-Native Retrieval Engine
Jiho Kim
  • Scaling PyTorch Knowledge Through Community: Lessons from PyTorch Korea
Rahul Vishwakarma, Shrey Modi
  • Encryption Hides the Data, Not the Agent: Trace-Privacy for PyTorch LLM Agents
Sahdev Zala
  • 2 in 1: Deep Dive into PyTorch Certified Associate (PTCA) and PyTorch Ambassador Program
Chintan Parikh, Gian Marco Iodice
  • From PyTorch to Edge: An Open-Source Workflow for Converting, and Optimizing On-Device Models
Sanhith Vandara
  • AI Power Efficiency Index (APEI): Framework for measuring Per-Token CO2 Emissions in Frontier LLMs
Tharun Adithya Srikrishnan
  • TraceLens Agent: Automated Performance Bottleneck Identification from PyTorch Profiler Traces
Ravi Gupta, Rishi Teja Madduri
  • Efficient Inference of Masked Diffusion Language Models on AMD XDNA2 NPU
Karan Jain
  • Multimodal Sensor Fusion for AI-Powered Robotics: Advancing Real-World Perception
Aswin Tony Kanikairaj
  • Human-in-the-Loop Requirement Classification for Enterprise API Workflow Generation with PyTorch
Duane Edgington
  • Perch 2.0 and Perch-Hoplite in Pure PyTorch: An End-to-End TensorFlow-Free Bioacoustics Pipeline
Nan Zhu
  • Scaling Multi-Model PyTorch Workloads for Closed-Loop Robotics Evaluation
Alvaro Moran, Jingya Huang
  • TorchTPU x Hugging Face: Native TPU Acceleration Across the HF Ecosystem
Frederik Gossen
  • Piecewise CUDA Graphs: Eager Breaks in CUDA Graph Capture for PyTorch LLM Inference
Kaustubh Tangsali
  • Don't Just Trust the Model, Test the Physics: Evaluating PyTorch Models with PhysicsNeMo-CFD
Kareem Khaled Fareed
  • cuxray: Optimize Kernels Without a GPU
Matthias Jouanneaux
  • Exploiting CUDA Locality Domains in PyTorch on Blackwell and Beyond
Bui Quang Trinh
  • Dynamic Geometric Computing: Stress-Gated Visible Feature Layers for Hidden-Regime Time Series
Abhilekh Verma
  • LLMs & Natural Language Processing – Unlocking the Future of AI Communication
  • Building Global Allies: How Male Mentors Can Accelerate Women in AI & Startups
Rini Susan V S
  • Your GPU Is Starving: The Hidden Cost of Loading LLMs at the Edge
Amit Samson Patole
  • Attested Inference: a signed, independently-verified receipt for every model output
Nishant Verma
  • AI-Driven Supply Chain Intelligence with PyTorch-Based Predictive Models
Qidong Zhao, Xiaoming Hu
  • Demystifying Nondeterministic PyTorch Model Edge Inference with the Google AI Edge Debugger
Chirabrata Senapati
  • Adaptive Consensus for Low-Latency Distributed Systems
  • Continuous Causal Inference for Real-Time Advertising Measurement
Vishal Shah, Ronak Kaoshik
  • Quantize-Then-Refine: Two-Stage Scoring for Memory-Efficient GPU Retrieval
Jeffrey Mahou
  • From Scheduled to Real Overlap: All-Gather/Reduce-Scatter in PyTorch Distributed
Andrew Bond
  • Moral Tensors and DecisionProofs: Compiling Language into an Auditable, Grounded Safety Layer
  • Atlas + Erebus: A Personal AI Research Platform on Two Prosumer GPUs
Muniker Aragon
  • VLA Models on Intel Arc Pro: Enabling Real-Time Robot Control with PyTorch XPU and OpenVINO
Rahul Unnikrishnan Nair, MinSung Kim
  • Stage-Specialized E-PD Disaggregation for SLO-Aware Multimodal Serving on Intel Arc Pro B50/B70
Isha Narula
  • AI in Analytics
Ivan Potapov
  • From PyTorch to a Phone: What Quantizes Cleanly On-Device, and What Doesn't
Radu Salavat, Nikhil Gupta
  • Importance Aware Attention for Faster PyTorch Inference on Arm CPUs
Haichen Zhang, Huamin Chen
  • Training Embedding Models Resiliently for Multimodal Model Inference Routing
Kartikaya Purohit
  • Real-Time Ranking Systems for E-Commerce Search and Discovery
Ziming Zhou
  • Before the Loss Spike: Training Alignment for Precise PyTorch Debugging at Scale
ceci lv
  • vLLM-plugin-FL -- Multi-Chip vLLM framework
Zhipeng Wang
  • Scaling Large-scale Distributed Training with DeepSpeed using Muon Optimizer
Sanskar Prasad, Aheli Poddar
  • KernelOPT: An Open-Source Multi-Agent Harness for GPU Kernel Optimization
Kaoutar El Maghraoui, Priyanka Naik
  • A Multi-Layer Profiling Toolkit for Out-of-Tree PyTorch Accelerators
Etai Lev Ran
  • REBAR: Routing Extensions for Burst-Aware RL
Abhishek Jain, Ashwin Sekhar
  • Efficient MoE LLM Inference on Arm with vLLM and OpenVINO
Eyal Chocron
  • Maestro: Benchmarking Concurrent GPU Operations as They Run in Real Workloads
Sohail Mohammad
  • A Practical Taxonomy of LLM Inference Bottlenecks: Prefill, Decode, Memory, and Scheduling
N Maajid Khan, Ashish Chopra
  • Improving PyTorch Efficiency on ARM with SVE-Accelerated Memory Primitives
Pranav Saji
  • weights_only Was Supposed to Save You: PyTorch Model RCE in 2026 and Safe Loading
Yidi Wu
  • Escape Hatches for torch.compile
Shuai Yang
  • CUDA Graph on Large Scale Recommender Systems
Roy Allela, Aravind Neelakantan
  • Train Across Your Entire GPU Fleet: Heterogeneous Mixed-Generation Training in Pure PyTorch
Andrew Madson
  • "Powering AI/ML with Python and Apache iceberg"
Paulo Aragao
  • Breaking torch.distributed on Purpose: Chaos Engineering for Large-Scale PyTorch Training
  • AllReduce, AllGather, or AllToAll? How Parallelism Shapes PyTorch Network Traffic
Aleksey Vlasenko
  • TPU Model Performance Auto-optimization
Rutuja Pathade
  • Diagnosing LLM Streaming Bottlenecks: Profiling TTFT, Decode Throughput, and Inference Telemetry
Shashank Agarwal
  • Building Self-Healing Infrastructure for AI Agents
Adit Modi
  • The 11-Minute Cold Start: A Visual Anatomy of Why GPU Pods Take Forever to Become Useful
Manvi Gupta
  • State Persistence in Containers
Linux Foundation Events
  • The health benefits of a cheese based diet