Klaus Ma
Principal SW. Eng at Nvidia
Beijing, China
Actions
Team leader, system architect, designer, software developer with 10+ years of experience across a variety of industries andtechnology bases, including cloud computing, machine learning, bigdata and financial services.
Founding Volcano & Flame; emeritus leader of Kubernetes SIG-Scheduling , CNCF TAG-Runtime and CNCF WG-Batch. Global Team Lead of IBM Spectrum Symphony CE & L3, Chief Architect of Batch Container at Huawei Cloud. Currently, Architect and R&D at Nvidia Technologies.
Area of Expertise
Topics
How to manage xPU automatically in large scale
With the development of the AI, more and more applications, e.g. LLM, asks for a better solution of network and storage to work together with GPU. xPU, including both DPU and IPU, is one of well-known solution which provides offload, hardware acceleration and so on. But an AI server usually includes serveral devices to maximize the performance, e.g. 8 DPU + 10 xPU, which introduces the challenge to manage all of those devices. Thanks to kubernetes, the GPU ares managed will by device plugin and operator; but xPU management is more complex which Including provisioning, storage and network.
This session will demonstrate how to manage those xPU in a large scale by some use cases and share the pain-point and solution accordingly; this session will also include part of the roadmap of the system, and general tips for xPU.
Building a Multi-Agent System with Flame and MCP, e.g., Simple Research Agent
As intelligent agents become increasingly widespread, a host of advanced technologies—such as Agentic RL and MCP—are being employed to enhance their capabilities. Throughout our work building agentic systems, we have identified several key challenges and build Flame to address them:
- Session-based isolation
- MCPs and agents discovery
- Session-level context of user for RL
- On-demand parallelism for multi-path reasoning
In this presentation, we showcase the "Simple Research Agent" (SRA) as a concrete example of how Flame and MCP solve these challenges. SRA is a multi-agent system that generates concise research reports. The supervisor agent formulates a plan based on the user prompt; the collector agent analyzes the prompt to gather additional context using sources like DuckDuckGo and web crawlers; finally, the writer agent employs large language models (LLMs) to synthesize and refine the report for the user.
Flame: a distributed engine for AI and Quant
With the development of AI and Quant, more and more elastic jobs were introduced into those areas with performance and security requirements, e.g. matrix multiplication, Monte Carlo. As an elastic job, there may be thousands of tasks which did not depends on each other, e g. matrix multiplication; so, any task can re-run/re-try at any time; the tasks may share dataset by cache or distributed filesystem. Considering the number of tasks, the performance, throughput and resource utilization is important; and security is also important for a multi-tenant platform.
For those scenarios, a distributed engine, named Flame, was introduced. It includes several features and enhancement, e.g. fair-share, preemption, pull-model, session/tasks, for performance, throughput.
This session will present the architect and features of Flame; it'll also demonstrate the improvement by matrix multiplication, Monte Carlo and so on.
Kubernetes Community Panel: A Decade of Evolution and Future Trends
Join us in celebrating the 10th anniversary of Kubernetes with a panel featuring some of the community's most influential contributors and maintainers from China. Over the past decade, Kubernetes has grown to the cornerstone of cloud-native infra, thanks to the dedication and innovation of its community members. In this panel, we will talk about our journeys with Kubernetes, share stories and experience, and discuss the future of Kubernetes in the next decade. Our panelists include current and previous owners, tech leads and maintainers. Feel free to join the panel to share your perspectives on the past and next decade of the Kubernetes community and ask anything about the community.
CNCF BSI-WG intro
Cloud Native Batch System Initiative Working Group is a group formed by passionate batch system maintainers looking to support batch workload in cloud native environments. In the passed year, there're several topics was discussed in this WG, e.g. batch landscape; we'd like to introduce our work into the community for the developer who are also interesting in batch system.
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top