Session

Fine-Tuning Frontier Models on Apple Silicon: A Practical Guide to MLX-Powered Local LLMs

Abstract: As organizations scramble to deploy generative AI at scale, the economics and privacy implications of cloud-based inference are becoming untenable for many use cases. This talk presents a comprehensive, hands-on exploration of fine-tuning and deploying frontier-class language models on Apple Silicon hardware using the MLX framework — an approach that delivers production-grade performance at a fraction of cloud costs while keeping data on-device. Drawing on real-world deployment experience, we'll walk through the full pipeline: from model selection and quantization strategies, through LoRA/QLoRA fine-tuning workflows, to production-ready inference optimization. We'll cover practical benchmarks comparing MLX against traditional frameworks, discuss the trade-offs between model size, quantization level, and task performance, and share lessons learned from deploying fine-tuned models in enterprise environments. The talk will include live demonstrations of model loading, fine-tuning, and inference on M-series chips, along with actionable recommendations for teams evaluating on-device AI deployment. Whether you're an ML engineer looking to reduce inference costs, a solutions architect designing privacy-first AI systems, or a researcher exploring efficient model adaptation, this session provides the technical depth and practical guidance needed to get started with MLX-powered local AI.

Key Takeaways:
•⁠ ⁠MLX delivers competitive inference performance on Apple Silicon with dramatically lower costs than cloud alternatives
•⁠ ⁠QLoRA fine-tuning on M-series chips makes frontier model adaptation accessible without GPU clusters
•⁠ ⁠On-device AI enables privacy-preserving deployments that cloud inference simply cannot offer
•⁠ ⁠Production deployment patterns for MLX include model quantization, batch optimization, and memory management strategies
•⁠ ⁠The enterprise case for local AI hinges on cost reduction, data sovereignty, and latency improvements

Session format
Breakout sessions (45–60 minutes)

Track
Operations & Internal AI

Level
Introductory and overview

Business Function Tags
Operations

Regarding the strategic relevance of this session, particularly given its focus on local deployments rather than traditional cloud infrastructure. The presentation is designed to contrast the high cost of cloud-only LLM utilization with the efficiency of edge computing.

In particular, the session will explore how to architect federated inferencing pipelines—demonstrating how organizations can strategically divide workloads, running heavier components in the cloud while routing targeted tasks to local edge devices. Furthermore, this talk will address how to ensure local inferencing is optimized for peak performance on modern hardware, allowing attendees to walk away with a clear blueprint for building balanced, cost-effective, and highly performant hybrid AI systems.

Learning Objective 1:
Implement Cost-Effective Local Inference: Attendees will learn how to evaluate and deploy frontier-class language models on Apple Silicon hardware, enabling production-grade AI performance while significantly reducing dependence on expensive cloud-based infrastructure.

Learning Objective 2:
Master On-Device Fine-Tuning Workflows: Participants will gain hands-on knowledge of the end-to-end MLX pipeline, including model selection, quantization strategies, and LoRA/QLoRA techniques, allowing them to adapt high-performance models to specific business tasks without the need for costly GPU clusters.

Learning Objective 3:
Design Privacy-First AI Architectures: Attendees will understand how to build and maintain data sovereignty by keeping sensitive information on-device, effectively addressing security and privacy concerns while improving system latency for enterprise-grade applications.



I believe this edge-to-cloud bridge will offer a unique and highly practical perspective for cloud professionals looking to optimize their generative AI budgets and architectures.

Note: The views and research presented during this session are my own and do not necessarily reflect the views or positions of AT&T.

Anjaneya Sastry Kappagantu

Expert Solution Architect, AT&T

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top