Pratik Parmar

Pratik Parmar

Developer Relations Professional who's passionate about helping developer communities across the globe!

Bengaluru, India

Actions

Pratik is a code whisperer and is passionate about tinkering with technologies. Currently, he'd working on the unique intersection of AI engineering and developer advocacy. When he's not helping the developer communities, he can be found roaming around in the mountains.

Area of Expertise

  • Information & Communications Technology

Topics

  • Machine Learning
  • Artificial Inteligence
  • python
  • TensorFlow
  • scrapy
  • Community Building
  • Community Engagement

Sovereign AI: Building Air-Gapped Intelligence That Scales

Organizations handling sensitive data face a critical challenge: deploying cutting-edge AI while maintaining absolute data sovereignty and regulatory compliance. Healthcare providers, financial institutions, government agencies, and enterprises all need AI capabilities that never expose their data beyond their infrastructure perimeter.
This session presents a practical framework for implementing sovereign AI through air-gapped deployments. We'll explore architectural decisions, fine-tuning strategies, and operational considerations for running powerful AI workloads entirely within controlled environments.

Attendees will gain a concrete methodology for evaluating, architecting, and deploying sovereign AI solutions that meet stringent security requirements while delivering production-grade capabilities.

Scaling LLM Inference for Production Traffic

Most teams start with `vllm serve model-name`, but is that actually enough for production when the traffic burst is unpredictable, the GPU fills up with the KV cache, and 100s of users are wondering why it is taking so long to get a response?

This talk takes you on a ride from a basic vLLM server to a scalable inference stack. Rather than a feature showcase, we'd rather take a problem at every stage, apply optimizations, and benchmark it to check if it actually helps.

Finally, to make it scalable and production-ready, we'd add KEDA to scale replicas on meaningful demand signals, not just GPU usage. Lastly, we’d talk about the harsh reality of inference: model-loading cold starts, warm capacity, and safe scale-down.

Pratik Parmar

Developer Relations Professional who's passionate about helping developer communities across the globe!

Bengaluru, India

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top