Speaker
Description
Running open source models locally or self-hosting is the new moat.
Today, in this rapidly moving AI age, data sovereignty is becoming more important. We don't want some vendor locking us in and getting access to our organisation's critical internal data. Most organisations are moving away from giant providers who provide closed models like a black box, and we have to pay huge bucks every month without even having track of usage and other metrics.
As API pricing keeps moving up and up, a few teams have already moved on to self-hosting LLMs and fine-tuning them in a way that aligns with organisational outcomes. This has led to huge demand for a new type of engineer called an "Inference Engineer." Unlike Infra Engineers, Inference Engineers handle the full deployment and serving lifecycle of LLMs and manage day-2 operations, just as traditional systems used to have.
In this talk, I will unpack how this shift is happening, as open weight models are giving good competition to frontier models, while GPUs are also becoming more accessible. I will also show how projects like KServe from the CNCF ecosystem have removed friction and give your team a level of abstraction.
I will then bring it back to the how by using KServe's deployment modes like Raw Deployment, Serverless, and the newer disaggregated LLM Inference Service as the infrastructure layer that makes running your models practical, not just possible.
I will also cover milestones and recent updates in the KServe project, its production adapters, and how they have evolved. This is not going to be a deep technical talk, but I feel everyone should get familiar with running some SLMs on local machines and playing with them.
If you are new to serving LLMs locally or have thought that it is not ideal to self-host and serve them in production at scale, this talk will give you an overview of how deploying models is different from deploying traditional services with the help of Kubernetes.
I'll pinpoint some of the trade-offs and gotchas that I learned the hard way, and I'll be open for some Q&A.
Session author's bio
Atharva is currently working as Open Source Developer and Maintainer at one of the projects of CNCF OpenEverest.
Worked as Google Summer of Code 2026 Mentee at OpenScienceLabs.
Working in Open Source for over year contributing in various CNCF Projects like Harbor, Kthena, OpenEverest
Loves working with golang, AI inferencing and cloud native ecosystems.
Any other info we should know?
Most AI inference talks assume you're already committed to self-serving and jump straight into the how. This one makes the business and data governance case first. Kubernetes has quietly become the default control plane for LLM inference. KServe sits exactly on top of Kubernetes and makes it understand new terms and practices for handling GPUs and multi-node GPUs at scale.
Engineers already running enterprise solutions that use the same technologies under the hood can learn how they work and go back to convince their managers about self-serving LLMs at scale using the industry-standard and widely trusted CNCF ecosystem.
For attendees, this talk will connect the dots between "Why Local LLMs make sense from a business perspective" and "How Kubernetes and KServe make production LLM serving practical."
| Agree to Privacy Policy and Notice | I agree |
|---|---|
| Level of Difficulty | Intermediate |
| Please confirm that there are included headshots of all speakers in their profiles | Yes |
| Social Media | https://x.com/AtharvaXDevs |
| In Person Attendance | Remote |