B2BangaloreB2Bangalore
HomeEventsFeedMarketplaceSponsorshipsSwagSaved
Log inSign up
  1. Home
  2. /Events
  3. /vLLM Inference Meetup Bengaluru
vLLM Inference Meetup Bengaluru cover image

vLLM Inference Meetup Bengaluru

SEP19
19 Sept 2026, 1:00 – 6:30 PM
1 – 6 PM
Red Hat India Pvt. Ltd. · Doddanekkundi, India

Community Experience

Be the first to share your experience from this event.

B2BangaloreB2Bangalore

Bangalore's B2B event intelligence platform. Find every flagship B2B event, see which target accounts will be there, and get instant registration links — built for India.

Get it onGoogle PlayEvent calendar forChrome

Product

  • Explore events
  • App home
  • Venue marketplace
  • List your venue
  • Sponsor marketplace
  • List your event
  • Platform
  • Companies
  • Saved events
  • Target accounts
  • MCP for AI tools

Company

  • Pricing
  • Newsletter
  • For organisers
  • Careers
  • About
  • Blog
  • WhatsApp community
  • Google Business Profile
  • [email protected]

Legal

  • Terms of Service
  • Privacy Policy
  • Refund & Cancellation

Browse Bangalore events by category

  • AI Events in Bangalore
  • Machine Learning Events in Bangalore
  • Data & Analytics Events in Bangalore
  • DevOps Events in Bangalore
  • Cloud Computing Events in Bangalore
  • Cybersecurity Events in Bangalore
  • Blockchain & Web3 Events in Bangalore
  • Developer Events in Bangalore
  • Product Management Events in Bangalore
  • Design & UX Events in Bangalore
  • Startup Events in Bangalore
  • Networking Events in Bangalore
  • Fintech Events in Bangalore
  • Marketing & Growth Events in Bangalore
  • SaaS Events in Bangalore
  • Women in Tech Events in Bangalore
  • Hackathons in Bangalore
  • Tech Workshops in Bangalore
  • Tech Conferences in Bangalore
  • Founder Dinners in Bangalore
  • Weekend Tech Events in Bangalore
  • Free Tech Events in Bangalore
  • Google Events in Bangalore
  • AWS Events in Bangalore
  • Microsoft & Azure Events in Bangalore
  • Business Events in Bangalore
  • Tech Events in Bangalore

© 2026 B2Bangalore · Bengaluru, Karnataka · India

[email protected]

About this event

Join us for a deep dive into the engine room of vLLM and llm-d AI inferencing, where we will focus on the architecture, optimizations, and raw engineering required to run inference at scale.

​Whether you’re looking to squeeze every last token out of your GPU cluster or you're curious about the latest commits to the vLLM and llm-d ecosystems, this is the room you want to be in.

​​What to Expect ​Deep Technical Sessions: Hear directly from the maintainers and core committers of vLLM and llm-d

​Scale in Production: Learn from industry leaders about deploying LLMs in production

​Hands-on learning & demos

​Networking: Stick around for food and soft drinks. It’s a great chance to chat with the speakers and exchange ideas with fellow developers and engineers.

​​Who Should Attend ​vLLM and llm-d users and contributors

​ML and infra engineers working on inference and serving

​Platform teams running GenAI in production

​Anyone curious about efficient inference across local, cloud, and Kubernetes

​​Event Agenda ​1:00 PM – 1:30 PM | Registration & Welcome Registration, networking, and welcome refreshments.

2:00 PM – 2:30 PM | vLLM & llm-d Updates Prasad Mukhedkar, Red Hat

​An overview of the latest updates and developments across vLLM and llm-d.

​2:30 PM – 3:00 PM | llm-d for Sovereign & Agentic Workloads Pravein Govindan Kannan, IBM

​An overview of llm-d’s distributed inference stack, along with ongoing optimization and benchmarking efforts for sovereign and agentic workloads.

​3:00 PM – 3:30 PM | Distributed Inference on ROCm with WideEP on vLLM & llm-d by Chaitanya Sri Krishna Lolla & Sirra Ajith, AMD

​Explore distributed inference on AMD ROCm using WideEP with vLLM and llm-d.

​3:30 PM – 4:00 PM | Break & Networking

​4:00 PM – 4:30 PM | Inside vLLM Semantic Router: Intelligent Routing for LLM Inference Aayush Saini, Red Hat

​Explore how vLLM Semantic Router uses intelligent, cache-aware routing based on request semantics, model capabilities, workload characteristics, and infrastructure constraints to optimize inference.

​4:30 PM – 5:00 PM | Scaling Inference at NxtGen Using the vLLM Ecosystem Abhishek Kumar, NxtGen

​Learn how NxtGen is leveraging the vLLM ecosystem to scale LLM inference workloads.

​5:00 PM – 5:15 PM | Break

​5:15 PM – 6:15 PM | Hands-on Workshop: vLLM Inference with AMD GPUs - Get hands-on experience running vLLM inference on AMD GPUs using open-source ROCm platforms.

​​ Important information ​Registration closes 24 hours before the event. We cannot admit unregistered attendees.

​Bring your laptop with SSH installed (GPU instances provided by the organizers)

​Please bring a photo ID to verify your registration on arrival.

Register

Free event

Event details

Format
meetup
Address
10th Floor East, NEUBAR, Carina Building, Ferns City, Doddanekkundi, Bengaluru, Karnataka 560037, India
Tickets
Free