Loading…
7-9 October, 2026
Prague, Czechia
View More Details & Registration
Important Note: Timing of sessions and room locations are subject to change.

The Sched app allows you to build your schedule but is not a substitute for your event registration. You must be registered for Open Source Summit Europe 2026 to participate in the sessions. If you have not registered but would like to join us, please go to the event registration page to purchase a registration.



Thursday October 8, 2026 16:40 - 17:20 CEST
As AI workloads move into production in cloud native environments, the open source community must define the future of AI inference infrastructure. This panel convenes experts from leading projects to assess the ecosystem and chart a path forward.

Key themes include:

- The evolution of cloud native inference (from serverless to large‑scale distributed serving)
- Performance and efficiency advances such as KV cache management, quantization and hardware acceleration
- The need for standardization across APIs, model formats and interoperability.

They’ll also examine how projects like vLLM, llm‑d, and KServe can collaborate to reduce fragmentation.

Panelists will address real‑world challenges in Kubernetes deployment, including multi‑tenancy, GPU sharing, autoscaling and observability. There will also explore emerging trends such as edge inference, multi‑modal models and the convergence of training and inference systems.
Speakers
avatar for Martin Hickey

Martin Hickey

Senior Technical Staff Member, IBM Research
Martin Hickey is a STSM at IBM Research, focused on Open Source, Cloud Native Computing, and AI. Martin has notable contributions to open source projects like vLLM, LMCache, Kubernetes, Helm, OpenTelemetry and OpenStack. Martin is a core maintainer for LMCache and an emeritus core... Read More →
avatar for Maroon Ayoub

Maroon Ayoub

Senior Principal Machine Learning Engineer, Red Hat
Maroon Ayoub is an AI infrastructure architect at Red Hat focused on distributed inference. He co-leads development of llm-d and specializes in scaling LLM inference with Kubernetes-native architectures, KV-cache orchestration, and open source integrations.
avatar for Nili Guy

Nili Guy

R&D and Senior Technical Staff Member, IBM
Nili is a Research Manager and Senior Technical Staff Member at IBM Research, co-creator of llm-d, and an expert in distributed inference and Kubernetes-native AI systems. She has led key open-source and productized inference initiatives across IBM’s AI platforms.
avatar for Cong Liu

Cong Liu

Software Engineer, Google
Cong is a software engineering working on Google Kubernetes Engine (GKE), building solutions for AI Inference workloads, such as efficient request scheduling for Large Language Models (LLMs). His past experience include GKE release channels and safe cluster upgrades.
avatar for hyunkyun moon

hyunkyun moon

ML Platform Engineer, Moreh
Hyunkyun Moon is an ML Platform Engineer at Moreh, where he builds high-performance LLM inference platforms using llm-d. He is an active contributor to open-source projects, including llm-d and vLLM. Previously at LINE Plus, he specialized in developing large-scale DBaaS using Kubernetes... Read More →
Thursday October 8, 2026 16:40 - 17:20 CEST
Forum Hall (Floor 2)
  Open AI & Data
  • Audience Experience Level Any

Sign up or log in to save this to your schedule, view media, leave feedback and see who's attending!

Share Modal

Share this link via

Or copy link