vLLM Comes to Tokyo: Inside Japan's First vLLM Community Night
Event |Apr 24, 2026 |EmbeddedLLM Team |3 min read

vLLM Comes to Tokyo: Inside Japan's First vLLM Community Night

A recap of vLLM Community Night in Tokyo — where core contributors, AWS, Fujitsu, and Shisa.AI took the stage to talk KV cache, compression, and production inference.

vLLM Japan 2026

After a year of building momentum across Singapore, Bangkok, and Kuala Lumpur, the vLLM community's next stop was Japan. On Friday, 24 April 2026, engineers, researchers, and AI builders packed into a venue in Roppongi-Itchome, Tokyo, for vLLM Community Night — a technical evening co-organized by Tokyo AI (TAI), Japan's largest AI community with over 4,000 members, and Embedded LLM, one of the leading contributors to the vLLM project.

If the earlier vLLM gatherings across Southeast Asia were about introducing the community to inference at scale, Tokyo's event went a level deeper. This wasn't an introductory session — it was five talks, back to back, aimed squarely at engineers already running vLLM in production or about to.

Setting the Stage

The evening was hosted by Ilya Kulyatin, founder of Tokyo AI and an AI-native systems entrepreneur, alongside Jiaqi Lim, Head of Communications and Marketing at Embedded LLM. Running from 6:00 PM to 9:00 PM with dinner and networking built in, the format mirrored what's become a signature of vLLM community events: real technical depth, delivered by the people actually building the tools, followed by unstructured time to connect.

The lineup was designed to move from theory to hard-won production reality — starting with a project update, working through the low-level mechanics of KV cache and compression, and closing with a full production post-mortem.

The Five Talks

Intro to vLLM and Project Update — Tun Jian Tan (vLLM Committer, Embedded LLM) The night opened with a state-of-the-project update from a vLLM committer, covering the latest features and where the engine's roadmap is headed — grounding the room in a shared understanding before diving into specifics.

Evolution of the KV Cache in vLLM — Tony Valderrama (Head of Product, Momento) Valderrama traced how the KV cache evolved from a simple in-engine optimization into a distinct, fully distributed system component. The talk covered current state-of-the-art approaches like LMCache and Mooncake, along with a roadmap for teams looking to adopt distributed caching incrementally rather than all at once.

Practical AI Model Compression with OneComp — Yuma Ichikawa (Senior Research Manager, Fujitsu) Ichikawa introduced OneComp, an open-source framework for post-training compression of generative AI models. The session walked through how the framework automates model inspection, mixed-precision planning, and progressive quantization — making compression a repeatable engineering process rather than a one-off tuning exercise.

Distributed Inference with vLLM on AWS — Toshinobu Akazawa (Solutions Architect, AWS) Akazawa covered the infrastructure choices behind large-scale vLLM deployments on AWS, from Amazon SageMaker HyperPod to AWS ParallelCluster, and the role EFA/SRD networking plays in low-latency GPU communication. The core of the talk focused on prefill-decode disaggregated inference — a pattern gaining traction for squeezing more throughput out of distributed serving setups.

vLLM in Production: From Quants to QPS — Leonard Lin (CTO, Shisa.AI) Closing out the night, Lin shared what it actually looks like to run all-Japan production inference across diverse model types and heterogeneous hardware. The talk didn't hold back on the messy parts — benchmarks, evals, quality-versus-performance tradeoffs, and the hardware-specific tricks Shisa.AI uses to tune their serving architecture for real-world query volumes.

Why Tokyo, Why Now

Pairing Tokyo AI's 4,000-plus-strong local community with Embedded LLM's position as a core vLLM contributor meant the event drew exactly the audience it was built for: ML engineers, infrastructure builders, open-source contributors, and product teams already deep in the weeds of serving LLMs at scale. It's a natural extension of the same pattern the vLLM community established across Southeast Asia in 2025 — bring the people who maintain the project together with the people deploying it in production, in the same room, and let the conversation go as technical as it needs to.

With Fujitsu, AWS, and Momento all represented on stage alongside the vLLM core team, Tokyo's event underlined something bigger than a single evening of talks: enterprise infrastructure providers and open-source maintainers are now building the future of inference together — and Japan is very much part of that conversation.

What's Next

vLLM Community Night in Tokyo is another marker in the project's growing footprint across Asia's AI hubs. As more regional teams push distributed inference, KV cache architecture, and model compression into production, expect the conversation — and the community behind it — to keep expanding.

Want in on the next one? Follow our LinkedIn page to be the first to know when the next vLLM community event lands near you.

← Back to all posts Embedded LLM · Apr 24, 2026