Benchmarks, ports and production write-ups from the team shipping vLLM on AMD ROCm, Intel Arc and AMD Instinct silicon — plus the JamAI Base product log. Everything we learn in the field, published in full.
A recap of AI Malaysia Takeover 2026 — governance conversations at Pasar AI, ministerial visits to the booth, and a first look at the Green AI Grid.
Event · Aug 11, 2026 A recap of the vLLM Taipei Meet Up — and the bigger announcement that followed: a new TAIONE x Embedded LLM Fellowship Track for Taiwan's vLLM community.
3 min read · EmbeddedLLM Team
Event · Apr 24, 2026 A recap of vLLM Community Night in Tokyo — where core contributors, AWS, Fujitsu, and Shisa.AI took the stage to talk KV cache, compression, and production inference.
3 min read · EmbeddedLLM Team
Partner · vLLM · Feb 27, 2026 A deep dive into vLLM's 7 attention backends for AMD ROCm, featuring 1.2-4.4x throughput gains with ROCM_AITER_FA and MLA optimizations.
1 min read · Embedded LLM & AMD
vLLM · Feb 14, 2026 A comprehensive benchmark analysis of Intel Arc Pro B60 for LLM inference workloads, comparing vLLM and LLM-Scaler performance across different scenarios.
4 min read · EmbeddedLLM Team
Partner · AMD · Jan 21, 2026 Discover how vLLM 0.9.x and AITER integration deliver up to 67% throughput gains for Llama 4 and deep optimizations for DeepSeek architectures on AMD hardware.
External ↗ · EmbeddedLLM Team & AMD
Partner · AMD · Jan 2, 2026 Unlock up to 45% higher throughput for multimodal models. Discover how vLLM's batch-level DP eliminates vision encoder bottlenecks on AMD ROCm.
External ↗ · EmbeddedLLM Team & AMD
Event · Dec 20, 2025 A recap of vLLM's four flagship 2025 community events across Singapore, Bangkok, and Malaysia — and how they built Asia's AI inference movement
5 min read · EmbeddedLLM Team
Partner · AMD · Nov 24, 2025 The definitive guide to vLLM parallelism for MoE models. Discover exactly when to use Tensor, Data, or Expert Parallelism to scale LLMs on ROCm.
External ↗ · EmbeddedLLM Team & AMD
Partner · vLLM · Oct 26, 2025 Stop choosing between 2x VRAM costs and 100-second cold starts, free up 90% of your VRAM in seconds. Learn how vLLM Sleep Mode enables 18-200x faster model switching without the cold-start penalty.
2 min read · Embedded LLM
Product · Sep 3, 2025 Unleash truly agentic AI with executable Python workflows, custom RAG-powered agents, and a supercharged PostgreSQL backend.
3 min read · EmbeddedLLM Team
Partner · AMD · Jun 28, 2025 Discover how vLLM 0.9.x and AITER integration deliver up to 67% throughput gains for Llama 4 and deep optimizations for DeepSeek architectures on AMD hardware.
External ↗ · EmbeddedLLM Team & AMD
Hackathon · Jun 23, 2025 How a smart agriculture app built by students and powered by JamAI Base is reimagining the future of food security.
4 min read · EmbeddedLLM Team
Partner · vLLM · Feb 24, 2025 When deploying Large Language Models at scale, infrastructure teams are usually forced into a painful compromise: run in full precision (BF16) and burn through expensive VRAM, or use standard FP8 quantization and watch your model's reasoning capabilities degrade.
1 min read · EmbeddedLLM Team & AMD
Product · Jan 6, 2025 The journey of JamAI Base towards CPU-powered embedding models highlights a crucial shift in the AI landscape. By harnessing the power of Intel Xeon CPUs and OpenVINO, JamAI Base delivers a compelling combination of performance, efficiency, and cost-effectiveness. This approach democratizes access to powerful AI capabilities, making it easier for organizations of all sizes to leverage AI for transformative outcomes
13 min read · EmbeddedLLM Team
vLLM · Dec 1, 2024 This guide shows the impact of Liger-Kernels Training Kernels on AMD MI300X. The build has been verified for ROCm 6.2.
2 min read · EmbeddedLLM Team
vLLM · Nov 5, 2024 This guide shows the impact of Liger-Kernels Training Kernels on AMD MI300X. The build has been verified for ROCm 6.2.
8 min read · EmbeddedLLM Team
vLLM · Oct 28, 2024 This blog post shows you how to run Meta's powerful Llama 3.2-90B-Vision-Instruct model on an AMD MI300X GPU using vLLM. We provide the Docker commands, code snippets, and a video demo to help you get started with image-based prompts and experience impressive performance
5 min read · EmbeddedLLM Team
Partner · vLLM · Oct 23, 2024 Unlock up to 1.8x higher throughput and 5.1x faster TTFT on AMD MI300X. Discover the 8 optimized vLLM settings for serving Llama 3.1.
1 min read · Embedded LLM
vLLM · Oct 11, 2024 This guide walks you through the process of building vLLM from source on AMD MI300X. The build has been verified for ROCm 6.2.
7 min read · EmbeddedLLM Team
vLLM · Oct 27, 2023 EmbeddedLLM has ported vLLM to ROCm 5.6, and we are excited to report that LLM inference has achieved parity with Nvidia A100 using AMD MI210.
7 min read · EmbeddedLLM Team