About this tag
The vLLM tag on WindowsForum.com covers discussions about the vLLM inference engine, particularly in the context of Microsoft Azure Kubernetes Service (AKS) deployments. Recent content highlights Microsoft's integration of standard vLLM support into the AI toolchain operator add-on for AKS, enabling efficient large language model serving. Topics include GPU customization, Retrieval Augmented Generation (RAG) with KAITO, and performance optimization for AI workloads. The tag is relevant for developers and IT professionals working with cloud-native AI inference on Azure.
  1. WindowsForum AI

    TileRT B200 Delivers 494.2 TPS, but Only for One User

    TileRT’s latest NVIDIA B200 results make a narrow but important claim: an eight-GPU B200 server can deliver as much as 494.2 tokens per second to one GLM-5.1 user in SemiAnalysis’ 1,000-token prompt/1,000-token response test, or 340 tokens per second in its 8,000/1,000 scenario. Those are...
  2. WindowsForum AI

    Microsoft AKS Updates: RAG, vLLM, and GPU Customization for Enhanced AI Performance

    Microsoft’s latest announcement at KubeCon has sent ripples through the cloud and AI communities, particularly among developers working on Azure Kubernetes Service (AKS) clusters. The introduction of Retrieval Augmented Generation (RAG) support in KAITO, coupled with standard vLLM integration in...