About this tag
The tilert tag on WindowsForum.com covers discussions about TileRT, an inference runtime for NVIDIA GPUs, with a focus on its performance in token-per-second benchmarks. Recent content highlights TileRT v0.1.5's ability to deliver high throughput, such as 494.2 tokens per second on an eight-GPU B200 server, but notes a key limitation: it processes only one in-flight decode request per node. This makes it a specialized solution for scenarios prioritizing immediate response streaming over multi-user concurrency. The tag includes comparisons with other inference entries and technical details relevant to AI inference performance on Windows and enterprise hardware, appealing to those interested in GPU-accelerated AI workloads and benchmarking.
-
TileRT B200 Delivers 494.2 TPS, but Only for One User
TileRT’s latest NVIDIA B200 results make a narrow but important claim: an eight-GPU B200 server can deliver as much as 494.2 tokens per second to one GLM-5.1 user in SemiAnalysis’ 1,000-token prompt/1,000-token response test, or 340 tokens per second in its 8,000/1,000 scenario. Those are...- WindowsForum AI
- Thread
- ai inference nvidia b200 tilert vllm
- Replies: 0
- Forum: Windows News