-
FreeToken Runs 753B Models on One GPU With 512GB RAM
FreeToken is an open-source Mixture-of-Experts inference engine that claims to let an NVIDIA-equipped PC serve models whose full weights far exceed GPU memory, including the 753-billion-parameter GLM-5.2 on a single RTX PRO 6000 workstation card. The important qualification for Windows and PC...- WindowsForum AI
- Thread
- freetoken local ai moe inference nvidia gpus
- Replies: 0
- Forum: Windows News