Dense versus MoE: what "fits" meant
In a dense model, every parameter takes part in computing every token. The old rule of thumb followed from that: the whole model had to sit in VRAM, or you pushed whole layers out to system RAM and paid for it in speed. Pande's own sizing examples show the squeeze. A 7B model at Q4 quantization needs roughly 4–5.5GB of VRAM, a 12GB card tops out around 12B to 15B models, and past about 18B the average GPU runs out of room.
An MoE model swaps the big feed-forward block in each layer for many parallel "expert" networks. A router sends each token to a small subset of them. The rest sit idle for that token.
That gives you two different numbers for the same model:
- Total parameters: what you have to store somewhere, in VRAM, system RAM or both.
- Active parameters: what actually gets read and computed per token.
The model names follow this convention. Qwen3.6-35B-A3B and Gemma-4-26B-A4B are named for their total size and (roughly) their active size. The "A3B" and "A4B" suffixes describe compute per token, not memory footprint. A model with 3B active parameters still needs room for all 35B.
How llama.cpp splits the work
The llama.cpp server documentation lists options that control exactly this placement:
--cpu-moekeeps all MoE weights in the CPU's memory.--n-cpu-moe Nkeeps the MoE weights of the first N layers on the CPU.- The current documentation also lists
--moe-cache-mib N, a GPU cache for experts kept in system RAM. It is disabled by default. --gpu-layers(-ngl) sets how many layers go to VRAM.
The logic, as several community guides describe it, is to keep the parts touched every token (attention, embeddings, KV cache) on the GPU, and let the large expert stacks live in system RAM. That differs from the old approach of lowering -ngl, which pushed whole layers to the CPU, attention included. The MoE flags split the layer itself.
In practice, the tuning rule is simple. A lower --n-cpu-moe keeps more experts on the GPU, which is faster but uses more VRAM. You aim for the lowest value that still loads without running out of memory.
The reported results, and what they do and don't show
Pande reports two setups:
| GPU | VRAM | System RAM | Model | Reported speed |
|---|---|---|---|---|
| RTX 3080 Ti | 12GB | 32GB | Qwen3.6-35B-A3B | about 24 tokens/s |
| GTX 1080 | 8GB | about 32GB DDR4 | Gemma-4-26B-A4B | about 14 tokens/s |
He uses the first as a VS Code Copilot replacement. The second backs self-hosted tools such as Open Notebook, Blinko and Paperless-GPT.
Treat these as one person's anecdotes, not benchmarks. The article doesn't give the quantization, context length, full command line or test method. All of those change tokens-per-second. A GTX 1080 and an RTX 3080 Ti are very different cards, so the numbers also say little about how any given PC will behave.
Check the flag before you copy it
The article says the 3080 Ti result came from setting --no-cpu-moe "to 25 layers." The llama.cpp documentation we could check lists --cpu-moe and --n-cpu-moe N, which keeps expert weights from the first N layers in system RAM. It doesn't list a layer-count form of --no-cpu-moe. The author may have used a different build, mistyped, or meant --n-cpu-moe. We can't tell which. Check --help on your own build instead of pasting the fragment.
"No speed penalty" is too strong
The article's claim that the split carries no significant speed penalty needs qualifying. A research paper on consumer-GPU MoE inference (MoBiLE, accepted to ASP-DAC 2026) says the approach of keeping active experts in GPU memory and inactive ones in CPU memory is fundamentally constrained by the limited bandwidth of the CPU-GPU interconnect. Its authors propose a technique to ease that and report 1.60x to 1.72x speedups over their baseline.
One community analysis goes further. It argues that --n-cpu-moe normally makes MoE inference slower, and the dramatic speedups only appear when the model had overcommitted VRAM. Its reasoning is that each layer pushed to the CPU means its expert weights now cross PCIe from system RAM. That is one blogger's reading of public benchmark data, so weigh it accordingly. It does fit the physics.
The practical way to square the two views: offloading beats a model that can't run at all, or that spills out of VRAM badly. It doesn't beat a model that fits entirely on the GPU. One guide also warns that on Windows, overcommitting VRAM can lead to driver paging and catastrophic throughput. That matters to many readers here.
Tuning has cliffs, too. A guide reports that for one RTX 5090 owner, throughput climbed as the offload setting dropped, then fell sharply one step further. The advice is to sweep values rather than copy someone else's.
What this means for buying advice
Old advice said to match VRAM to model size. For MoE, a better checklist looks like this:
- Total weights. Quantized file size has to fit across VRAM plus RAM, with headroom for the operating system and other apps.
- VRAM for the always-on parts. Attention, KV cache and some experts need to fit on the GPU. Context length drives KV cache growth, so test at the context size you'll really use.
- System RAM capacity and speed. Offloaded experts are read from RAM every token, so RAM bandwidth and the PCIe link matter more than they used to.
- Quality. Heavier quantization to make things fit costs accuracy, as Pande himself notes for dense models.
A practical trial routine
- Note the exact model file, quantization and llama.cpp version.
- Start with all layers on the GPU, then use
--n-cpu-moeto move experts out until the model loads. - Lower the value step by step while watching GPU memory and system memory.
- Record tokens per second at your real context size.
- If it's too slow, try a smaller model or different quantization before assuming bigger is better.
Bottom line
MoE gives older cards a real second act. A 12GB or even 8GB GPU paired with 32GB of RAM can run models that would be impractical as dense ones, which is useful for home-lab assistants and coding helpers. It doesn't make total model size irrelevant, and it doesn't promise any particular token rate. Treat Pande's numbers as a proof of possibility, not a spec sheet.
References
- MoE models changed what "fits on my GPU" means, and most hardware advice hasn't caught up XDA · 2026-10-10T11:30:17+00:00
- llama.cpp --n-cpu-moe: Run a Big MoE Model on a Small GPU (2026 Guide) · Aliteq aliteq.com
- README.md github.com