Why 8GB cards were stuck at 4–9B
Larger models usually spill into system RAM on an 8GB card, which causes a steep drop in generation speed. The author says a decent 4-bit build of Qwen3.6 27B needs more than double his VRAM. PrismML's own comparison agrees. Its reference Qwen3.6-27B Q4_K_XL build is listed at 17.6GB for weights alone, before any context memory is added.
That is why models in the 4–9B range have been the practical ceiling for this class of GPU.
What Bonsai 27B is
Bonsai 27B comes from PrismML, a startup that emerged from Caltech researchers. It is built on Qwen3.6 27B, so it is still a Qwen model, only heavily compressed. PrismML announced it on July 14, 2026. It is a multimodal model that takes images as well as text, and it is released under the Apache 2.0 license.
The 1-bit variant stores each weight as −1 or +1, with one shared scale per group of 128 weights. That works out to about 1.125 bits per weight. A conventional 16-bit version of the 27.8B-parameter model would need roughly 54GB, which is the vendor's figure. PrismML puts the 1-bit variant at about 3.9GB.
A few details matter here:
- The low-bit storage covers the language network: embeddings, attention, MLP projections and the output head.
- The vision tower is separate and uses 4-bit HQQ. Its optional projector adds about 0.63GB, and it is only needed for image input.
- The author's LM Studio download was 4.4GB with the vision files bundled. That is a package size, not the memory used while running.
Why a small file doesn't guarantee a fit
LM Studio's own model page says the headline footprints describe the language-model weights only. Actual runtime memory also includes the KV cache, activations, runtime buffers and the vision tower, and it grows with context length.
PrismML's model card gives measured peak memory for the llama.cpp Q1_0 build with no KV-cache compression:
| Context | Peak memory |
|---|---|
| 4K tokens | 5.2GB |
| 10K tokens | 5.6GB |
| 100K tokens | 11.6GB |
The same card says a 4-bit KV cache cuts the context-dependent portion by about 4x. With it, the 100K peak drops to about 6.8GB, and the full 262K window fits in about 9.4GB. Those are vendor figures for a particular runtime, so treat them as estimates.
The author lands in the same place. He puts full context at around 10GB even with a compressed cache. He leaves headroom for other programs, so he expects roughly 30–50K tokens before running out of room. That is his estimate for his own setup, not a universal limit. Windows itself, a browser and a game launcher all take VRAM too.
What the author found
These are one person's informal tests, not controlled benchmarks.
Where it held up
- Logic: Given a notebook-bundle pricing problem with three options, it worked through the valid combinations and picked the cheapest.
- Summaries: It produced concise summaries and tool lists from longer documents without the text being split up.
- Instruction following: It satisfied all five stacked rules in a product-description prompt, including banned words and a price in his currency.
Where it fell short
- Coding: A folder-watcher script looked finished but never tracked which files it had already handled. It would have re-read the whole folder every second, including its own index file.
- Images: On a design-app screenshot it got hex values and dimensions right. The element list repeated three times, and some element names were read as settings.
- Speed: The pricing prompt ran at 27 tokens per second, against 35 for Gemma 4 E4B with the same settings.
His split is clear. He still uses Qwen 3.5 9B for tool calling, and recommends Qwen 3.6-35B-A3B if your hardware can run it. He keeps Gemma 4 for image work. He starts from PrismML's recommended sampling settings: temperature 0.7, top-p 0.95, top-k 20, with repeat and presence penalties off.
What PrismML's benchmarks say
PrismML ran its own evaluation on an H100 using EvalScope and vLLM, in thinking mode, across 15 benchmarks. The 1-bit model averaged 76.11, against 85.07 for its Qwen3.6-27B FP16 reference. That is 89.5%.
| Category | FP16 | 1-bit Bonsai 27B |
|---|---|---|
| Math | 95.33 | 91.66 |
| Coding | 88.74 | 81.88 |
| Agentic / tool calling | 80.00 | 66.03 |
| Instruction following | 78.47 | 65.74 |
| Vision | 72.61 | 59.57 |
HumanEval+ is 89.63 for Bonsai against 95.12 for FP16. The author's coding result is a good reminder that a strong score on short benchmark problems doesn't prevent a practical bug.
The author's order of strengths matches PrismML's numbers. Math holds up best. Tool calling and vision lose the most. PrismML's own limitations section says agentic coding on long, multi-file tasks isn't yet a strong target of this release. It also points to the ternary build, at about 94.6% of FP16, if quality matters more than size.
One caveat on the author's instruction-following result. PrismML's category score there is 65.74 against 78.47 for FP16, so a clean pass on one five-rule prompt doesn't mean it follows every prompt well.
Bonsai 2 and LM Studio
The author notes a newer Bonsai 2 that is ternary only, and says LM Studio can't run its files yet. Outside coverage from September agrees. One guide says Bonsai 2 27B was released on September 17, 2026, with a 5.95GB PTQ1_0 pack and a 7.21GB PQ2_0 pack. It adds that neither Ollama nor LM Studio loads these files, because LM Studio ships stock llama.cpp builds that reject the new tensor types. That guide says the 5.95GB pack leaves room for about 16K of context on an 8GB card, or about 6K with the vision pack loaded.
Those compatibility claims come from third-party guides, and runner support changes quickly. Check your exact LM Studio version before downloading.
Practical notes for 8GB Windows users
- Pick the right file: In LM Studio you need the Q1_0 language model. For image input you also need a vision projector file. A Hugging Face thread shows a user's load failure fixed by using Bonsai-27B-Q1_0.gguf plus an mmproj file. The "dspark" files in the repository are speculative-decoding drafters and won't run on their own in LM Studio.
- Start with a short context: Raise it only while watching VRAM use, and stay well under what a spec sheet says is possible.
- Treat vision as a separate memory cost: The projector is only needed for images.
- Match the model to the job: Use it for reasoning and summaries. For tool-calling workflows or image analysis, the author's smaller Qwen and Gemma models remain his picks.
Bottom line
For an 8GB card, Bonsai 27B is a real option. It can run a 27B-class model without hitting system RAM, as long as you keep the context modest and skip heavy image work. The vendor's own scores show a measurable quality drop, and the author's tests show the same weak spots in tool use and vision. Treat it as a useful extra on the shelf next to Qwen and Gemma rather than a replacement.
References
- I dropped Qwen and Gemma for a 27B model that runs smoothly on my 8GB graphics card XDA · 2026-10-10T14:30:17+00:00
- PrismML — PrismML Announces 1-bit Bonsai 27B – The First 27B Model to Run on a Phone prismml.com
- Bonsai 27B Review: Running a massive local LLM on an 8GB GPU archyde.com