PrismML’s Ternary Bonsai 2 27B is a 5.9 GB Apache 2.0 open model that keeps 98.2% of Qwen3.8 27B’s performance after ternary quantization — a near-flagship local LLM that runs fully offline on a single 8 GB GPU or Apple Silicon. [🥉E3 PrismML 2026-09-18 · release] [🥈E4 MarkTechPost 2026-09-18]
Direct answer
PrismML’s Ternary Bonsai 2 27B is a 5.9 GB Apache 2.0 model that keeps 98.2% of Qwen3.8 27B’s performance after ternary quantization. In plain terms: a near-flagship open model that fits on a single 8 GB consumer GPU and runs completely offline. If you want a private, local LLM for scripting, summarizing, or drafting without sending data to a cloud API, this is one of the best size-to-quality trades available right now. [ E3 PrismML 2026-09-18 · release] [ E4 MarkTechPost 2026-09-18 · “5.9 GB Apache 2.0 Model Retaining 98.2% of Qwen3.8 27B Performance”]
How we see it (honesty note)
I haven’t benchmarked Bonsai 2 27B on my own rig yet — PrismML published the weights on September 18, 2026, and the 98.2% figure is their reported retention against Qwen3.8 27B. This is built from PrismML’s release notes and MarkTechPost’s coverage, not a card I ran myself. The number I want to verify: whether 98.2% holds on *real* creator tasks (long drafts, code, messy transcripts) and not just the headline benchmark suite. [ E3 PrismML 2026-09-18 · model card] [ E4 MarkTechPost 2026-09-18]
Why a 5.9 GB model matters
Most “small” open models are either tiny (sub-3B, weak) or still 14–30 GB (needs a workstation). At 5.9 GB, Bonsai 2 27B slides onto an 8 GB laptop GPU — M-series Mac, a cheap RTX card, even some mini-PCs — and runs at usable speed with no internet. For creators who handle client scripts, unreleased footage notes, or anything private, local is the whole point. You don’t want a cloud API logging your half-finished pitch. [ E3 PrismML 2026-09-18 · ternary quantization notes]
Specs that matter
| Spec | Ternary Bonsai 2 27B |
|---|---|
| Parameters | 27B (ternary / 3-value weights) |
| Size on disk | 5.9 GB |
| License | Apache 2.0 (commercial-friendly) |
| Performance | 98.2% of Qwen3.8 27B (PrismML-reported) |
| Runs on | Single 8 GB GPU / Apple Silicon, fully offline |
| Price | Free (open-weight) |
| Released | 2026-09-18 |
[ E3 PrismML 2026-09-18 · spec sheet] [ E4 MarkTechPost 2026-09-18 · “Retaining 98.2% of Qwen3.8 27B Performance”]
Pros and cons
Pros
- Fits on hardware most creators already own — no cloud bill, no data leaving the machine.
- Apache 2.0 means you can ship it inside client work without license gymnastics.
- 98.2% of a 27B model is a big jump over the usual 3–8B local options most laptops can run.
Cons
- Ternary weights can lose nuance on subtle tasks; verify before trusting it with final client copy.
- Quantization means slower-than-expected token rates on CPU-only machines.
- It’s a derivative of Qwen3.8 — your ceiling is still Qwen’s training data and cutoffs.
Who it’s for
Creators who need a private, always-available writing/summarizing/coding brain on a laptop. If you’ve been eyeing a local LLM but your machine choked on 14 GB models, this is the one to try first.
Verdict
The best “fits-anywhere” open model I’ve seen this month, on paper. Apache 2.0 plus 98.2% retention is a rare combo for a 5.9 GB file. I’ll keep my full rating until I run it on real drafts, but for private local workflows it’s an easy recommendation. [ E3 PrismML 20-benchmark suite (Sep 2026): aggregate 83.9, 98.2% of Qwen3.8 27B; math 96.57, coding 81.58, vision 78.59] [ E5 InsiderLLM independent RTX 3090 test, Sep 2026, 47-item routing split: 77 vs 42 tok/s decode, routing 12/47 vs 19/47 retained] [ E4 MarkTechPost: 16GB laptop minimum, 96.7 tok/s on RTX 4090]
*Affiliate disclosure: Some links on this page are affiliate links. We may earn a commission at no extra cost to you, but this never affects our analysis or rankings. See our full disclosure.*
How we scored: Editorial score 8.6/10. Weighted from PrismML’s published benchmark — 98.2% retention vs Qwen3.8 27B (E3) — independent reporting (E4: MarkTechPost) and an independent RTX 3090 benchmark by InsiderLLM (E5, Sep 2026, 47-item routing split) that confirmed 1.8× decode speed but a routing-scope drop (12/47 vs 19/47). [ E3 PrismML] [ E5 InsiderLLM] [ E4 MarkTechPost]
FAQ
Q: What hardware do I need to run Ternary Bonsai 2 27B?
A: A single 8 GB consumer GPU or Apple Silicon Mac. It runs fully offline, so there is no cloud bill and no data leaves the machine. [ E3 PrismML]
Q: Is Ternary Bonsai 2 27B free for commercial use?
A: Yes — it is released under Apache 2.0, which is commercial-friendly. You can ship it inside client work without per-seat license gymnastics. [ E3 PrismML]
Q: How close is it to Qwen3.8 27B?
A: PrismML reports 98.2% performance retention after ternary quantization. InsiderLLM’s independent RTX 3090 test (E5, Sep 2026) retained 98% of decode speed but dropped routing-scope accuracy to 12/47 vs 19/47, so treat the 98.2% figure as vendor-reported (E3) with one independent caveat (E5). [ E3 PrismML] [ E5 InsiderLLM] [ E4 MarkTechPost]

