Only one Qwen3.8-27B "distill" truly fits a 12 GB RTX 5070 — Prism ML's 5.9 GB Bonsai 2, and it is a quantization, not a distillation

Asked (summary):

What are the current best Qwen 3.8 27B distilled models? Go deep and find ones that will run on a 12 GB RTX 5070 with good performance — list names, where to get them, and who made them (e.g. Prism).

As of 9 October 2026, the evidence covers 793 extracted records from three research passes over the Qwen3.8-27B derivative ecosystem — Prism/low-bit builds, quantization repositories, and fine-tuned/distillation-style releases — with heavy deduplication; the broad derivative pass searched over 1,500 web results and surfaced hundreds of candidate mentions. The report is organized around a six-model shortlist, not every upload. The key distinction: "runs" is not the same as "fully GPU-resident with good performance."

12 GB compatibility matrix: file size vs the card's ceiling

Smallest quality-credible quant per model, GB on disk. The 12 GB line is the RTX 5070's VRAM; weights are not the whole bill — KV cache and runtime overhead push real use higher (base Q2 shows a measured ~14 GB floor even at 4K context).
YES — full GPUHYBRID — CPU offload neededUNVERIFIED 12 GB fit

What the chart cannot say

Bonsai 2 is a 1.76-bit ternary post-training quantization of the unchanged base, not a teacher-output distillation — yet it retains 98.2% of FP16 (84.78 vs 86.32, vendor-reported) and a community user measured 26.5 tok/s on a 12 GB RTX 3060. platform-monkey.com
ThinkingCap cuts thinking tokens 37.2% across 12 benchmarks for under a point of accuracy (86.65% → 85.79%) — but its smallest cited good GGUF, IQ4_XS, is 15.5 GB. kaitchup.substack.com
Swift 1.5 is the one fine-tune where accuracy rises as tokens fall — LiveCodeBench just under 77% → 81.7% with output tokens down from ~11,200 to ~8,400 — and independent Aider runs showed Pass1 30.8 vs 27.1 at half the wall-clock. youtube.com
Standard Q2 base quants weigh ~10.2 GB on disk, yet batiai measured a 14 GB total-memory floor even with context squeezed to 4,096 — weights alone never settle the 12 GB question. huggingface.co/batiai

The shortlist, in recommendation order

1
Prism ML
Bonsai 2 27B (Ternary-Bonsai-2-27B-gguf)
YES — full GPU

The only well-supported Qwen3.8-27B derivative whose weights (5.9 GB) genuinely fit wholly in 12 GB with context headroom. 98.2% of FP16 benchmark average (84.78 vs 86.32, vendor), 143 tok/s on an RTX 5090 (vendor) and 26.5 tok/s community-measured on a 12 GB RTX 3060.

Runtime / quant: llama.cpp (Prism fork) · PTQ1_0 GGUF 5.95 GB · MLX 2-bit companion

Caveat: Not a teacher-output distillation — a 1.76-bit ternary post-training quantization of the unchanged base. Quality numbers are vendor-reported; the 12 GB throughput figure is independent (community).

Download: huggingface.co/prism-ml

2
BottleCap AI
ThinkingCap-Qwen3.8-27B
HYBRID — CPU offload

Best general-purpose efficient-reasoning fine-tune: 37.2% fewer thinking tokens across 12 benchmarks while macro accuracy moves only 86.65% → 85.79% (publisher release measurements).

Runtime / quant: Official GGUF repo; smallest cited good quant IQ4_XS at 15.5 GB (Q4_K_M 17.4 GB) — exceeds 12 GB, so partial CPU/RAM offload on an RTX 5070

Caveat: PolyForm Small Business license, not Apache. GGUF evals cover only five benchmarks; quants not proven lossless.

Download: huggingface.co/bottlecapai

3
Yucus AI (UkisAI)
Swift 1.5 (Swift-Qwen3.8-27B)
UNVERIFIED

Most aggressive generalist/coding efficient-reasoning option: up to 58.5% fewer thinking tokens; LiveCodeBench rises from just under 77% to 81.7% while output tokens fall from ~11,200 to ~8,400. Independent Aider evals: Pass1 30.8 vs 27.1 base, half the time per case.

Runtime / quant: Many runtimes: GGUF, MLX, NVF4, AMD builds, free research API. IQ2-family GGUFs (9.1–11.0 GB) are listed as fitting 12 GB, but KV cache costs 64 KiB/token

Caveat: Treat true 12 GB full-GPU fit as unverified in the collected evidence; the quality-grade Q4_K_M is 18.0 GB.

Download: huggingface.co/ukisai

4
Qwen Pi team
Qwen Pi (Pi coding-agent checkpoint)
UNVERIFIED

Best specialist for the minimal Pi coding-agent harness: its medium setting reportedly matches base X-high on Terminal-Bench with about 41% fewer output tokens. SFT on successful real Pi sessions plus RL with a reasoning-efficiency reward.

Runtime / quant: Apache 2. No quant sizes or 12 GB measurements appear in the collected evidence

Caveat: Narrow: base model still wins on GPQA at medium; not proven to transfer to other agent harnesses. 12 GB fit unverified.

Download: youtube.com

5
DavidAU
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF
HYBRID — CPU offload

Strongest adventurous uncensored/coding fine-tune by its publisher's numbers: mxfp8 ARC-C 0.735 and ARC-E 0.882 (first fine-tune past those marks per the card), thinking tokens cut 1/2 to 1/10. Includes light Claude Opus reasoning traces and GPT-5 Polaris data.

Runtime / quant: Q4_K_M (18.05 GB) minimum recommended for reliable tool calling, Q6 preferred — CPU offload on 12 GB. Solstice-AI UltraOptimised IQ4_XS is 16.58 GB

Caveat: Benchmark claims are primarily publisher-produced; label accordingly.

Download: huggingface.co/DavidAU

6
DavidAU
Qwen3.8-27B-Cold-Fusion-GAIN-V1.1
HYBRID — CPU offload

Less extreme DavidAU alternative with reduced overthinking and a detailed published benchmark table; the card reports 99% of BF16 power retained at 8-bit and 4-bit and 'cook 1' exceeding all 7 core Qwen 3.8 benchmarks.

Runtime / quant: NEO-MAX-MTP GGUF repo; good 4-bit quants still exceed 12 GB

Caveat: Publisher-produced benchmarks; not a full-GPU 12 GB fit at its good 4-bit quant.

Download: huggingface.co/DavidAU

Why ordinary base-model quants are not the answer

batiai, LM Studio-style and Unsloth standard Q2–Q4 releases of the unchanged base are quantizations, not distilled models — and they do not deliver true 12 GB full-GPU operation either. Unsloth's 2-bit tier (11–13 GB) is called "workable on a 12 GB card in a pinch, with a real quality cost"; batiai's measurements show even the smallest Q2 at a ~14 GB memory floor at 4K context, and bartowski's IQ2-family quants show perplexity climbing from 6.76 (Q4_K_M) to 8.41 (IQ2_XXS). Sub-3-bit conventional quants collapse on sustained reasoning — which is precisely the regime Bonsai 2's ternary-plus-Hadamard method was built to survive.

All seven rows

#MakerModel12 GB verdictSmallest good quantKey evidence (measured vs publisher)Source
1Prism MLBonsai 2 27BYES — full GPUPTQ1_0 · 5.9 GB98.2% of FP16 (84.78/86.32, vendor); 26.5 tok/s on 12 GB RTX 3060 (community)huggingface.co/prism-ml
2BottleCap AIThinkingCap-Qwen3.8-27BHYBRID — CPU offloadIQ4_XS · 15.5 GB−37.2% thinking tokens, 86.65% → 85.79% (publisher)huggingface.co/bottlecapai
3Yucus AISwift 1.5UNVERIFIEDQ4_K_M · 18.0 GB (IQ2_M 10.7 GB listed)LiveCodeBench ~77% → 81.7%, up to −58.5% thinking tokens (publisher); Aider Pass1 30.8 vs 27.1 (independent)huggingface.co/ukisai
4Qwen Pi teamQwen PiUNVERIFIEDno quant size citedMedium ≈ base X-high on Terminal-Bench, ~41% fewer output tokens (publisher)youtube.com
5DavidAUTURBO Fable Cold Fusion 735-882HYBRID — CPU offloadQ4_K_M · 18.05 GB minmxfp8 ARC-C 0.735, ARC-E 0.882; thinking tokens −1/2 to −1/10 (publisher)huggingface.co/DavidAU
6DavidAUCold-Fusion-GAIN-V1.1HYBRID — CPU offloadQ4_K_M min, size n/a99% of BF16 at 4/8-bit; exceeded 7 core Qwen 3.8 benchmarks (publisher)huggingface.co/DavidAU
—bartowski / batiai / UnslothBase Qwen3.8-27B quants (not distills)NO true 12 GB full-GPUQ2_K · 10.2–10.8 GBMeasured ~14 GB total-memory floor even at 4K context (batiai); 2-bit 'workable in a pinch, real quality cost'huggingface.co/bartowski

Method: compiled 2026-10-09 from 793 extracted web records across three research passes (Prism/low-bit models, Qwen3.8-27B quantizations, fine-tuned/distilled candidates); the broad derivative pass searched 1,500+ web results and required heavy deduplication. File sizes are GB on disk per each repository's own table; verdicts compare the smallest quality-credible quant against 12 GB VRAM, noting KV-cache overhead. Benchmark figures are labelled vendor/publisher or independent/community as sourced. Duplicate records, mirrors, and sub-shortlist candidates (e.g. barozp and rico03 Opus distills) were cut for space. This is a best-current shortlist, not an exhaustive benchmark of every Hugging Face upload.

This report was generated automatically by Keenable SELECT at a user's request, from publicly available web sources linked herein. Keenable does not review, verify, or endorse its contents and makes no representation as to accuracy, completeness, or timeliness; AI-based extraction may contain errors. Nothing in this report is investment, legal, financial, or other professional advice. All trademarks and referenced content remain the property of their respective owners; no affiliation or endorsement is implied. To report an error, rights concern, or request removal: legal@keenable.ai.

Keenable SELECTAsk your own question
Made with Keenable SELECT