What are the current best Qwen 3.8 27B distilled models? Go deep and find ones that will run on a 12 GB RTX 5070 with good performance — list names, where to get them, and who made them (e.g. Prism).
As of 9 October 2026, the evidence covers 793 extracted records from three research passes over the Qwen3.8-27B derivative ecosystem — Prism/low-bit builds, quantization repositories, and fine-tuned/distillation-style releases — with heavy deduplication; the broad derivative pass searched over 1,500 web results and surfaced hundreds of candidate mentions. The report is organized around a six-model shortlist, not every upload. The key distinction: "runs" is not the same as "fully GPU-resident with good performance."
The only well-supported Qwen3.8-27B derivative whose weights (5.9 GB) genuinely fit wholly in 12 GB with context headroom. 98.2% of FP16 benchmark average (84.78 vs 86.32, vendor), 143 tok/s on an RTX 5090 (vendor) and 26.5 tok/s community-measured on a 12 GB RTX 3060.
Best general-purpose efficient-reasoning fine-tune: 37.2% fewer thinking tokens across 12 benchmarks while macro accuracy moves only 86.65% → 85.79% (publisher release measurements).
Most aggressive generalist/coding efficient-reasoning option: up to 58.5% fewer thinking tokens; LiveCodeBench rises from just under 77% to 81.7% while output tokens fall from ~11,200 to ~8,400. Independent Aider evals: Pass1 30.8 vs 27.1 base, half the time per case.
Best specialist for the minimal Pi coding-agent harness: its medium setting reportedly matches base X-high on Terminal-Bench with about 41% fewer output tokens. SFT on successful real Pi sessions plus RL with a reasoning-efficiency reward.
Strongest adventurous uncensored/coding fine-tune by its publisher's numbers: mxfp8 ARC-C 0.735 and ARC-E 0.882 (first fine-tune past those marks per the card), thinking tokens cut 1/2 to 1/10. Includes light Claude Opus reasoning traces and GPT-5 Polaris data.
Less extreme DavidAU alternative with reduced overthinking and a detailed published benchmark table; the card reports 99% of BF16 power retained at 8-bit and 4-bit and 'cook 1' exceeding all 7 core Qwen 3.8 benchmarks.
| # | Maker | Model | 12 GB verdict | Smallest good quant | Key evidence (measured vs publisher) | Source |
|---|---|---|---|---|---|---|
| 1 | Prism ML | Bonsai 2 27B | YES — full GPU | PTQ1_0 · 5.9 GB | 98.2% of FP16 (84.78/86.32, vendor); 26.5 tok/s on 12 GB RTX 3060 (community) | huggingface.co/prism-ml |
| 2 | BottleCap AI | ThinkingCap-Qwen3.8-27B | HYBRID — CPU offload | IQ4_XS · 15.5 GB | −37.2% thinking tokens, 86.65% → 85.79% (publisher) | huggingface.co/bottlecapai |
| 3 | Yucus AI | Swift 1.5 | UNVERIFIED | Q4_K_M · 18.0 GB (IQ2_M 10.7 GB listed) | LiveCodeBench ~77% → 81.7%, up to −58.5% thinking tokens (publisher); Aider Pass1 30.8 vs 27.1 (independent) | huggingface.co/ukisai |
| 4 | Qwen Pi team | Qwen Pi | UNVERIFIED | no quant size cited | Medium ≈ base X-high on Terminal-Bench, ~41% fewer output tokens (publisher) | youtube.com |
| 5 | DavidAU | TURBO Fable Cold Fusion 735-882 | HYBRID — CPU offload | Q4_K_M · 18.05 GB min | mxfp8 ARC-C 0.735, ARC-E 0.882; thinking tokens −1/2 to −1/10 (publisher) | huggingface.co/DavidAU |
| 6 | DavidAU | Cold-Fusion-GAIN-V1.1 | HYBRID — CPU offload | Q4_K_M min, size n/a | 99% of BF16 at 4/8-bit; exceeded 7 core Qwen 3.8 benchmarks (publisher) | huggingface.co/DavidAU |
| — | bartowski / batiai / Unsloth | Base Qwen3.8-27B quants (not distills) | NO true 12 GB full-GPU | Q2_K · 10.2–10.8 GB | Measured ~14 GB total-memory floor even at 4K context (batiai); 2-bit 'workable in a pinch, real quality cost' | huggingface.co/bartowski |
Method: compiled 2026-10-09 from 793 extracted web records across three research passes (Prism/low-bit models, Qwen3.8-27B quantizations, fine-tuned/distilled candidates); the broad derivative pass searched 1,500+ web results and required heavy deduplication. File sizes are GB on disk per each repository's own table; verdicts compare the smallest quality-credible quant against 12 GB VRAM, noting KV-cache overhead. Benchmark figures are labelled vendor/publisher or independent/community as sourced. Duplicate records, mirrors, and sub-shortlist candidates (e.g. barozp and rico03 Opus distills) were cut for space. This is a best-current shortlist, not an exhaustive benchmark of every Hugging Face upload.