An 11-row decision matrix compares five open-source distillation/training tools and five low-memory quantized-inference stacks, plus one scope check on the PRISM robot-planning paper. The 11 rows are drawn from seven independent hosts, and each row has its own source URL. Memory figures are in GB of weights unless stated; KV cache and runtime overhead are additional.
Hugging Face's on-policy knowledge distillation with generalized Jensen-Shannon divergence loss; supports LoRA, gradient checkpointing, DeepSpeed/FSDP, and huge teachers (100B+) on an external vLLM server. Checkpoints reload with from_pretrained().
Logit-based KD and quantization-aware distillation (QAD) via Megatron-Bridge — the pick when the distilled student must ship quantized on NVIDIA infrastructure. Exports Megatron and Hugging Face formats.
PyTorch-native KD recipes with LoRA/QLoRA, FSDP2 and QAT+LoRA; model support explicitly includes Gemma 2 27B and up to 405B. SDFT is an interesting self-distillation option but not the default start.
| Tool | Category | Best for | Memory guidance | Model support | Source |
|---|
Decision matrix of 11 rows gathered from 11 pages on 7 independent hosts: 5 open-source distillation/training tools, 5 quantized-inference options for 27B-class models, and 1 scope check on the PRISM robot-planning paper. Each row carries its own source URL; no URL backs more than one row. Memory figures are approximate weight footprints in GB; KV cache, context and runtime overhead are additional. Long method and deployment details were shortened in the table — hover a box in the diagram or follow a source link for the full text. Official documentation (huggingface.co, github.com project pages) is labeled "official docs"; hardware figures from aixcove.com, intuitionlabs.ai, kingy.ai and grokipedia.com are third-party guidance. MLX-Flash is experimental.