Seven small open models for Russian fine-tuning: Meno-Lite, T-lite and YandexGPT-5-Lite lead the shortlist

Asked

“What are the best recent small (<9B) opensource model for fine-tuning with russian language support?”

Seven shortlisted open-weight models under 9B parameters, 1.47B to 8B, compared on Russian specialization, context length, checkpoint type, fine-tuning readiness, license clarity and Russian benchmark evidence as of 28 September 2026. Evidence is predominantly official Hugging Face model cards (six of seven rows) plus one arXiv paper — not broad independent comparative testing.

Decision matrix: pick by use case, not by a single score

Russian benchmarks on these cards use different datasets and scales, so cells grade the evidence rather than rank raw scores. Hover or tap any cell for the full detail from the source.

strong / confirmed verify or partial not resolved on source informational
Meno-Lite-0.1 (7B, Apache 2.0, 128K) is the best fit for Russian RAG and document work — LIBRA passkey 0.98 at 128K, but its card warns complex reasoning degrades beyond 32K. huggingface.co
T-lite-it-2.1 (8B) is the strongest ready-made Russian assistant and tool-use candidate: Ru Arena Hard 83.9, ruIFeval 75.9, ruBFCL 56.5, with its Qwen3-8B base in the model tree. huggingface.co
YandexGPT-5-Lite-Pretrain (8B) is the cleanest Russian-first base for domain LoRA or full tuning — 15T-token RU/EN pretraining, llama-like architecture and a documented torchtune example. huggingface.co
QVikhr-3-4B (Ru Arena General 78.2) is the economical Russian specialist at 4B, while Phi-4-mini (3.8B, MIT, 128K) is the safest permissively licensed multilingual pick with an official TRL/Accelerate SFT script — though its card reports no Russian-specialist benchmarks. huggingface.co · huggingface.co
Below 2B, Gamayun 1.5B is the quality-oriented Russian option (MERA 38.1, RuBIN 35.2) and Zarya-1.7B/2B an experimental MIT-licensed hybrid with released training code — a research option, not a production default. arxiv.org · huggingface.co
ModelParamsCheckpointContextLicenseBest forKey caveatSource

Method: 7 shortlisted open-weight models under 9B parameters, one row per model, extracted as of 2026-09-28 from official Hugging Face model cards (6 rows) and one arXiv paper (Gamayun); each row has its own source URL. Cells grade evidence — Russian focus, context tokens, checkpoint type, fine-tuning readiness, license clarity, Russian benchmark reporting — because benchmark scores across cards use different datasets and scales and are not directly comparable. “Verify terms” marks licenses the extracted rows did not resolve. Full benchmark strings and long support notes were shortened for space; hover the matrix for the source text.

This report was generated automatically by Keenable SELECT at a user's request, from publicly available web sources linked herein. Keenable does not review, verify, or endorse its contents and makes no representation as to accuracy, completeness, or timeliness; AI-based extraction may contain errors. Nothing in this report is investment, legal, financial, or other professional advice. All trademarks and referenced content remain the property of their respective owners; no affiliation or endorsement is implied. To report an error, rights concern, or request removal: legal@keenable.ai.

Keenable SELECTAsk your own question
Made with Keenable SELECT