Как автоматически размечать качество подсказок бота (шкала A–E, генерирует DeepSeek, судит Claude; κ=0,641 на 140 парах одного человека, бинарно 0,584) без второго живого разметчика и не обмануть себя: жюри из моделей и агрегация, разрыв круга калибровки, слабые метки из копирования/сходства/времени/исхода, active learning на 30–50 пар, синтетические A–E и дрейф, метрики и пороги, инструменты для маленькой команды, стоимость 1 000 пар и месяца.
Отчёт построен на 60 отобранных исследованиях и инженерных блогах (жюри и судьи, weak supervision, active learning, синтетика, метрики; 47 исследовательских строк и 13 блогов/вторичных), 20 официальных страницах инструментов, объединённых в 10 продуктов, и 22 строках официальных API-цен. Все данные собраны 2026-09-13. Свидетельство о пользе жюри смешанное: панели иногда улучшают точность, но коррелированные ошибки резко уменьшают эффективное число голосов, а в одном клиническом примере лучший одиночный судья обошёл жюри.
Панель судей несёт меньше информации, чем кажется: девять судей ≈ 2–2.5 эффективных голоса из-за коррелированных ошибок — поэтому базовое жюри ограничено тремя разными семействами моделей. arxiv.org
Жюри не гарантирует прироста над лучшим одиночным судьёй: в клиническом кейсе κ=.74 у жюри против κ=.75 у Gemini — рост относительно текущих κ=0.641/0.584 нельзя обещать без доменного held-out теста. arxiv.org
Внутренний консенсус моделей без внешнего человеческого якоря вводит в заблуждение: Dawid–Skene оценивал точность GPT в 90.7% там, где внешняя проверка её не подтверждала — anchor из 140 пар остаётся неприкосновенным. lrec.elra.info
Универсального «процента ручной проверки» не существует: rule of three даёт при нуле ошибок в 50 случайных проверках лишь верхнюю 95%-границу ≈6%; практический ритм — 10–15 пар в неделю плюс месячный аудит 30–50. jatinbansal.com
Три независимые семейства (Claude Sonnet 5, Gemini 3.1 Pro, Grok 4.20 reasoning), общий контекст и rubric, JSON-ответ с вероятностями и evidence spans. DeepSeek — только challenger: он генерирует подсказки, возможна self-preference. Majority vote — baseline; веса/матрицы ошибок обучать только на calibration-fold. Конфликт с D или отсутствие уверенного A+B → abstain, не оптимистичное большинство.
Коррелированные ошибки: панель из 9 судей несёт ~2–2.5 эффективных голоса; разрыв с независимым идеалом 8–22 п.п.
Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels — neff≈2.0–2.5; Condorcet gap is 8–22 percentage points (pp; permutation p<10^-4)
Смешанное свидетельство: в клиническом кейсе лучший одиночный судья κ=.75, жюри κ=.74 — панель не гарантирует прироста
aipsy-judge: A Specialized, Psychologist-Corrected Local Judge for the Psychological Safety of Conversational AI — Gemini kappa=.75, jury kappa=.74
Специализированная панель (судья на критерий) даёт +10.9 п.п. к лучшему одиночному судье
Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring — +10.9 pp
Reference-free панели показывают коррелированные ложноотрицательные; консенсус-риск нужно диагностировать на калибровочном probe
JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification — FN-only correlations 0.402 and 0.368; lifts 3.13x and 18.13x
Инженерный опыт: три разные линии моделей и стоп — дальнейшие судьи возвращают коррелированные ошибки
We Audited Our Own LLM-Judge Panel — $0.002 a candidate
Неизменяемый human anchor из 140 пар делится на calibration и untouched test по диалогам; авто-метки никогда не становятся gold, provenance обязательна. Раз в квартал — независимый второй человек на 30–50 пар, чтобы оценить человеческий потолок. Универсального «процента ручной проверки» нет: rule of three даёт при 0 ошибок в n случайных проверках верхнюю 95%-границу ≈3/n (n=50 → ~6%, n=100 → ~3%), кластеризация диалогов её ухудшает.
Внутренний консенсус LLM без внешнего человеческого якоря серьёзно вводит в заблуждение (Dawid–Skene переоценивает точность)
Multi-Source Emotion Annotation in Children’s Language: When LLM Consensus Diverges from Human Judgment — Dawid-Skene estimates GPT-5.2 valence accuracy at 90.7%, whereas evaluation against human gold yields 71.0%
Без корректного reference судьи согласны с экспертами только там, где сами могут решить задачу — «нет бесплатных меток»
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding — 54% of the question-response pairs received unanimously consistent annotations
Уточнение annotation guidelines подняло согласие 0.593 → 0.84 — вклад rubric сравним с вкладом модели
Automating Annotation Guideline Improvements using LLMs: A Case Study — inter-annotator agreement going from 0.593 to 0.84 when improving WNUT-17’s guidelines
Клинический пример: согласие модель–человек κ≈.68–.75 — реалистичный ориентир потолка, не 1.0
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis — Cohen’s Khuman×gemini = .75, Khuman×qwen = .68, Khuman×kimi = .56, Khuman×jury = .74
Логировать копирование, точный вставленный фрагмент, последующие правки, тему, длину, исход. Abstaining labeling functions: точная копия/малые правки → слабый плюс к A/B; существенные добавления → C; topic mismatch → E; противоречие по числам/дате/ставке → D. Время и бизнес-исход — ковариаты и диагностика, не прямая истина. Объединять label model / Dawid–Skene как диагностический baseline.
Слабая супервизия по поведению пользователя улучшает выявление ошибочных меток на 36.7%
LabelAId: Just-in-time AI Interventions for Improving Human Labeling Quality and Domain Knowledge in Crowdsourcing Systems — 36.7%
Отбор подмножества слабых меток (cut statistic) даёт до +19% точности пайплайнов weak supervision
Training Subset Selection for Weak Supervision — up to 19% (absolute) on benchmark tasks
Редкие turn-level свидетельства отказа (1–10% диалогов) достаточны для обучения предиктора при правильной архитектуре
When Evidence is Sparse: Weakly Supervised Early Failure Alerting in Dialogs and LLM-Agent Trajectories — 1-10%
Имплицитный фидбэк смещён: менее 5% ответов получают отметку — селекция портит eval-набор без строгого разделения сигналов
The Thumbs-Up Button That Poisoned Your Eval Set Through the Back Door — fewer than five percent of responses receive any user feedback at all
LLM-аннотация исхода диалога коррелирует с оценкой клиента 0.47 против 0.36 у сентимента; тон и удовлетворённость расходятся в 44%
What sentiment analysis can't see: Measuring whether customers were helped, and what went wrong, across 70,000 support conversations — 70,450 support conversations; correlation of 0.47 versus 0.36; tone and satisfaction disagreed in 44% of conversations
Стратификация: ~40% расхождения жюри/низкий margin, 20% границы B/C и C/D, 20% редкие темы, новые KB-записи и числовые ответы, 20% чисто случайные — иначе нельзя оценить реальную частоту ошибок. Дедупликация и bootstrap — по диалогам.
Отбор по границе rubric даёт 50–80% экономии бюджета аннотаций при том же качестве; но uncertainty-only переиндексирует edge cases
Human-in-the-Loop Feedback Loops for LLM Systems — 50–80% reductions in annotation budget
LLM-метки + выборочная человеческая проверка неопределённых строк: качество полной ручной разметки при 6–7% усилий
Human-in-the-Loop Feedback Loops for LLM Systems — 6–7% of the human annotation effort
Контрпример: active learning не даёт надёжного преимущества над случайной выборкой при LLM-аннотаторе
Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection — \28 for GPT-5.2 Batch API vs. \316 for Prolific
ActiveLab: консенсус-метки при 5× меньшем числе аннотаций в задачах перепроверки меток
ActiveLab: Active Learning with Data Re-Labeling — 5x fewer total annotations
Canary-набор вне обучения: A — парафраза проверенной пары; B — только стиль; C — удалить обязательный слот; D — заменить ставку/дату/статус/знак; E — ответ другой темы. Paired metamorphic tests, seed и версия фиксируются, реалистичность проверяется вручную; фиксированная доля canary/anchor интерливится при каждом релизе судьи.
Фиксированный human-labeled anchor, пересчитываемый текущим судьёй, отделяет дрейф судьи от дрейфа системы (anytime-valid)
Who Drifted: the System or the Judge? Anytime-Valid Attribution in LLM Evaluation Pipelines — On two real judge changes, a silent version bump is detected as judge drift in 60/60 runs with zero judge-to-system misattribution, and a contaminating strict-prompt change is correctly attributed on 110 of 120 runs at guard width 300 -- while the industry-default rolling z-test false-alarms on 75% of drift-free streams.
Синтетические контрастные пары + rejection sampling: судья вырос 75.4 → 88.3 на RewardBench без человеческих меток
Synthetic Eval Bootstrapping: How to Build Ground-Truth Datasets When You Have No Labeled Data — 75.4 to 88.3
Метаморфические отношения (инвариантность смысла, негация) масштабируются до сотен тысяч тестов
Synthetic Eval Bootstrapping: How to Build Ground-Truth Datasets When You Have No Labeled Data — 560,000
Critique shadowing против human-labeled validation поднял точность авто-судьи 68% → 94% (блог, не переносить цифру как факт)
Synthetic Eval Bootstrapping: How to Build Ground-Truth Datasets When You Have No Labeled Data — 68% to 94%
Pairwise Cohen κ (модель–человек, человек–человек); Fleiss κ для панели; Krippendorff α ordinal только для A–D и nominal для A–E; weighted κ для A–D; бинарный A+B vs C+D; coverage/abstention, macro-F1, confusion matrix, recall D и C+D. CI — stratified/cluster bootstrap по диалогам. Предрегистрированный внутренний gate (не универсальный стандарт): нижняя граница 95% CI binary κ ≥ baseline и 0.60; recall D ≥ .95, C+D ≥ .90; ordinal α ≥ .67, для жёстких решений ~.80.
Скан 24 работ: поле не имеет единой практики отчётности agreement-метрик — фиксируйте набор заранее
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why — 24 papers; 11 graded or continuous evaluations; 7 pairwise-preference evaluations; 3 binary-verdict evaluations; 2 judge ensembles; 1 scale-matched mix
Krippendorff α = 0.78 между людьми на агентных задачах — ориентир достижимого согласия
Counsel: A Meta-Evaluation Dataset for Agentic Tasks — Krippendorff’s alpha 0.78; strongest judge achieving 88% agreement on location and 65% on reasoning
Trust-or-escalate: калибровка уверенности судьи снижает ECE на 50% и даёт управляемую эскалацию к человеку
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement — reducing ECE by 50% and improving AUROC by 13% for GPT-4
Согласие эксперт–LLM гетерогенно: weighted κ от 0.17 до 0.86 по 21 измерению — считать по классам, не в среднем
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need? — 0.17 to 0.86 of weighted kappa across 21 evaluation dimensions
Конформный risk control для панелей: acceptance .973 при контролируемом риске .098 — формальный каркас для abstention
Multi-Expert Conformal Risk Control for Pairwise LLM Judging in Open-Ended Dialogue — Acceptance Rate .973, Accuracy .887, Risk .098 for Decision Voting on ESConv (Homogeneous)
Label Studio OSS или Argilla — человеческая очередь и ревью; Langfuse — если нужна self-hosted трассировка с annotation queues; Promptfoo или Inspect — версионированные regression/canary evals; LangSmith/Braintrust — облачная альтернатива, публичные цены частично не указаны; OpenAI Evals/API — программные graders, не замена UI разметки. Пропуски в официальных страницах отмечены «не указано», не додуманы.
Режим «человек редактирует вердикт LLM» повышает согласие 0.64 → 0.71 — аргумент за review-очередь, а не чистую автоматизацию
Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks — from 0.64 to 0.71
Rubric-based оценка даёт больше выполненных оценок и более полезные объяснения, чем pairwise (n=15, 131 оценка)
Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences — n=15 practitioners completing 6 tasks yielding 131 evaluations; judge strategy impact on number of evaluations F(1, 14) = 12.64, p = 0.003; pairwise strategy led to significantly fewer evaluations than the rubric strategy with mean difference = 0.72, p = 0.007; positional bias appeared in 35% of 196 total evaluations; perceived helpfulness when positional bias was present vs absent t(28) = 2.98, p = 0.003; rubric explanations rated higher than pairwise explanations with mean difference = 0.73
Практика панелей: crowd-vs-expert согласие 72–83%, expert-vs-expert 79–90% — ориентиры для ожиданий от ревью
Weak judges, strong panel: an ensemble approach to LLM eval — crowd-vs-expert agreement of 72–83%, and expert-vs-expert at 79–90%
По официальным ценам на 2026-09-13, при 3 000 input + 200 output токенов на судью: Claude Sonnet 5 $8.00, Gemini 3.1 Pro Preview $8.40, Grok 4.20 reasoning $4.25 — итого ≈$20.65 за 1 000 пар на три судьи. При 150–200 пар/нед (≈650–866 пар/мес) ≈$13.4–17.9/мес; с арбитром на ~30% конфликтов, ретраями и canary — порядка $15–20/мес. Замена Grok на deepseek-reasoner даёт ≈$18.49/1 000, но ломает независимость от генератора подсказок. Не включены платформа, embeddings/NLI и человеческое время; оценки чувствительны к длине reasoning-вывода и preview-статусу.
Средняя модель с правильным дебиасингом может обойти frontier-судью за долю цены (κ=0.549 у Gemini 2.5 Flash)
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines — 71.0%, kappa=0.549
Инженерный ориентир: панель малых судей из разных семейств — около $0.002 за кандидата
We Audited Our Own LLM-Judge Panel — $0.002 a candidate
Self-preference bias судьи измерим и снижается калиброванными многомерными rubric — довод против DeepSeek-судьи над DeepSeek-подсказками
Quantifying and Mitigating Self-Preference Bias of LLM Judges — achieving a twofold reduction in prediction error compared to uncalibrated baselines
Внутренняя схема команды: строки — предложенные abstaining labeling functions, объединяемые label model / Dawid–Skene как диагностическим baseline, не как ground truth.
Пороги — внутренние, предрегистрированные до запуска, настраиваются по стоимости false positive/false negative; это не универсальный стандарт из литературы.
«Не найдено / публично не указано» означает отсутствие сведений на собранных официальных страницах, а не отсутствие функции.
Официальные API-цены на 2026-09-13. Допущение: 3 000 input + 200 output токенов на каждого судью и пару. Базовое жюри: Claude Sonnet 5 ($8.00) + Gemini 3.1 Pro Preview ($8.40) + Grok 4.20 reasoning ($4.25) ≈ $20.65 за 1 000 пар. При 150–200 новых пар в неделю (×4.33 ≈ 650–866 пар/мес) это ≈ $13.4–17.9/мес; с арбитром на ~30% конфликтов, ретраями и canary — ориентир $15–20/мес. Не включены платформа разметки, хранение, embeddings/NLI и человеческое время; оценка чувствительна к длине reasoning-вывода и preview-статусу моделей.
Замена Grok на deepseek-reasoner снижает сумму до ≈$18.49/1 000 пар, но DeepSeek генерирует сами подсказки — возможна self-preference, поэтому он предложен только как challenger.
| Тип | Тема | Работа | Количественная деталь | Уровень | Источник |
|---|---|---|---|---|---|
| блог | active learning | Human-in-the-Loop Feedback Loops for LLM Systems | 6–7% of the human annotation effort | низк. | jatinbansal.com |
| блог | active learning | Human-in-the-Loop Feedback Loops for LLM Systems | 50–80% reductions in annotation budget | низк. | jatinbansal.com |
| блог | active learning | ActiveLab: Active Learning with Data Re-Labeling | 5x fewer total annotations | низк. | cleanlab.ai |
| блог | жюри и судьи | Quorum Drift: When LLM Juries Slowly Converge on Wrong Answers | 43.2% consensus rate and mean panel variance of 1,753.6 on 7,063 Armalo jury judgments | низк. | trust.armalo.ai |
| блог | жюри и судьи | Weak judges, strong panel: an ensemble approach to LLM eval | crowd-vs-expert agreement of 72–83%, and expert-vs-expert at 79–90% | низк. | orq.ai |
| блог | жюри и судьи | We Audited Our Own LLM-Judge Panel | $0.002 a candidate | низк. | workloft.ai |
| блог | жюри и судьи | LLM Jury-on-Demand Framework | +6.7% accuracy gain in crowd-comparative evaluation, +11–22 points in violation detection | низк. | emergentmind.com |
| блог | жюри и судьи | Building a Reliable LLM-as-a-Judge: Bias and Calibration | over 80% agreement with humans | низк. | aitechconnect.in |
| блог | жюри и судьи | LLM-as-a-judge: when a model can grade another model | 50% error rates | низк. | artifipedia.com |
| блог | имплицитный фидбэк | The Thumbs-Up Button That Poisoned Your Eval Set Through the Back Door | fewer than five percent of responses receive any user feedback at all | низк. | tianpan.co |
| блог | синтетика и дрейф | Synthetic Eval Bootstrapping: How to Build Ground-Truth Datasets When You Have No Labeled Data | 68% to 94% | низк. | tianpan.co |
| блог | синтетика и дрейф | Synthetic Eval Bootstrapping: How to Build Ground-Truth Datasets When You Have No Labeled Data | 75.4 to 88.3 | низк. | tianpan.co |
| блог | синтетика и дрейф | Synthetic Eval Bootstrapping: How to Build Ground-Truth Datasets When You Have No Labeled Data | 560,000 | низк. | tianpan.co |
| исследование | active learning | Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection | \28 for GPT-5.2 Batch API vs. \316 for Prolific | низк. | arcxiv.org |
| исследование | active learning | Which Examples to Annotate for In-Context Learning? Towards Effective and Efficient Selection | 4.4% accuracy points | сред. | arxiv.org |
| исследование | active learning | Revisiting Active Learning under (Human) Label Variation | — | низк. | pith.science |
| исследование | active learning | Next Generation Active Learning: Mixture of LLMs in the Loop | — | выс. | ojs.aaai.org |
| исследование | active learning | LLMs Capture Emotion Labels, Not Emotion Uncertainty: Distributional Analysis and Calibration of Human-LLM Judgment Gaps | — | сред. | arxiv.org |
| исследование | жюри и судьи | Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need? | 0.17 to 0.86 of weighted kappa across 21 evaluation dimensions | сред. | arxiv.org |
| исследование | жюри и судьи | Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need? | 0.17 to 0.86 of weighted kappa | сред. | arxiv.org |
| исследование | жюри и судьи | Automating Annotation Guideline Improvements using LLMs: A Case Study | inter-annotator agreement going from 0.593 to 0.84 when improving WNUT-17’s guidelines | выс. | aclanthology.org |
| исследование | жюри и судьи | JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification | FN-only correlations 0.402 and 0.368; lifts 3.13x and 18.13x | сред. | arxiv.org |
| исследование | жюри и судьи | Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks | from 0.64 to 0.71 | выс. | aclanthology.org |
| исследование | жюри и судьи | More Human, More Efficient: Aligning Annotations with Quantized SLMs | 0.5774 | сред. | arxiv.org |
| исследование | жюри и судьи | Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring | +10.9 pp | сред. | arxiv.org |
| исследование | жюри и судьи | Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences | n=15 practitioners completing 6 tasks yielding 131 evaluations; judge strategy impact on number of evaluations F(1, 14) | сред. | arxiv.org |
| исследование | жюри и судьи | Margin-Adaptive Confidence Ranking for Reliable LLM Judgement | — | сред. | arxiv.org |
| исследование | жюри и судьи | Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement | reducing ECE by 50% and improving AUROC by 13% for GPT-4 | сред. | arxiv.org |
| исследование | жюри и судьи | CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation | 27-40% variance reduction at budget E=1 | сред. | arxiv.org |
| исследование | жюри и судьи | Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring | 49 annotated prompt/response pairs; 1.7B-7B parameters; agreement rates range from 51.5% (mistral:7b) to 69.1% (phi4-min | сред. | arxiv.org |
| исследование | жюри и судьи | Multi-Source Emotion Annotation in Children’s Language: When LLM Consensus Diverges from Human Judgment | Dawid-Skene estimates GPT-5.2 valence accuracy at 90.7%, whereas evaluation against human gold yields 71.0% | выс. | lrec.elra.info |
| исследование | жюри и судьи | No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding | 54% of the question-response pairs received unanimously consistent annotations | сред. | arxiv.org |
| исследование | жюри и судьи | Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why | 24 papers; 11 graded or continuous evaluations; 7 pairwise-preference evaluations; 3 binary-verdict evaluations; 2 judge | сред. | arxiv.org |
| исследование | жюри и судьи | Multi-Expert Conformal Risk Control for Pairwise LLM Judging in Open-Ended Dialogue | Acceptance Rate .973, Accuracy .887, Risk .098 for Decision Voting on ESConv (Homogeneous) | сред. | arxiv.org |
| исследование | жюри и судьи | A Unified Framework for Rubric-Based LLM Evaluation | 5,586 individual criterion-level judgments (931 criteria 3 systems 2 judges) | сред. | arxiv.org |
| исследование | жюри и судьи | A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges | 3.5 and 2.6 percentage-point gains in accuracy and precision, respectively, over majority voting | сред. | arxiv.org |
| исследование | жюри и судьи | Auditing and Debiasing LLM-as-a-Judge Evaluation with a Bias-Aware Panel | more than 1,000,000 simulated pairwise judgments; single judge agreement 87.6%; 0.73-rate position, verbosity, and self | низк. | ijisae.org |
| исследование | жюри и судьи | Mitigating Biases to Embrace Diversity: A Comprehensive Annotation Benchmark for Toxic Language | — | выс. | aclanthology.org |
| исследование | жюри и судьи | Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines | 71.0%, kappa=0.549 | сред. | arxiv.org |
| исследование | жюри и судьи | Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators | MAE shifts from 1.62 to 1.16 on HANNA and from 0.78 to 0.86 on SummEval | сред. | arxiv.org |
| исследование | жюри и судьи | Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels | neff≈2.0–2.5; Condorcet gap is 8–22 percentage points (pp; permutation p<10^-4) | сред. | arxiv.org |
| исследование | жюри и судьи | Quantifying and Mitigating Self-Preference Bias of LLM Judges | achieving a twofold reduction in prediction error compared to uncalibrated baselines | сред. | arxiv.org |
| исследование | жюри и судьи | Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm | r = 0.050 on 2,565 cells | сред. | arxiv.org |
| исследование | жюри и судьи | LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora | 2,377 essays, 12 judges, 4 providers, 5 version contrasts, Judge severity spans 219 points on ENEM's 0-1000 scale, panel | сред. | arxiv.org |
| исследование | жюри и судьи | Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis | Cohen’s Khuman×gemini = .75, Khuman×qwen = .68, Khuman×kimi = .56, Khuman×jury = .74 | сред. | arxiv.org |
| исследование | жюри и судьи | JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification | — | сред. | arxiv.org |
| исследование | жюри и судьи | PERSONAJUDGE: Simulating Individual Human Preference Judgments with Evaluator-Specific Demonstration Data | up to 9.9 points | низк. | pith.science |
| исследование | жюри и судьи | Counsel: A Meta-Evaluation Dataset for Agentic Tasks | Krippendorff’s alpha 0.78; strongest judge achieving 88% agreement on location and 65% on reasoning | сред. | arxiv.org |
| исследование | жюри и судьи | aipsy-judge: A Specialized, Psychologist-Corrected Local Judge for the Psychological Safety of Conversational AI | Gemini kappa=.75, jury kappa=.74 | сред. | arxiv.org |
| исследование | жюри и судьи | Calibrate, Don’t Curate: Label-Efficient Estimation from Noisy LLM Judges | On RewardBench2, retaining all judges achieves negative log-likelihood (NLL) of versus under top-5 selection, halving th | сред. | arxiv.org |
| исследование | жюри и судьи | When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability | — | низк. | pith.science |
| исследование | жюри и судьи | LLMEval-3: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models | 90% agreement with human experts | сред. | arxiv.org |
| исследование | имплицитный фидбэк | What sentiment analysis can't see: Measuring whether customers were helped, and what went wrong, across 70,000 support conversations | 70,450 support conversations; correlation of 0.47 versus 0.36; tone and satisfaction disagreed in 44% of conversations | сред. | arxiv.org |
| исследование | имплицитный фидбэк | User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal | 48% | низк. | alphaxiv.org |
| исследование | синтетика и дрейф | Who Drifted: the System or the Judge? Anytime-Valid Attribution in LLM Evaluation Pipelines | On two real judge changes, a silent version bump is detected as judge drift in 60/60 runs with zero judge-to-system misa | сред. | arxiv.org |
| исследование | слабая супервизия | LabelAId: Just-in-time AI Interventions for Improving Human Labeling Quality and Domain Knowledge in Crowdsourcing Systems | 36.7% | сред. | arxiv.org |
| исследование | слабая супервизия | Metadata Predictability Is Not Evidence Dependence: An Intervention-Based Audit for Weak-Label Benchmarks | 0.643 | сред. | arxiv.org |
| исследование | слабая супервизия | [2606.05414] When Evidence is Sparse: Weakly Supervised Early Failure Alerting in Dialogs and LLM-Agent Trajectories | 1-10% | низк. | academ.us |
| исследование | слабая супервизия | When Evidence is Sparse: Weakly Supervised Early Failure Alerting in Dialogs and LLM-Agent Trajectories | 1-10% | сред. | arxiv.org |
| исследование | слабая супервизия | Training Subset Selection for Weak Supervision | up to 19% (absolute) on benchmark tasks | сред. | arxiv.org |
Уровень подтверждения: выс. — peer-reviewed (ACL/AAAI/LREC); сред. — один arXiv-препринт; низк. — инженерный блог или вторичный источник. Цифры из блогов не переносятся как установленные факты.
Метод: 60 строк исследований и блогов (каждая с URL и хостом; уровень подтверждения присвоен по типу площадки: peer-reviewed, arXiv-препринт, блог), 20 официальных страниц инструментов, сведённых к 10 продуктам, и 22 строки официальных API-цен; всё собрано 2026-09-13. Стоимость за 1 000 пар рассчитана из опубликованных цен при допущении 3 000 input + 200 output токенов на судью; калькулятор пересчитывает формулу, а не фиксированный счёт. Пороги κ/α — внутренние предрегистрированные ориентиры команды, не универсальный стандарт. Для места сокращены поля limitation и даты публикаций отдельных работ; отсутствующие в официальных страницах цены отмечены «публично не указано».