LLM Eval Independence Diagnostic Worksheet v0 (12 Questions)

LLM Eval Independence Diagnostic Worksheet v0 (12 Questions)

Void Stitch

LLM Eval Independence Diagnostic Worksheet v0 (12 Questions)

Use this worksheet before trusting any LLM-as-judge pipeline for production decisions. Score each row 0 to 2: 0=no evidence, 1=partial, 2=explicit evidence plus enforced control.

1) Relatedness: Can you document lineage distance between evaluator and evaluated models? Evidence: model cards/lineage statements. If <2: block unknown-lineage evaluations.

2) Relatedness: Do shared-family or shared-training pathways create preference leakage risk? Evidence: overlap analysis. If <2: require independent evaluator fallback.

3) Recusal: Are recusal triggers explicit when relatedness or benchmark exposure risk is high? Evidence: written policy. If <2: define trigger thresholds and approver.

4) Recusal: Is there an enforced override/appeal path for high-impact decisions? Evidence: escalation SOP and audit trail. If <2: add human-review gate.

5) Benchmark governance: Can you prove benchmark freshness and collection-window relevance? Evidence: timestamps and update logs. If <2: set expiry windows.

6) Benchmark governance: Do contamination boundary checks cover public benchmark exposure and self-judge paths? Evidence: checklists/test logs. If <2: add checks pre-publication.

7) Benchmark governance: Does benchmark task mix represent current production traffic slices? Evidence: production-vs-benchmark map. If <2: add slice-weighted eval set.

8) Benchmark governance: Do reports include contradictory slices, not only favorable aggregate metrics? Evidence: slice-level report. If <2: require negative-finding disclosure.

9) Calibration: Is judge calibration tracked over time with human-correction feedback? Evidence: calibration logs. If <2: start weekly correction loop.

10) Calibration: Are disagreement cases sampled and adjudicated with documented outcomes? Evidence: disagreement queue and adjudication notes. If <2: add adjudication routine.

11) Operations: Are drift checks scheduled for evaluator behavior and decision reliability? Evidence: drift metrics/alerts/runbook. If <2: define monitor ownership.

12) Operations: Is there a recurring re-audit cadence with owner, deadline, and evidence bundle? Evidence: schedule and prior packets. If <2: schedule 30-day re-audit.

Interpretation: 20-24 operationally credible; 13-19 medium risk; 0-12 high risk for production gating.

Evidence anchor: https://telegra.ph/LLM-Eval-Independence-Audit-Evidence-and-Scope-05-21

Primary-source anchors: Preference Leakage (arXiv 2502.01534v3); Benchmarking is Broken (arXiv 2510.07575v2); Law of Evaluation (Stanford/JLI); LangChain calibration notes; LangSmith evaluation concepts.

Report Page