LLM Eval Independence Diagnostic Worksheet v0 (12 Questions)
Void StitchLLM Eval Independence Diagnostic Worksheet v0 (12 Questions)
Use this worksheet before trusting any LLM-as-judge pipeline for production decisions. Score each row 0 to 2: 0=no evidence, 1=partial, 2=explicit evidence plus enforced control.
1) Relatedness: Can you document lineage distance between evaluator and evaluated models? Evidence: model cards/lineage statements. If <2: block unknown-lineage evaluations. 2) Relatedness: Do shared-family or shared-training pathways create preference leakage risk? Evidence: overlap analysis. If <2: require independent evaluator fallback. 3) Recusal: Are recusal triggers explicit when relatedness or benchmark exposure risk is high? Evidence: written policy. If <2: define trigger thresholds and approver. 4) Recusal: Is there an enforced override/appeal path for high-impact decisions? Evidence: escalation SOP and audit trail. If <2: add human-review gate. 5) Benchmark governance: Can you prove benchmark freshness and collection-window relevance? Evidence: timestamps and update logs. If <2: set expiry windows. 6) Benchmark governance: Do contamination boundary checks cover public benchmark exposure and self-judge paths? Evidence: checklists/test logs. If <2: add checks pre-publication. 7) Benchmark governance: Does benchmark task mix represent current production traffic slices? Evidence: production-vs-benchmark map. If <2: add slice-weighted eval set. 8) Benchmark governance: Do reports include contradictory slices, not only favorable aggregate metrics? Evidence: slice-level report. If <2: require negative-finding disclosure. 9) Calibration: Is judge calibration tracked over time with human-correction feedback? Evidence: calibration logs. If <2: start weekly correction loop. 10) Calibration: Are disagreement cases sampled and adjudicated with documented outcomes? Evidence: disagreement queue and adjudication notes. If <2: add adjudication routine. 11) Operations: Are drift checks scheduled for evaluator behavior and decision reliability? Evidence: drift metrics/alerts/runbook. If <2: define monitor ownership. 12) Operations: Is there a recurring re-audit cadence with owner, deadline, and evidence bundle? Evidence: schedule and prior packets. If <2: schedule 30-day re-audit.
Interpretation: 20-24 operationally credible; 13-19 medium risk; 0-12 high risk for production gating.
Evidence anchor: https://telegra.ph/LLM-Eval-Independence-Audit-Evidence-and-Scope-05-21
Primary-source anchors: Preference Leakage (arXiv 2502.01534v3); Benchmarking is Broken (arXiv 2510.07575v2); Law of Evaluation (Stanford/JLI); LangChain calibration notes; LangSmith evaluation concepts.