Pair the groundedness judge with the checklist judge in SingleTurnAuditor - #82
Conversation
…itor The groundedness judge is blind to expected_behavior by design, so its severity measured what the answer relied on, not whether it was right: a wrong answer in the model's own words scored pass (review point on #69). When a marked scenario also carries expected_behavior, the runner now makes a second judge call with the checklist judge on the same exchange and combines the two: the stricter severity wins, both halves are kept under provenance and correctness, and the checklist fills issues_found, summary and recommendations, which the provenance output left empty. correctness_judge=None restores the provenance-only judgment; scenarios without expected_behavior are unchanged. The correctness judge is shown the expectations (its rubric) and never the marks (tested). Built with Claude Code (Fable 5.1)
|
@avalyset please have a look when you have time |
|
First, a correction to what I told you on #69, since your description repeats it. I said The equivalence is conditional rather than structural: it holds because underscores, ß/casefold Reviewed and run locally: 9 passed isolated, 1,099 passed, 19 skipped on the full suite at One thing worth changing. PROVENANCE_FINDINGS duplicates knowledge that already lives in What does survive is smaller than a consolidation: two length guards with different thresholds |
The combiner listed the three groundedness findings by name, a copy of the keys of context_findings.FINDING_SEVERITY, which is what derive_severity iterates over. A finding added there would have dropped out of issues_found silently. The combiner now reads FINDING_SEVERITY directly, PROVENANCE_FINDINGS is derived from it, and a test asserts agreement and that a register addition is reported. Review point from avalyset on #82. Built with Claude Code (Fable 5.1)
|
That closes the review point. The combiner reads |
single_turn.py conflicted with #95, which routed SingleTurnAuditor through the Target and main's run arguments. Imports are the union of both sides. The correctness call is kept, and the on_turn "judge" event fires after combine_judgments, so it reports the final judgment once.
With judge params now applied to the groundedness call, the correctness
call was the one judge call that ignored them: judge_params={"temperature":
0} reached one half of the verdict and not the other. _judge_correctness
now takes the same params and evidence spans.
The test runs both halves through run_async and checks that each judge
call gets the temperature and that on_turn reports one judge event.
Description
Follow-up to #69, for the review point that the groundedness judge's severity measured
provenance rather than correctness: the judge is blind to
expected_behaviorby design, so awrong answer in the model's own words scored
passas long as it quoted no document.When a marked scenario also carries
expected_behavior,SingleTurnAuditornow makes a secondjudge call on the same exchange with the checklist judge (#80) and combines the two:
ERROR on one side does not pull the verdict down, and both ERROR gives ERROR;
judgment["provenance"]andjudgment["correctness"], withseverity_componentsgiving the two severities side by side;issues_found,summaryandrecommendationscome from the checklist half (the one thatquotes the transcript) with each fired provenance finding appended, which also fills the fields
the provenance output left empty (review point 12 on Add context_grounding pack with marked documents and groundedness judge (#64) #69).
The correctness judge is shown the expectations, which are its rubric, and never the marks: the
conversation entry carries no
documentskey and the description is the scenario's own (testedwith the same leak canaries as
test_single_turn.py, minus the word that legitimately occurs inthe description).
SingleTurnAuditor(correctness_judge=None)restores the provenance-only judgment;correctness_judge="<name>"swaps in another registry judge. Scenarios withoutexpected_behavior, and every multi-turn path, are unchanged. Cost: one extra judge call permarked scenario with expectations, on this pack only.
Tests
1099 passed, 19 skipped.
tests/test_single_turn_correctness.py(9): stale-context reliancecounts even when the checklist passes; a wrong answer in the model's own words is no longer a
pass; a right, well-grounded answer passes; both halves kept and default fields filled; the
correctness judge sees expectations but no marks;
correctness_judge=None; no expectations meansno second call; unknown judge name fails at construction;
combine_judgmentsedge cases.Live check
context_groundingpack, Haiku 4.5 as target and judge, Norwegian, one run each way:Score 100 provenance-only versus 50 paired. No provenance finding fired on any scenario: the
answers did not restate a document closely enough for attribution to reach the threshold, so
without the correctness half this pack could not fail at all, which is the review point made
concrete. Two of the three correctness judgments had at least one unverified quote
(
evidence_complete: false); under the default policy those violations still count and areflagged.
Note on the
spans.pyquestionUnchanged here. avalyset's point stands that
context_attribution.normaliseandspans.normalisediffer in substance; unifying them is a separate PR with the divergence measured first.
Built with Claude Code (Fable 5.1)