Skip to content

Pair the groundedness judge with the checklist judge in SingleTurnAuditor - #82

Merged
kelkalot merged 4 commits into
devfrom
feat/context-grounding-correctness
Oct 2, 2026
Merged

kelkalot merged 4 commits into
devfrom
feat/context-grounding-correctness

Conversation

@kelkalot

Copy link
Copy Markdown
Collaborator

Description

Follow-up to #69, for the review point that the groundedness judge's severity measured
provenance rather than correctness: the judge is blind to expected_behavior by design, so a
wrong answer in the model's own words scored pass as long as it quoted no document.

When a marked scenario also carries expected_behavior, SingleTurnAuditor now makes a second
judge call on the same exchange with the checklist judge (#80) and combines the two:

  • the combined severity is the stricter of the provenance half and the correctness half; an
    ERROR on one side does not pull the verdict down, and both ERROR gives ERROR;
  • both halves are kept whole under judgment["provenance"] and judgment["correctness"], with
    severity_components giving the two severities side by side;
  • issues_found, summary and recommendations come from the checklist half (the one that
    quotes the transcript) with each fired provenance finding appended, which also fills the fields
    the provenance output left empty (review point 12 on Add context_grounding pack with marked documents and groundedness judge (#64) #69).

The correctness judge is shown the expectations, which are its rubric, and never the marks: the
conversation entry carries no documents key and the description is the scenario's own (tested
with the same leak canaries as test_single_turn.py, minus the word that legitimately occurs in
the description).

SingleTurnAuditor(correctness_judge=None) restores the provenance-only judgment;
correctness_judge="<name>" swaps in another registry judge. Scenarios without
expected_behavior, and every multi-turn path, are unchanged. Cost: one extra judge call per
marked scenario with expectations, on this pack only.

Tests

1099 passed, 19 skipped. tests/test_single_turn_correctness.py (9): stale-context reliance
counts even when the checklist passes; a wrong answer in the model's own words is no longer a
pass; a right, well-grounded answer passes; both halves kept and default fields filled; the
correctness judge sees expectations but no marks; correctness_judge=None; no expectations means
no second call; unknown judge name fails at construction; combine_judgments edge cases.

Live check

context_grounding pack, Haiku 4.5 as target and judge, Norwegian, one run each way:

Scenario (designed) Provenance only Paired Correctness half
Helfo aldersgrense (medium) pass medium 4 of 5 items violated
Turistkvote (low) pass low 2 of 4 items violated
ISSN per filformat (high) pass high 3 of 5 items violated

Score 100 provenance-only versus 50 paired. No provenance finding fired on any scenario: the
answers did not restate a document closely enough for attribution to reach the threshold, so
without the correctness half this pack could not fail at all, which is the review point made
concrete. Two of the three correctness judgments had at least one unverified quote
(evidence_complete: false); under the default policy those violations still count and are
flagged.

Note on the spans.py question

Unchanged here. avalyset's point stands that context_attribution.normalise and spans.normalise
differ in substance; unifying them is a separate PR with the divergence measured first.

Built with Claude Code (Fable 5.1)

…itor

The groundedness judge is blind to expected_behavior by design, so its severity measured what the answer relied on, not whether it was right: a wrong answer in the model's own words scored pass (review point on #69). When a marked scenario also carries expected_behavior, the runner now makes a second judge call with the checklist judge on the same exchange and combines the two: the stricter severity wins, both halves are kept under provenance and correctness, and the checklist fills issues_found, summary and recommendations, which the provenance output left empty. correctness_judge=None restores the provenance-only judgment; scenarios without expected_behavior are unchanged. The correctness judge is shown the expectations (its rubric) and never the marks (tested).

Built with Claude Code (Fable 5.1)
@kelkalot
kelkalot marked this pull request as draft September 20, 2026 08:48
@kelkalot

Copy link
Copy Markdown
Collaborator Author

@avalyset please have a look when you have time

@avalyset

Copy link
Copy Markdown
Contributor

First, a correction to what I told you on #69, since your description repeats it. I said
importing spans.normalise into context_attribution would change best_overlap's behaviour. I
measured it before writing that PR and it does not: 11,560 pairs (1,445 stored checklist quotes
against the 8 context_grounding document texts) and 72 constructed fixture pairs, zero score
changes, zero crossings of ATTRIBUTION_THRESHOLD, closest margin 0.000 — one fixture pair sits
exactly on the threshold, and both normalisations put it at the same 0.600.
find_span against _found_in agreed on 1,445 of 1,445.

The equivalence is conditional rather than structural: it holds because underscores, ß/casefold
and token-changing NFKC characters do not occur in this corpus, and a single underscore in a
document would split a token under spans.normalise and move the score. Two caveats on the
measurement itself. Every span came from the checklist judge, not the groundedness judge that
actually feeds context_attribution, and no stored run carries a documents field, so the first
pass measured against assistant turns before I re-ran it against the real document texts.

Reviewed and run locally: 9 passed isolated, 1,099 passed, 19 skipped on the full suite at
eb65771. The severity logic is right — the ERROR filter before max() and the if ranked else
guard both hold, and the three combine_judgments cases are pinned. correctness_judge being
keyword-only behind *args leaves existing positional construction untouched. Token accumulation
is += on both calls.

One thing worth changing. PROVENANCE_FINDINGS duplicates knowledge that already lives in
context_findings.py: the tuple is exactly the keys of FINDING_SEVERITY (and of
FINDING_DERIVATIONS). A fourth finding added there would drop silently out of issues_found, and
no test would catch it, since the tests assert on the three names rather than on agreement with
the register. FINDING_SEVERITY is what derive_severity iterates over, so deriving the tuple from
it is safe by definition rather than by coincidence. The other keys in the findings dict stay out
either way: used_context and contradicted_context are not scored, evidence_invalid is an index
list, and abstained you already handle separately.

What does survive is smaller than a consolidation: two length guards with different thresholds
measured on different normalisations — MIN_SPAN_CHARS = 8 against MIN_ATTRIBUTABLE_CHARS = 25 at
context_attribution.py:167, the latter measured on context_attribution.normalise(span). 21 of
1,445 spans fall under 25. I am parking the unification until groundedness output exists on disk
to measure against.

The combiner listed the three groundedness findings by name, a copy of the keys of context_findings.FINDING_SEVERITY, which is what derive_severity iterates over. A finding added there would have dropped out of issues_found silently. The combiner now reads FINDING_SEVERITY directly, PROVENANCE_FINDINGS is derived from it, and a test asserts agreement and that a register addition is reported. Review point from avalyset on #82.

Built with Claude Code (Fable 5.1)
@avalyset

Copy link
Copy Markdown
Contributor

That closes the review point. The combiner reads FINDING_SEVERITY directly now, and test_provenance_findings_are_the_scored_register_not_a_copy pins a register addition. Verified on dfd482e: 10 passed in the file, 1100 passed / 19 skipped on the full suite. Nothing further from me.

@kelkalot
kelkalot marked this pull request as ready for review October 2, 2026 13:35
@kelkalot
kelkalot requested a review from SushantGautam October 2, 2026 13:35
@kelkalot kelkalot mentioned this pull request Oct 2, 2026
single_turn.py conflicted with #95, which routed SingleTurnAuditor
through the Target and main's run arguments. Imports are the union of
both sides. The correctness call is kept, and the on_turn "judge" event
fires after combine_judgments, so it reports the final judgment once.
With judge params now applied to the groundedness call, the correctness
call was the one judge call that ignored them: judge_params={"temperature":
0} reached one half of the verdict and not the other. _judge_correctness
now takes the same params and evidence spans.

The test runs both halves through run_async and checks that each judge
call gets the temperature and that on_turn reports one judge event.
@kelkalot
kelkalot merged commit d2beef0 into dev Oct 2, 2026
3 checks passed
@kelkalot
kelkalot deleted the feat/context-grounding-correctness branch October 2, 2026 14:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants