Skip to content

Vertex-backed eval metrics drop rubric verdicts and explanations (multi_turn_*_v1, safety_v1)聽#7350

Description

@DalmPhilipe

** Please make sure you read the contribution guide and file the issues in the right place. **
Contribution guide.

馃敶 Required Information

Is your feature request related to a specific problem?

The metrics listed on https://adk.dev/evaluate/criteria/ return very different levels of detail depending on where the judge runs:

  • Metrics judged inside ADK (rubric_based_*, hallucinations_v1, final_response_match_v2) return the judge's reasoning in rubric_scores, so a failed case explains itself.
  • Metrics judged by the Vertex AI Gen AI evaluation service (multi_turn_task_success_v1, multi_turn_tool_use_quality_v1, multi_turn_trajectory_quality_v1, safety_v1, response_evaluation_score) return only a number.

All of them go through _VertexAiEvalFacade._get_score() in src/google/adk/evaluation/vertex_ai_eval_facade.py, which reads only summary_metrics[0].mean_score. The Vertex response also contains, per metric, explanation and (for the adaptive-rubric metrics) rubric_verdicts, with each generated rubric, its verdict and its reasoning. ADK discards them.

This hurts most with the multi_turn_*_v1 metrics: Vertex generates new rubrics on every run, so when a case scores 0.8 there is no way to know which criterion failed or why without calling Vertex again by hand. In our runs, several failures turned out to be spurious (for example a conditional rubric like "if the agent calls X to verify..." failed when the condition never happened), and we could only find that out after recovering the verdicts.

It also hides the reason behind safety_v1 failures. With the change below applied locally, a correct customer-service answer that only asks the user for their phone line number scored 0.0, and the discarded explanation was "Violated policies: PII & Demographic Data". Without it, the result is a bare 0.0.

Checked on google-adk 2.10.0 and on main: main already fixes the content-less event issue ("fix(eval): skip content-less events when mapping Vertex multi-turn turns"), but still reads only the mean score.

Describe the Solution You'd Like

For Vertex-backed metrics, map what Vertex returns into the existing ADK result types, the same way ADK-native judges do:

  • each item in metric_results[<metric>].rubric_verdicts becomes a RubricScore in PerInvocationResult.rubric_scores and EvaluationResult.overall_rubric_scores:
    • rubric_id: the rubric description, or its id when there is no description
    • rationale: verdict.reasoning
    • score: 1.0 if verdict.verdict else 0.0
  • metric_results[<metric>].explanation (and error_message when there is no score) is kept as well, e.g. as an extra RubricScore with only rationale, or in a dedicated field.

No config change needed; adk eval --print_detailed_results and the saved evalset_result.json would then show the criteria like they already do for rubric_based_*.

Impact on your work

We use multi_turn_task_success_v1 and multi_turn_tool_use_quality_v1 as the accuracy and tool-selection metrics of our multi-turn and user-simulation suites. Without the verdicts, every failure needs manual investigation, and we cannot tell a real regression from a spurious generated rubric. Not blocking: we have a workaround (below), but we'd like to delete it.

Willingness to contribute

Are you interested in implementing this feature yourself or submitting a PR?
(Yes)


馃煛 Recommended Information

Describe Alternatives You've Considered

A custom metric registered under the same name in custom_metrics, which calls vertexai.Client().evals.evaluate(...) directly with the same AgentData, and maps rubric_verdicts and explanation into rubric_scores. It works, but it duplicates the facade logic and has to track changes in both ADK and the vertexai SDK.

Proposed API / Implementation

Tested locally by patching the installed google-adk 2.10.0 (plus the content-less event fix from main) and running adk eval with no custom metrics:

# vertex_ai_eval_facade.py, in _VertexAiEvalFacade
def _get_rubric_scores(self, eval_result: object) -> list[RubricScore]:
  scores = []
  for case in getattr(eval_result, "eval_case_results", None) or []:
    for candidate in getattr(case, "response_candidate_results", None) or []:
      results = getattr(candidate, "metric_results", None) or {}
      result = results.get(str(self._metric_name)) or results.get(
          getattr(self._metric_name, "name", None)
      )
      if result is None and len(results) == 1:
        result = next(iter(results.values()))
      if result is None:
        continue
      for v in getattr(result, "rubric_verdicts", None) or []:
        rubric = getattr(v, "evaluated_rubric", None)
        text = getattr(
            getattr(getattr(rubric, "content", None), "property", None),
            "description",
            None,
        )
        scores.append(
            RubricScore(
                rubric_id=text or getattr(rubric, "rubric_id", None) or "rubric",
                rationale=getattr(v, "reasoning", None),
                # Vertex omits `verdict` when a rubric fails.
                score=1.0 if getattr(v, "verdict", None) else 0.0,
            )
        )
      explanation = getattr(result, "explanation", None) or getattr(
          result, "error_message", None
      )
      if explanation:
        scores.append(RubricScore(rubric_id="explanation", rationale=explanation))
  return scores

Then pass rubric_scores=self._get_rubric_scores(eval_case_result) to the PerInvocationResult built by _SingleTurnVertexAiEvalFacade and _MultiTurnVertexiAiEvalFacade, and overall_rubric_scores= to the EvaluationResult of the multi-turn facade.

Results of the local run:

Metric Score What the change surfaced
multi_turn_task_success_v1 (static multi-turn case) 1.0 3 generated rubrics with reasoning
multi_turn_task_success_v1 (user-simulation case) 1.0 5 generated rubrics
multi_turn_tool_use_quality_v1 (user-simulation case) 1.0 11 generated rubrics
safety_v1 (correct answer asking for the phone line) 0.0 Violated policies: PII & Demographic Data
response_evaluation_score N/A 400 INVALID_ARGUMENT ... Error parsing JSON (see below)

Open points for the PR:

  • The single-turn facade builds no overall_rubric_scores. Explanations show per invocation only. Aggregating them (or leaving them per invocation) needs a decision.
  • Unit tests would mock _perform_eval with a result carrying rubric_verdicts and explanation.

Additional Context

  • google-adk 2.10.0, Python 3.11, GOOGLE_CLOUD_LOCATION=global.
  • Related: the content-less event fix already on main.
  • Related, but a different code path: RubricBasedEvaluator drops per-invocation rubric verdicts when invocations carry different rubrics聽#7301 (RubricBasedEvaluator drops per-invocation rubric verdicts). That one is about ADK-native rubric judges; this request is about the Vertex-backed metrics in vertex_ai_eval_facade.py, which never read the verdicts at all.
  • Side finding: with the change applied, response_evaluation_score turned out to fail on every call with 400 INVALID_ARGUMENT ("Error parsing JSON ... Input: {Evaluation Steps: ..."), apparently while Vertex parses its own judge output. Today ADK reports it only as N/A. Happy to file it separately.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

eval[Component] This issue is related to evaluationneeds review[Status] The PR/issue is awaiting review from the maintainer

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions