You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Metrics judged inside ADK (rubric_based_*, hallucinations_v1, final_response_match_v2) return the judge's reasoning in rubric_scores, so a failed case explains itself.
Metrics judged by the Vertex AI Gen AI evaluation service (multi_turn_task_success_v1, multi_turn_tool_use_quality_v1, multi_turn_trajectory_quality_v1, safety_v1, response_evaluation_score) return only a number.
All of them go through _VertexAiEvalFacade._get_score() in src/google/adk/evaluation/vertex_ai_eval_facade.py, which reads only summary_metrics[0].mean_score. The Vertex response also contains, per metric, explanation and (for the adaptive-rubric metrics) rubric_verdicts, with each generated rubric, its verdict and its reasoning. ADK discards them.
This hurts most with the multi_turn_*_v1 metrics: Vertex generates new rubrics on every run, so when a case scores 0.8 there is no way to know which criterion failed or why without calling Vertex again by hand. In our runs, several failures turned out to be spurious (for example a conditional rubric like "if the agent calls X to verify..." failed when the condition never happened), and we could only find that out after recovering the verdicts.
It also hides the reason behind safety_v1 failures. With the change below applied locally, a correct customer-service answer that only asks the user for their phone line number scored 0.0, and the discarded explanation was "Violated policies: PII & Demographic Data". Without it, the result is a bare 0.0.
Checked on google-adk 2.10.0 and on main: main already fixes the content-less event issue ("fix(eval): skip content-less events when mapping Vertex multi-turn turns"), but still reads only the mean score.
Describe the Solution You'd Like
For Vertex-backed metrics, map what Vertex returns into the existing ADK result types, the same way ADK-native judges do:
each item in metric_results[<metric>].rubric_verdicts becomes a RubricScore in PerInvocationResult.rubric_scores and EvaluationResult.overall_rubric_scores:
rubric_id: the rubric description, or its id when there is no description
rationale: verdict.reasoning
score: 1.0 if verdict.verdict else 0.0
metric_results[<metric>].explanation (and error_message when there is no score) is kept as well, e.g. as an extra RubricScore with only rationale, or in a dedicated field.
No config change needed; adk eval --print_detailed_results and the saved evalset_result.json would then show the criteria like they already do for rubric_based_*.
Impact on your work
We use multi_turn_task_success_v1 and multi_turn_tool_use_quality_v1 as the accuracy and tool-selection metrics of our multi-turn and user-simulation suites. Without the verdicts, every failure needs manual investigation, and we cannot tell a real regression from a spurious generated rubric. Not blocking: we have a workaround (below), but we'd like to delete it.
Willingness to contribute
Are you interested in implementing this feature yourself or submitting a PR?
(Yes)
馃煛 Recommended Information
Describe Alternatives You've Considered
A custom metric registered under the same name in custom_metrics, which calls vertexai.Client().evals.evaluate(...) directly with the same AgentData, and maps rubric_verdicts and explanation into rubric_scores. It works, but it duplicates the facade logic and has to track changes in both ADK and the vertexai SDK.
Proposed API / Implementation
Tested locally by patching the installed google-adk 2.10.0 (plus the content-less event fix from main) and running adk eval with no custom metrics:
Then pass rubric_scores=self._get_rubric_scores(eval_case_result) to the PerInvocationResult built by _SingleTurnVertexAiEvalFacade and _MultiTurnVertexiAiEvalFacade, and overall_rubric_scores= to the EvaluationResult of the multi-turn facade.
safety_v1 (correct answer asking for the phone line)
0.0
Violated policies: PII & Demographic Data
response_evaluation_score
N/A
400 INVALID_ARGUMENT ... Error parsing JSON (see below)
Open points for the PR:
The single-turn facade builds no overall_rubric_scores. Explanations show per invocation only. Aggregating them (or leaving them per invocation) needs a decision.
Unit tests would mock _perform_eval with a result carrying rubric_verdicts and explanation.
Side finding: with the change applied, response_evaluation_score turned out to fail on every call with 400 INVALID_ARGUMENT ("Error parsing JSON ... Input: {Evaluation Steps: ..."), apparently while Vertex parses its own judge output. Today ADK reports it only as N/A. Happy to file it separately.
** Please make sure you read the contribution guide and file the issues in the right place. **
Contribution guide.
馃敶 Required Information
Is your feature request related to a specific problem?
The metrics listed on https://adk.dev/evaluate/criteria/ return very different levels of detail depending on where the judge runs:
rubric_based_*,hallucinations_v1,final_response_match_v2) return the judge's reasoning inrubric_scores, so a failed case explains itself.multi_turn_task_success_v1,multi_turn_tool_use_quality_v1,multi_turn_trajectory_quality_v1,safety_v1,response_evaluation_score) return only a number.All of them go through
_VertexAiEvalFacade._get_score()insrc/google/adk/evaluation/vertex_ai_eval_facade.py, which reads onlysummary_metrics[0].mean_score. The Vertex response also contains, per metric,explanationand (for the adaptive-rubric metrics)rubric_verdicts, with each generated rubric, its verdict and its reasoning. ADK discards them.This hurts most with the
multi_turn_*_v1metrics: Vertex generates new rubrics on every run, so when a case scores 0.8 there is no way to know which criterion failed or why without calling Vertex again by hand. In our runs, several failures turned out to be spurious (for example a conditional rubric like "if the agent calls X to verify..." failed when the condition never happened), and we could only find that out after recovering the verdicts.It also hides the reason behind
safety_v1failures. With the change below applied locally, a correct customer-service answer that only asks the user for their phone line number scored0.0, and the discarded explanation was"Violated policies: PII & Demographic Data". Without it, the result is a bare0.0.Checked on google-adk 2.10.0 and on
main:mainalready fixes the content-less event issue ("fix(eval): skip content-less events when mapping Vertex multi-turn turns"), but still reads only the mean score.Describe the Solution You'd Like
For Vertex-backed metrics, map what Vertex returns into the existing ADK result types, the same way ADK-native judges do:
metric_results[<metric>].rubric_verdictsbecomes aRubricScoreinPerInvocationResult.rubric_scoresandEvaluationResult.overall_rubric_scores:rubric_id: the rubric description, or its id when there is no descriptionrationale:verdict.reasoningscore:1.0ifverdict.verdictelse0.0metric_results[<metric>].explanation(anderror_messagewhen there is no score) is kept as well, e.g. as an extraRubricScorewith onlyrationale, or in a dedicated field.No config change needed;
adk eval --print_detailed_resultsand the savedevalset_result.jsonwould then show the criteria like they already do forrubric_based_*.Impact on your work
We use
multi_turn_task_success_v1andmulti_turn_tool_use_quality_v1as the accuracy and tool-selection metrics of our multi-turn and user-simulation suites. Without the verdicts, every failure needs manual investigation, and we cannot tell a real regression from a spurious generated rubric. Not blocking: we have a workaround (below), but we'd like to delete it.Willingness to contribute
Are you interested in implementing this feature yourself or submitting a PR?
(Yes)
馃煛 Recommended Information
Describe Alternatives You've Considered
A custom metric registered under the same name in
custom_metrics, which callsvertexai.Client().evals.evaluate(...)directly with the sameAgentData, and mapsrubric_verdictsandexplanationintorubric_scores. It works, but it duplicates the facade logic and has to track changes in both ADK and thevertexaiSDK.Proposed API / Implementation
Tested locally by patching the installed google-adk 2.10.0 (plus the content-less event fix from
main) and runningadk evalwith no custom metrics:Then pass
rubric_scores=self._get_rubric_scores(eval_case_result)to thePerInvocationResultbuilt by_SingleTurnVertexAiEvalFacadeand_MultiTurnVertexiAiEvalFacade, andoverall_rubric_scores=to theEvaluationResultof the multi-turn facade.Results of the local run:
multi_turn_task_success_v1(static multi-turn case)multi_turn_task_success_v1(user-simulation case)multi_turn_tool_use_quality_v1(user-simulation case)safety_v1(correct answer asking for the phone line)Violated policies: PII & Demographic Dataresponse_evaluation_score400 INVALID_ARGUMENT ... Error parsing JSON(see below)Open points for the PR:
overall_rubric_scores. Explanations show per invocation only. Aggregating them (or leaving them per invocation) needs a decision._perform_evalwith a result carryingrubric_verdictsandexplanation.Additional Context
GOOGLE_CLOUD_LOCATION=global.main.RubricBasedEvaluatordrops per-invocation rubric verdicts). That one is about ADK-native rubric judges; this request is about the Vertex-backed metrics invertex_ai_eval_facade.py, which never read the verdicts at all.response_evaluation_scoreturned out to fail on every call with400 INVALID_ARGUMENT("Error parsing JSON ... Input: {Evaluation Steps: ..."), apparently while Vertex parses its own judge output. Today ADK reports it only as N/A. Happy to file it separately.