Skip to content

Built-in google_search calls are not recorded as OpenTelemetry span attributes (grounding_metadata is ignored by telemetry) #7345

Description

@spandankeche

🔴 Required Information

Is your feature request related to a specific problem?

google_search is a built-in, server-side tool: the search runs inside the Gemini call, so there is no local function call and no execute_tool span. That part is expected.

However, the only record of the search is LlmResponse.grounding_metadata, and nothing in google/adk/telemetry/ reads it:

  • Content capture on (ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS, default on): the search appears only inside the serialized gcp.vertex.agent.llm_response JSON on the call_llm span. There's no attribute to query or filter on.
  • Content capture off (ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS=false, which adk deploy agent_engine --otel_to_cloud sets by default, see cli_deploy.py): nothing in any span shows a search happened.

The data is kept on the session event (event.grounding_metadata), so it is not lost. The gap is only in traces, which is where people debug production agents. Today you can't answer "did this answer come from a search, and what was searched?" from traces alone.

Repro, live on Vertex AI (InMemoryRunner + LlmAgent(tools=[google_search]); spans captured with InMemorySpanExporter; GOOGLE_GENAI_USE_VERTEXAI=TRUE, location global):

Model Content capture grounding_metadata on event Search visible in span attributes?
gemini-2.5-flash on 2 web_search_queries, 2 chunks only inside the gcp.vertex.agent.llm_response JSON
gemini-2.5-flash off 2 web_search_queries, 2 chunks no
gemini-3.5-flash on 9 web_search_queries, 3 chunks only inside the gcp.vertex.agent.llm_response JSON
gemini-3.5-flash off 8 web_search_queries, 3 chunks no

Spans emitted in every case: invocation, invoke_agent root, call_llm, generate_content <model>. The response contained only text parts; the search is not returned as a tool_call/tool_response part.

Repro script
import asyncio, os, sys
os.environ["ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS"] = sys.argv[2]  # "true" / "false"

from google.genai import types
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter

exporter = InMemorySpanExporter()
provider = TracerProvider()
provider.add_span_processor(SimpleSpanProcessor(exporter))
trace.set_tracer_provider(provider)

from google.adk.agents.llm_agent import LlmAgent
from google.adk.runners import InMemoryRunner
from google.adk.tools.google_search_tool import google_search

agent = LlmAgent(name="root", model=sys.argv[1],
                 instruction="Always use Google Search before answering.",
                 tools=[google_search])

async def main():
  runner = InMemoryRunner(agent=agent, app_name="live")
  s = await runner.session_service.create_session(app_name="live", user_id="u")
  events = [e async for e in runner.run_async(
      user_id="u", session_id=s.id,
      new_message=types.Content(role="user", parts=[types.Part(
          text="What is the latest google-adk release on PyPI?")]))]
  print("grounded:", any(e.grounding_metadata for e in events))
  for span in exporter.get_finished_spans():
    for k, v in (span.attributes or {}).items():
      if "grounding" in k or "grounding" in str(v):
        print(span.name, k)

asyncio.run(main())

Run as python repro.py gemini-2.5-flash false. It prints grounded: True and no span lines.

Describe the Solution You'd Like

When a model response carries grounding_metadata, record it as attributes on the existing model-call span, only when it is present, so ungrounded spans are unchanged.

  • Always recorded (no user content): whether the call was grounded, the number of search queries, and the number of grounding chunks.
  • Only when message content capture is enabled (queries are derived from user input): the search query text and the source URIs.

Impact on your work

Without this, you can't use traces to debug or monitor search-grounded agents in deployments that turn content capture off, including the Agent Engine --otel_to_cloud default. For example, you can't tell whether a wrong answer happened because no search ran, the queries were poor, or the sources were poor.

Willingness to contribute

Yes. I can send a PR with unit tests once the naming and gating are agreed.


🟡 Recommended Information

Describe Alternatives You've Considered

  • Emit a synthetic execute_tool google_search span. Rejected: the search runs inside the model call, so ADK has no real start or end time. A made-up span would mislead latency analysis.
  • Rely on gcp.vertex.agent.llm_response. Only available with content capture on, can't be queried as an attribute, and is off by default for adk deploy agent_engine --otel_to_cloud.
  • Enable opentelemetry-instrumentation-google-genai. It has an open TODO to report grounding_metadata, so it doesn't record this today either.
  • Set include_server_side_tool_invocations. Larger behaviour change; belongs with Feature Request: Native AFC & Multi-Tool Combination Support (Built-in Grounding + Custom Python Callables) on Vertex AI #5772, not here.

Proposed API / Implementation

Two naming options. I'd like maintainer input on which fits:

A. Reuse incubating OTel GenAI attributes where they exist

  • gen_ai.retrieval.query.text: the spec marks it as possibly sensitive, so content-gated.
  • gen_ai.retrieval.documents: JSON documents with id/score, content-gated.
  • Mismatches:
    • query.text is a single string, but Gemini issues several queries per call (8–9 in the gemini-3.5-flash runs above).
    • Grounding chunks have a uri/title but no ID and no per-document score.
    • These attributes appear intended for a separate retrieval operation span, not a generate_content span.

B. ADK-specific experimental attributes, following _set_context_cache_attributes

  • adk.experimental.grounding.{grounded, query_count, chunk_count}, always recorded.
  • adk.experimental.grounding.{web_search_queries, source_uris}, content-gated.
  • Behind a new @experimental_telemetry(gate="grounding").

My suggestion is B for the counts (there's no semconv equivalent), and a maintainer decision on whether the content-gated fields should use A or B.

# telemetry/tracing.py, called from trace_call_llm and trace_inference_result
@experimental_telemetry(gate="grounding")
def _set_grounding_attributes(span, grounding_metadata, capture_content: bool):
  if grounding_metadata is None:
    return
  ...

Open questions:

  1. Should this use A, B, or the mix above?
  2. Should the content-gated fields follow should_add_content_to_legacy_spans or the OTel GenAI content-capture setting?

Additional Context

Environment: google-adk 2.10.0 (main @ def458b6), google-genai 2.24.0, opentelemetry-sdk 1.42.1, Python 3.11.15, Windows 11, Vertex AI backend.

Scope: the other built-in grounding tools (VertexAiSearchTool, EnterpriseWebSearchTool, GoogleMapsGroundingTool, url_context) reach the model through the same path (process_llm_request appends to config.tools). A mock-model repro shows the same result for VertexAiSearchTool. I have only verified google_search live.

Prior art checked:

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

needs review[Status] The PR/issue is awaiting review from the maintainertracing[Component] This issue is related to OpenTelemetry tracing

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions