You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit fdda17c
Browse filesBrowse the repository at this point in the historyBrowse files
feat: add after_tool_callback to screen tool output
Input and model output results that report `MATCH_FOUND` are now blocked even when `invocation_result` is not `SUCCESS` and `block_on_screening_failure` is `False`.
Merge #6969Fixes#6966
PiperOrigin-RevId: 994696801
Copy file name to clipboardExpand all lines: docs/guides/integrations/model_armor/index.md
+40-21Lines changed: 40 additions & 21 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -1,18 +1,18 @@
1
1
# Model Armor
2
2
3
-
`ModelArmorPlugin` screens user inputand model output against [Google Cloud Model Armor](https://cloud.google.com/security-command-center/docs/model-armor-overview) templates. When a filter matches, or when screening cannot complete, the offending content is replaced with a safe message before it reaches the model or the user.
3
+
`ModelArmorPlugin` screens user input, model output, and tool output against [Google Cloud Model Armor](https://cloud.google.com/security-command-center/docs/model-armor-overview) templates. When a filter matches, or when screening cannot complete, the offending content is replaced with a safe message before it reaches the model or the user.
4
4
5
5
## Introduction
6
6
7
7
Model Armor is a Google Cloud service that inspects text for prompt injection and jailbreak attempts, harmful content, and sensitive data. You define what to look for in a *template* — a named, server-side policy — and the service returns a verdict for each piece of text you send it.
8
8
9
-
The integration is two public types: `ModelArmorPlugin`, a `BasePlugin` subclass driven by `PluginManager`, and `ModelArmorConfig`, which says which templates to screen against and what to do about a match. The plugin reads text off the `LlmRequest` and `LlmResponse`, calls Model Armor, and returns a replacement`LlmResponse` when content should be blocked.
9
+
The integration is two public types: `ModelArmorPlugin`, a `BasePlugin` subclass driven by `PluginManager`, and `ModelArmorConfig`, which says which templates to screen against and what to do about a match. The plugin reads text off the `LlmRequest`, `LlmResponse`, and tool results, calls Model Armor, and returns a replacement when content should be blocked.
10
10
11
11
Key features:
12
12
13
-
-**Inputand output screening**, each governed by its own template, and each optional.
13
+
-**Input, output, and tool output screening**, each governed by its own template, and each optional.
14
14
-**Block screening failures by default**: by default screening failures are blocked rather than delivered.
15
-
-**Blocked responses are marked** with `custom_metadata['model_armor_blocked']` so your application can detect them.
15
+
-**Blocked input and model output responses are marked** with `custom_metadata['model_armor_blocked']` so your application can detect them (blocked tool outputs return a plain `{'error': ...}` dict to the model and do not carry this marker).
16
16
17
17
## Get started
18
18
@@ -76,22 +76,34 @@ Credentials come from Application Default Credentials.
76
76
3. The text is sent to Model Armor's `SanitizeModelResponse` method.
77
77
4. It acts on the result (below).
78
78
79
+
### `after_tool_callback` - screening tool output
80
+
81
+
`after_tool_callback` runs after each tool call:
82
+
83
+
1. If `tool_output_template_name` is unset, it returns immediately and nothing is screened.
84
+
2. It extracts and serializes text from the tool result (skipping raw bytes and non-text media parts), splitting oversized text into overlapping 65,536-character chunks.
85
+
3. Each chunk is sent to Model Armor's `SanitizeUserPrompt` method.
86
+
4. It acts on each result (below).
87
+
79
88
### Acting on a result
80
89
81
-
|`invocation_result`| Meaning | Plugin behavior |
82
-
| :--- | :--- | :--- |
83
-
|`SUCCESS`| Every filter ran. | Check `filter_match_state`. |
84
-
| Anything else | Some or all filters were skipped, failed, or the field was unset. | Screening failure. |
90
+
|`filter_match_state`|`invocation_result`| Meaning | Plugin behavior |
91
+
| :--- | :--- | :--- | :--- |
92
+
|`MATCH_FOUND`| Any | At least one filter tripped. | Block content. |
93
+
| Anything else |`SUCCESS`| Every filter ran with no match. | Pass through untouched. |
94
+
| Anything else | Anything else | Some or all filters were skipped, failed, or the field was unset. | Screening failure. |
85
95
86
-
When screening completes successfully, a `filter_match_state`of `MATCH_FOUND` means at least one filter tripped, and the content is blocked. Anything else passes through untouched.
96
+
`filter_match_state`is checked first, so a partial result that still reports `MATCH_FOUND` is always blocked.
87
97
88
98
A screening failure is routed through `block_on_screening_failure` and blocked by default.
89
99
90
100
### The blocked response
91
101
92
-
Blocking returns an `LlmResponse` carrying the message for the direction that
93
-
was screened: `input_blocked_message` for user input, `output_blocked_message`
94
-
for model output.
102
+
Blocking user input or model output returns an `LlmResponse` carrying the
103
+
message for the direction that was screened: `input_blocked_message` for user
104
+
input, `output_blocked_message` for model output. Blocking tool output returns
105
+
`{'error': tool_output_blocked_message}` to the model as the tool result instead
106
+
of an `LlmResponse`.
95
107
96
108
### Template paths and regional endpoints
97
109
@@ -124,27 +136,30 @@ Options introduced by `ModelArmorPlugin` (those inherited from `BasePlugin` are
124
136
| :--- | :--- | :--- | :--- |
125
137
|`prompt_template_name`|`str \| None`|`None`| Template used to screen user input. Unset means input is not screened. |
126
138
|`response_template_name`|`str \| None`|`None`| Template used to screen model output. Unset means output is not screened. |
139
+
|`tool_output_template_name`|`str \| None`|`None`| Template used to screen tool output. Unset means tool output is not screened. |
127
140
|`input_blocked_message`|`str`|`"I'm sorry, but I can't help with that request."`| Replacement text shown when user input is blocked. |
128
141
|`output_blocked_message`|`str`|`"I'm sorry, but I can't help with that request."`| Replacement text shown when model output is blocked. |
142
+
|`tool_output_blocked_message`|`str`|`"Tool output was blocked by Model Armor."`| Replacement error message returned to the model when tool output is blocked. |
129
143
|`block_on_screening_failure`|`bool`|`True`| Whether to block content that could not be screened. |
130
144
131
-
At least one of the two template names must be set.
145
+
At least one of `prompt_template_name`, `response_template_name`, or `tool_output_template_name` must be set.
-`prompt_template_name`: Screens user input prompts before forwarding to the model.
141
155
-`response_template_name`: Screens model responses before delivering to the user.
156
+
-`tool_output_template_name`: Screens tool outputs before returning to the model.
142
157
143
-
If both are set they must reside in the same GCP location — see [Template paths and regional endpoints](#template-paths-and-regional-endpoints).
158
+
If multiple templates are set they must reside in the same GCP location — see [Template paths and regional endpoints](#template-paths-and-regional-endpoints).
#### `input_blocked_message`, `output_blocked_message`, and `tool_output_blocked_message`
146
161
147
-
Defines the replacement text returned to the user when a prompt or response is blocked. Screening failures reuse the message for the direction that failed.
162
+
Defines the replacement text returned to the user when a prompt or response is blocked, or returned to the model in `{'error': tool_output_blocked_message}` when tool output is blocked. Screening failures reuse the message for the direction that failed.
148
163
149
164
#### `block_on_screening_failure`
150
165
@@ -174,7 +189,7 @@ config = ModelArmorConfig(
174
189
175
190
### Detecting blocks in your application
176
191
177
-
Blocked responses carry a marker, so a UI can render them differently from a real answer:
192
+
Blocked input and model output responses carry a `custom_metadata['model_armor_blocked']`marker, so a UI can render them differently from a real answer (blocked tool outputs return a plain `{'error': ...}` dict to the model and do not carry this marker):
178
193
179
194
```python
180
195
asyncfor event in runner.run_async(...):
@@ -184,7 +199,11 @@ async for event in runner.run_async(...):
184
199
185
200
## Limitations
186
201
187
-
-**Tool output is not screened.** Only the most recent `user` content with text parts is sent for screening. Tool results are added to the request as `user` content whose only part is a `function_response` and doesn't reach Model Armor.
202
+
-**Live streaming and long-running tools are not screened.** Live streaming tools send yielded chunks directly to the live request queue and long-running tools deliver their final `function_response` on a later turn, so their outputs skip `after_tool_callback`.
203
+
204
+
-**Media in tool results is not screened.** Raw `bytes` values and non-text `Part` objects (such as `inline_data` or `file_data` media) in tool results are skipped during text extraction and are not sent to Model Armor.
205
+
206
+
-**Plugin registration order matters.**`PluginManager` stops at the first plugin that returns a non-`None` value from a callback. Register `ModelArmorPlugin` ahead of plugins that return a value from `after_tool_callback` (such as `MultimodalToolResultsPlugin`) so tool output screening is not bypassed.
188
207
189
208
-**Enforcement mode is limited.** The Model Armor plugin is currently limited to logging detection results and blocking content. Future extensions could include replacing or redacting text.
0 commit comments