Skip to content

fix(schema): decode Python string literals in static schema - #3182

Open
breken-ai wants to merge 1 commit into
replicate:mainfrom
breken-ai:fix/schema-python-string-literals
Open

breken-ai wants to merge 1 commit into
replicate:mainfrom
breken-ai:fix/schema-python-string-literals

Conversation

@breken-ai

Copy link
Copy Markdown

Summary

  • Make parseStringLiteral (static schema generator) return the value Python gives a string literal: strip r/u prefixes and quotes, decode backslash escapes in non-raw strings, and join implicitly concatenated literals. Bytes and f-strings are still rejected.
  • Add table tests for the literal parser and a parser test that checks description, regex and default of an Input(...).

Why

parseStringLiteral returned the source text between the quotes. The bundled .cog/openapi_schema.json then differs from the predictor's real values:

code: str = Input(
    description="A numeric code. "
    "Digits only.",
    regex="^\\d+$",
)

On main this gives:

  • description: A numeric code. "\n "Digits only. The inner quotes and indentation end up in the schema and API docs. Splitting long descriptions across lines like this is common.
  • pattern: ^\\d+$ (two backslashes) instead of ^\d+$. coglet validates inputs against this schema, so "123" is rejected even though the predictor's own regex accepts it.
  • r"""...""" came back as "".."", and u"..." was not recognised. For default=, that made the default None.

Test plan

  • go test ./pkg/schema/python -run 'TestParseStringLiteral|TestInputStringArgumentsUsePythonStringValues' -count=1 fails on main and passes with this change
  • go test ./pkg/schema/... ./pkg/config/... -count=1 (includes the fuzz seed corpora)
  • go build ./...
  • golangci-lint run ./pkg/schema/... (v2.10.1): 0 issues

I found and fixed this with an AI coding agent, and I checked the change and the test results above.

parseStringLiteral returned the source text between the quotes, so the
generated OpenAPI schema differed from the Python values:

- implicitly concatenated strings ("a " "b") kept the inner quotes and
  newline, garbling multi-line descriptions;
- escape sequences were not decoded, so regex="^\\d+$" became the pattern
  `^\\d+$` and coglet's schema validation rejected valid input;
- r"""...""" and u"..." literals were mis-parsed or dropped.

Decode prefixes, quotes, escapes and implicit concatenation as Python does.
@breken-ai
breken-ai requested a review from a team as a code owner September 25, 2026 10:27

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant