Skip to content

fix(llm): surface invalid structured output - #996

Open
foma-agent wants to merge 1 commit into
plastic-labs:mainfrom
foma-agent:fix/invalid-structured-output
Open

foma-agent wants to merge 1 commit into
plastic-labs:mainfrom
foma-agent:fix/invalid-structured-output

Conversation

@foma-agent

@foma-agent foma-agent commented Aug 8, 2026 •

Copy link
Copy Markdown

Summary

  • raise StructuredOutputError when PromptRepresentation JSON is malformed or has no recognized schema fields instead of silently returning an empty representation
  • keep schema-valid empty objects and repairable/truncated JSON working
  • validate provider raw text consistently across OpenAI, Anthropic, and Gemini so SDK-defaulted models cannot hide wrong keys
  • keep error telemetry payload-safe: failure class, model, byte length, and SHA-256 digest, without raw model output

Verification

  • uv run pytest -q — 1818 passed, 25 skipped
  • uv run basedpyright — 0 errors, 0 warnings
  • changed-file pre-commit suite — clean
  • exact-head Claude Code review — 0 findings

Closes #993

Summary by CodeRabbit

  • Bug Fixes

    • Improved structured-output validation across Anthropic, Gemini, and OpenAI responses.
    • Added safer handling for malformed, incomplete, or schema-incompatible JSON.
    • Errors now provide diagnostic details without exposing invalid response contents.
    • Empty structured representations are accepted when valid.
    • Gemini streaming now reports blocked responses that contain no text.
  • Tests

    • Expanded coverage for parsing, validation, repair, truncation, and blocked responses.

@coderabbitai

coderabbitai Bot commented Aug 8, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

Structured-output parsing now validates schema relevance across Anthropic, Gemini, and OpenAI backends. Repairable truncated JSON remains supported. Unrecoverable or schema-irrelevant responses raise sanitized StructuredOutputError diagnostics instead of returning empty fallback representations.

Changes

Structured output validation

Layer / File(s) Summary
Validation and repair rules
src/llm/structured_output.py, tests/llm/test_structured_output.py
Shared validation rejects schema-irrelevant payloads, preserves valid empty representations and truncated JSON repair, and reports model and payload-hash diagnostics.
Backend parsing integration
src/llm/backends/anthropic.py, src/llm/backends/gemini.py, src/llm/backends/openai.py, tests/llm/test_backends/*
The backends route structured responses through shared validation and repair handling. OpenAI preserves model-specific raw-content parsing.
Regression coverage
tests/utils/test_length_finish_reason.py
Tests now require StructuredOutputError for unrecoverable responses across the supported backends.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant LLMBackend
  participant StructuredOutputValidator
  participant JSONRepair
  participant Caller
  LLMBackend->>StructuredOutputValidator: Validate parsed or raw response
  StructuredOutputValidator-->>LLMBackend: Return valid structured output
  StructuredOutputValidator->>JSONRepair: Handle decoding or validation failure
  JSONRepair-->>LLMBackend: Return repaired output
  JSONRepair-->>Caller: Raise StructuredOutputError when repair fails
Loading

Possibly related PRs

Suggested reviewers: vvoruganti

Poem

A rabbit checks each JSON key,
No silent loss escapes its eye.
Empty fields may still be true,
Broken shapes raise errors through.
Hashes guard the message trail,
And retries wake when repairs fail. 🐇

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 30.43% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the primary change: invalid structured output now surfaces as an error.
Linked Issues check ✅ Passed The changes satisfy #993 by propagating invalid output errors, preserving valid empty and repairable JSON, enforcing schema relevance, and protecting telemetry.
Out of Scope Changes check ✅ Passed The changes remain focused on structured-output validation, provider consistency, safe error telemetry, and tests required by #993.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
src/llm/structured_output.py (1)

38-38: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Add Google-style docstrings to the changed functions.

  • src/llm/structured_output.py#L38-L38: Document repair behavior, arguments, return value, and raised errors.
  • src/llm/structured_output.py#L114-L117: Document supported payload types and error behavior.
  • src/llm/structured_output.py#L137-L158: Document schema relevance and empty-object detection behavior.

As per coding guidelines, In Python code, use Google-style docstrings.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/llm/structured_output.py` at line 38, The changed functions in
src/llm/structured_output.py at lines 38-38, 114-117, and 137-158 require
Google-style docstrings: document the first function’s JSON repair behavior,
arguments, return value, and raised errors; document the second function’s
supported payload types and error behavior; and document the third function’s
schema relevance and empty-object detection behavior.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@src/llm/structured_output.py`:
- Line 38: The changed functions in src/llm/structured_output.py at lines 38-38,
114-117, and 137-158 require Google-style docstrings: document the first
function’s JSON repair behavior, arguments, return value, and raised errors;
document the second function’s supported payload types and error behavior; and
document the third function’s schema relevance and empty-object detection
behavior.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: e6462f48-668b-4ff1-b47b-9bb4afe38bdb

📥 Commits

Reviewing files that changed from the base of the PR and between d191c10 and 10772f9.

📒 Files selected for processing (9)
  • src/llm/backends/anthropic.py
  • src/llm/backends/gemini.py
  • src/llm/backends/openai.py
  • src/llm/structured_output.py
  • tests/llm/test_backends/test_anthropic.py
  • tests/llm/test_backends/test_gemini.py
  • tests/llm/test_backends/test_openai.py
  • tests/llm/test_structured_output.py
  • tests/utils/test_length_finish_reason.py

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Deriver silently discards a batch when structured output fails validation

1 participant