Skip to content

fix: enable native Anthropic prompt caching - #204

Open
eliebak wants to merge 2 commits into
mainfrom
fix/anthropic-prompt-caching
Open

eliebak wants to merge 2 commits into
mainfrom
fix/anthropic-prompt-caching

Conversation

@eliebak

@eliebak eliebak commented Sep 30, 2026 •

Copy link
Copy Markdown

Direct Anthropic requests currently use Chat Completions, which does not support prompt caching. Route api.anthropic.com through the native Messages API and enable automatic five-minute caching, with a separate system-prompt breakpoint for reuse across compaction.

Add optional api_format, prompt_cache (5m, 1h, or off), and max_output_tokens provider settings. Preserve native content and tool history, handle Anthropic retries, context overflows, and context-window discovery, and count cache writes toward token budgets while excluding cache reads. Full prompt usage still includes both for context management.

Validation:

  • uv run pytest tests/: 338 passed; 5 opt-in live tests skipped.
  • Live Anthropic tests: 5 passed, covering cache lifetimes, growing history, caching disabled, tool execution, compaction, and ACP stdio follow-ups.
  • Markdownlint, Ruff lint, and formatting checks passed.

Note

Medium Risk
Changes the default inference path for Anthropic endpoints and token-budget accounting; mistakes could affect cost limits or break tool/compaction flows, though behavior is heavily tested.

Overview
Anthropic traffic now uses the native Messages API when provider.base_url is api.anthropic.com (or api_format is "anthropic"), instead of Chat Completions, so prompt caching works. OpenAI-style chat history is converted in a new transport layer; responses are normalized back to ChatCompletion, including anthropic_content for faithful tool replay.

The ACP provider contract gains api_format (auto / openai / anthropic), prompt_cache (5m, 1h, off), and max_output_tokens. Requests set ephemeral cache controls (plus a system-prompt breakpoint for reuse after compaction). Tree max_total_tokens billing treats cache reads as non-new work while still counting full prompt usage for compaction thresholds; Anthropic retries, overflow detection, and model context-window discovery (max_input_tokens) are wired through the shared client path.

Documentation describes caching behavior and opt-in live tests; anthropic is added as a dependency with unit and live test coverage.

Reviewed by Cursor Bugbot for commit 5864229. Bugbot is set up for automated code reviews on this repo. Configure here.

@eliebak
eliebak marked this pull request as ready for review September 30, 2026 01:51

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want reviews to match your repository better? Bugbot Learning can learn team-specific rules from PR activity. A team admin can enable Learning in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 5864229. Configure here.

Comment thread src/rlm/compaction.py
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant