feat: context window renders processed and self turns as symmetric context tags - #967
Draft
gorkem2020 wants to merge 4 commits into
Draft
feat: context window renders processed and self turns as symmetric context tags#967gorkem2020 wants to merge 4 commits into
gorkem2020 wants to merge 4 commits into
Conversation
…ant misattribution
With "User:"/"Assistant:" line prefixes only the first line of a message
carries a speaker marker, so a multi-paragraph assistant reply sheds its
marker after the first paragraph and the extractor attributes
assistant-authored plans and preferences to the user. Wrap every message
wholly in <user_message>/<assistant_message> blocks built from the hook's
role-tagged message-loop order, neutralize literal speaker tags typed inside
message content, and snap extractMaxChars truncation to a tag boundary so a
sliced transcript never leads with headless text.
The extraction prompt becomes a {system, user} split so the format teaching
and grounding rules ride the system half while the transcript rides the user
half: captureAssistant=false grounds memories exclusively in user blocks
(assistant lines never enter the transcript); captureAssistant=true keeps
assistant blocks as attributable sources with explicit attribution rules.
completeJson gains an optional per-call system prompt to carry the split.
Port the extraction lane's speaker-tagged conversation structure into the reflection distiller input: session messages render as <user_message>/ <assistant_message> blocks instead of role-colon lines, and the INPUT code fence is removed (any code block inside the conversation terminated the fence early and leaked the rest of the transcript out of the input frame). Clipping now snaps to whole tagged blocks via trimTranscriptToTagBoundary instead of slicing mid-message, and the prompt teaches the tag grammar up front. Stored session-summary rows keep the legacy labeled role-colon shape via an explicit format switch: a stored row must never carry literal speaker tags that a later recall could replay into a prompt as fake transcript structure.
…r extraction
In steady state (history-carrying sessions extracting every turn) each
extraction transcript contains only that call's own turns, so the extractor
never has the conversational context to resolve references ("yes exactly,
that one"). Add a rolling pair window: the last N user turns, with their
assistant replies interleaved in true order, stay in the transcript across
extractions. Retained turns are context for the extractor; the watermark
keeps them from re-becoming extraction sources.
N = autoCaptureContextTurns: 0 (the default) disables retention entirely and
preserves stock behavior; 1-10 sets the window size, decoupled from the
extractMinMessages warm-up gate. The window is bounded at set time by
trimTurnsToUserCap (never trimming this call's own unextracted turns out of
their transcript), and dedupePairWindow repairs double-included pairs when a
below-threshold deferral preserves the same exchange on two paths.
…symmetric context tags With autoCaptureContextTurns > 0 the retained window now carries real conversational context with mechanically exact source marking: - Already-processed turns (the retained prior window) render as <context_user_message>/<context_assistant_message> blocks, marked at retrieval from the rolling buffer: processed-ness needs no new state, retained prior-window turns are by definition the already-extracted set. - Under captureAssistant=false, assistant replies now enter the transcript window as context blocks: they join the rolling window for reference resolution (pronouns, follow-ups, corrections) while staying out of the eligible set: they cannot fire the gate, move the watermark, or ground a memory. Self replies always render as <context_assistant_message> (self is never a source). - The NEW delta alone keeps the extractable <user_message> (and, under captureAssistant=true, <assistant_message>) tags, so 'Extract ONLY from <user_message> blocks' is mechanically exact instead of prose about earlier turns. The prompt teaches exactly the tags that can appear in each mode; context tags only exist when the window is enabled, and the plain modes are byte-identical. The spoof neutralizer and tag-boundary trimmer cover the context tags too.
This was referenced Jul 22, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
With the rolling context window (#966) the transcript now carries retained turns, but they wear the same tags as the new delta. The extractor re-proposes candidates from already-processed turns every call (burning dedup/judge work each time), and under
captureAssistant=falseassistant replies are absent entirely, so the window still cannot resolve "that one" references against what the assistant said.Change
The window renders with mechanically exact source marking:
<context_user_message>/<context_assistant_message>blocks, marked at retrieval from the rolling buffer. Processed-ness needs no new state: retained prior-window turns are by definition the already-extracted set.captureAssistant=false, assistant replies now enter the transcript window as context blocks: they join the rolling window for reference resolution (pronouns, follow-ups, corrections) while staying out of the eligible set, so they cannot fire the gate, move the watermark, or ground a memory. Self replies always render as<context_assistant_message>(self is never a source).<user_message>(and, undercaptureAssistant=true,<assistant_message>) tags, so "Extract ONLY from<user_message>blocks" is mechanically exact instead of prose about earlier turns.Tests
Prompt-teaching cases for both window modes in
test/extraction-transcript-speaker-tags.test.mjs; full-pipeline cases intest/pair-window-retention.test.mjsproving self replies ride ascontext_assistant_message, processed user turns wearcontext_user_message, and the new delta keepsuser_message.Stacked on #966; the diff includes its predecessors until they merge.