Skip to content

gap_check resolves early when a scaffolded sub-question is answered correctly #50

Description

@aaronashby

Overview

When a wrong gap_check answer causes the tutor to scaffold/mini-lesson the student through a prerequisite topic (one or more intermediate sub-questions before its real check-in question), judgeTurn's gap_check branch and advance's gap_check case have no way to distinguish "the student answered one of the tutor's intermediate scaffolding sub-questions correctly" from "the student answered the tutor's real gap-check probe correctly." Any correct GAP_ATTEMPT immediately resolves the current gap (stateMachine.ts, the gap_check case) and, if it's the last remaining gap, transitions straight to solve — even mid-scaffold, while the tutor's own reply is still asking about the sub-question.

This is the same class of bug already fixed for solve in a prior ticket ("Solve judge can't tell a coincidental value match from the real final answer"), but on the gap_check side, and currently unfixed.

Reproduction (manual, observed 2026-08-27)

  1. Trigger a gap on a prerequisite topic (mastery below GAP_THRESHOLD).
  2. Answer the tutor's gap-check question incorrectly — the tutor scaffolds/teaches through the topic with one or more sub-questions.
  3. Answer one of those intermediate sub-questions correctly.
  4. Observed: the session immediately transitions out of gap_check into solve, while the tutor's own reply is still mid-scaffold on the sub-question (not a resolved-sounding wrap-up). A second correct answer to the same sub-question ("acknowledging" it) then properly surfaces the actual problem.

Root cause hypothesis

  • judge.ts's gap_check prompt judges only whether the student answered "the question the tutor just asked" correctly — with no concept of whether that question is the tutor's real check-in probe or an intermediate scaffold step introduced while teaching.
  • advance's gap_check case treats any correct GAP_ATTEMPT as resolving the current gap, unconditionally.
  • Unlike solve, which now has an explicit isFinalAttempt signal (a UI toggle, not inferred from text) precisely because a live-model eval showed the model can't reliably distinguish "final answer" from "sub-step that happens to match" from conversation text alone, gap_check has no equivalent signal.

Notes for whoever picks this up

  • Worth reading the prior solve-side fix's history before choosing an approach here: a judge-prompt-only fix (asking the model to weigh the tutor's last message) was tried first and shown unreliable via a live-model eval (missed the disqualifier 1-2 times out of 5); the fix that actually worked was an explicit, non-inferred signal instead of a prompt tweak.
  • gap_check doesn't have as clean a "single final answer" concept as solve — a gap's mini-lesson can have an open-ended number of scaffolding sub-steps before the tutor is done with it, so the same toggle mechanic may not translate directly. Worth deciding whether the fix should be judge-side (e.g., only resolve a gap on a correct answer and an explicit "I'm done teaching this" signal from the tutor's own structured output) or state-machine-side.
  • Consider whether a live-model eval is warranted here too before committing to a prompt-only fix, given the prior experience on the solve side.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions