Skip to content

Fix the three silent failures behind the empty harvest - #60

Open
ThinkingSpade wants to merge 4 commits into
mainfrom
feat/harvest-net-and-grading
Open

ThinkingSpade wants to merge 4 commits into
mainfrom
feat/harvest-net-and-grading

Conversation

@ThinkingSpade

Copy link
Copy Markdown
Owner

Found by reading prod (KV key counts, D1, wrangler tail) rather than by reading code. Cron was ticking every 15 minutes and had stored 659 rows, while three separate things were broken — two of which logged nothing at all.

Adjacent terms had never worked

deriveAdjacentTerms borrowed getChatAgentModel(), which enables the reasoning channel. Measured live:

max_tokens finish reasoning text chars usable terms
400 (was) length 399 0 0
1500 (now) stop 427 956 97

Prod KV held zero harvest-vocab: keys and every stored row matched one of 5 seed words. It billed OpenRouter on every call to return nothing.

Reasoning is deliberately kept on — without it MiniMax leaks its <think> trace inline and parseAdjacentTerms would accept those words as vocabulary, which decides what gets harvested and what costs availability credits. The slug is pinned rather than read from OPENROUTER_MODEL, which was the only environment difference behind a prod-only Missing Authentication header.

The backfill was permanently stuck

Probed live: 2026-08-17+ → 200, 2026-08-16 and older → 401 "You cannot download file". The subscription only grants dates from its start, so backfill is impossible. The newest unharvested date is picked first, so every tick chose 08-16, failed, recorded nothing and retried — ~95 wasted requests/day, no row written since 04:45 UTC. Now a permanent skip, distinct from a bad key and from a transient failure.

DR grading had never succeeded

0 of 659 rows graded, zero ahrefs-dr: keys in KV, hidden behind a bare catch {}. Now logged, capped at 3 attempts, newest-first, plus a free explicit "Grade now" that loops bounded batches.

Sizing derived, not guessed

Cloudflare Free allows 50 queries per invocation, spent by D1 and KV and fetches. The harvest moved to its own cron trigger, a tick harvests exactly one project, and the daily cap is min(insert budget, what grading keeps up with) = 248, chosen so DR stays current instead of growing an ungraded tail. Widened to com/net/org/co and 50 terms.

Migrations 0044–0046 were applied to prod before this push, since a branch push deploys here.

🤖 Generated with Claude Code

…empty

Found by reading prod, not code: cron ticked every 15 minutes and stored 659
rows while three separate things were broken, two of them logging nothing.

Adjacent industry terms had NEVER worked. deriveAdjacentTerms borrowed
getChatAgentModel(), which enables the reasoning channel; measured against the
live API that spent 399 of 400 output tokens on reasoning and returned ONE text
token, so the parser saw an empty string and every harvest ran on 5 seed words
while OpenRouter billed for each call. Prod additionally threw "Missing
Authentication header" where local succeeded, the only difference being the
OPENROUTER_MODEL secret. The task now pins its own slug, keeps the reasoning
channel (without it MiniMax leaks its <think> trace into the answer, and the
parser would happily accept those words as vocabulary), and gets a budget that
reasoning cannot starve -- measured 1,500 tokens yields ~97 usable terms.
Empty or trace-like replies are now failures, never cached.

The backfill was permanently stuck. WhoisFreaks 401s every date before the
subscription began: 08-17 and newer return 200, 08-16 and older do not. Since
the newest unharvested date is picked first, every tick chose 08-16, failed,
recorded nothing and retried -- about 95 wasted requests a day, with no row
written since 04:45 UTC. That refusal is now a permanent skip, kept distinct
from a bad key and from a transient failure, and surfaced rather than swallowed.

DR grading had never succeeded once: 0 of 659 rows graded and zero ahrefs-dr
keys in KV, invisible behind a bare catch. It now logs, gives up after three
attempts instead of retrying forever, grades newest first, and has an explicit
free "Grade now" that loops bounded batches so the column fills on demand.

Sizing is now derived rather than guessed. Cloudflare allows 50 queries per
invocation on the Free plan and D1, KV and fetches all spend it, so the harvest
moved to its own cron trigger, a tick harvests exactly one project, and the
daily cap is the smaller of what inserts allow and what grading can keep up
with -- 248, chosen so Domain Rating stays current rather than accumulating an
ungraded tail. Widened to com/net/org/co and 50 vocabulary terms.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Aug 21, 2026 •

Copy link
Copy Markdown

Deploying with  Cloudflare Workers  Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

Status Name Latest Commit Updated (UTC)
✅ Deployment successful!
View logs
flyrocketseo 17a90c7 Aug 21 2026, 03:27 PM

ThinkingSpade and others added 3 commits August 21, 2026 09:09
The harvest was moved onto its own cron ("7,22,37,52 * * * *") so it would not
share the Free-plan 50-query invocation budget with the rank checks. The
trigger was committed and deployed and NEVER FIRED: over two scheduled windows
a prod tail saw only "*/15 * * * *". Because the handler harvested solely on
the new expression, the harvest stopped running altogether -- a worse failure
than the budget contention it was meant to avoid.

Cause: Workers Builds deploys through the versions API, and
`wrangler versions upload`/`deploy` do not apply the `triggers` block --
only `wrangler deploy` does. A newly declared cron is therefore committed,
built, deployed, and silently never registered. `wrangler triggers deploy`
would register it by hand, but a feature that dies unless someone remembers an
out-of-band CLI step is not a feature that works.

So the two jobs are now separated by TICK inside the single trigger that
already exists: the top of each hour runs the rank checks, the other three run
the harvest. Splitting a trigger that is already registered cannot fail the
same way. Rank checks are due-based -- they select configs whose nextRunAt has
passed and advance by the config's own interval -- so an hourly tick delays a
daily check by at most an hour.

The daily cap follows the arithmetic down: 72 harvest ticks a day instead of
96 means grading resolves 552 rows a day, so the cap is 184 per project rather
than 248, still the smaller of what inserts allow and what grading keeps up
with.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three failures in this cron each hid behind a catch-and-continue for days while
the tick itself reported "Ok", and a trigger that had silently stopped firing
looked exactly like work that ran and found nothing to do. A tail cannot tell
those apart: an absent log line is equally consistent with "did nothing" and
"never ran". One line naming the tick and the unit it selected distinguishes
them, so it stays.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The dispatch log added last commit proved the tick fires, but a tick that
completes successfully and does nothing is indistinguishable from one that
threw somewhere uninstrumented -- both closures only logged on the exception
path, and the caller discarded harvestDroppedDomains's return value entirely.
Four ticks after the fenced-claim fix landed, D1 still showed zero attempts
and zero new harvest_runs rows, with no error logged either.

One line per tick now reports what actually happened: matched/harvested/
skipped/failed counts for a harvest attempt, attempted/graded/failed/remaining
for a grading pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant