Skip to content

Harvest the daily deleted-domain feed into a persistent shortlist - #59

Merged
ThinkingSpade merged 6 commits into
mainfrom
feat/deleted-domain-harvest
Aug 21, 2026
Merged

ThinkingSpade merged 6 commits into
mainfrom
feat/deleted-domain-harvest

Conversation

@ThinkingSpade

Copy link
Copy Markdown
Owner

Turns the WhoisFreaks subscription into a stored, graded shortlist of dropped .com domains in the project's industries — and the industries around it.

Why the first run looked broken

The Expired Domains tab returned five vending companies. The vocabulary was five words taken from the project's own keywords (breakroom, office, coffee, dallas, vending). A client's keywords can only describe the client's own vertical, so a harvest built from them returns more of that vertical — by construction.

Adjacent terms now reach the industries around the business: the verticals its customers are in, the venues it operates in, and topics it could credibly publish about. A vending operator serves schools, so an education domain is a legitimate target even though education is not its trade.

Measured on one real day of the feed:

Matches
Seed terms only 147
Seed + adjacent industries 1,234

Top matched terms: water 146, logistics 127, school 110, hotel 99, wellness 99, clinic 89, fitness 86, education 52, university 31.

Two endpoint corrections, found live

The documented endpoint does not exist. whoisfreaks.com/api/v3/dropped-domains returns their marketing site's Next.js 404 page — HTML, identically for every date, which is why it parses as a failure only by luck. The real feed is files.whoisfreaks.com/v3.1/download/domainer/dropped, found by rendering their JavaScript docs. A test pins it.

And the body is gzipped newline-delimited names, not JSON. Decoded with DecompressionStream, CRLF-tolerant, TLD filtered client-side — there is no server-side TLD filter, so the download is the whole day (~2 MB, ~240k rows, ~84k of them .com).

whois=true returns 413 "Please upgrade your plans", so the API enforces the subscribed tier. Also pinned.

Cost model — deliberately ungated

Harvesting and DR grading are free: a flat subscription and a keyless endpoint. There is no spend gate on either, and adding one would only train the user to click past the gates that do matter. Availability is the exception — 5 APIVerve credits per domain, behind an explicit action that states its cost and is capped at 25 per click.

Design notes

  • Cron is self-limiting, not scheduled. It rides the existing 15-minute tick; datesToHarvest returns nothing once a day is stored, so real work happens about once a day per project. One day per tick — a 2 MB file should not be held several at a time.
  • Rows outlive the subscription. The feed is a tap that gets turned off after a harvest window; the shortlist is the asset. Nothing downstream depends on the feed being live.
  • The matcher matches the STEM, never the TLD — otherwise every .coffee domain hits "coffee" and one TLD floods the harvest. Longest term wins, so "breakroom" explains a row rather than "room". Terms under four characters are ignored because "co" matches a large share of the internet.
  • domain_rating null means UNGRADED, never "no authority" — a real 0 stores as 0. Only a thrown lookup leaves a row for a later tick to retry.
  • Vocabulary cached 30 days per project, and an empty answer is never cached: one transient model failure would otherwise lock the harvest into the narrow vocabulary for a month.

Refactor included

resolveDomainRating moved out of serverFunctions/ahrefs.ts into a lib taking its cache as a parameter. That file statically imports cloudflare:workers, so the rating logic could be reached from neither a cron nor a test. Behaviour is unchanged — same cache prefix, TTL, and 0-vs-null rule.

Verification

  • pnpm ci:check clean; pnpm test green — 315 files, 3181 tests.
  • Migration 0041 already applied to prod D1 ahead of this code; additive and nullable.
  • Verified live end to end against the real subscription: 242,478 domains for 2026-08-19 → 83,676 .com → 1,234 matches, each attributed to the term that hit.

Not verified

The panel is not browser-verified — the harvest table is empty until this deploys and a day is pulled.

🤖 Generated with Claude Code

Migration 0041 creates harvested_domains, ALREADY APPLIED to prod D1 ahead of
this code. Rows persist deliberately: the WhoisFreaks subscription is a tap
that gets turned off after a month, and the shortlist it produces is the thing
of value -- nothing downstream may depend on the feed still being live.

The client pulls the DELETED feed, not the expiring one. Expiring domains are
still in redemption or pending-delete and nobody can register them; deleted
ones are registerable at registration price today, which is the whole point.
TLDs are filtered server-side so a day's pull is .com rather than 1,529 TLDs of
mostly junk.

403 maps to AUTH_FAILED rather than a credits code, unlike the APIVerve
mapping: this is a flat subscription with no per-call balance to exhaust.

domain_rating is nullable and null means UNGRADED or no answer, never 'no
authority' -- a real DR of 0 stores as 0. That distinction silently broke a
ranking verdict in this repo once already.
Runs off the existing 15-minute cron and is self-limiting rather than
schedule-driven: datesToHarvest returns nothing once a day is stored, so real
work happens about once a day per project and every other tick is two cheap
reads. One date per tick, because a day's file is ~2 MB and ~240,000 rows and
holding several at once would be pointless.

The vocabulary matcher is the only thing standing between 84,000 .com names a
day and what gets stored, so it is pure and tested. It matches the STEM, never
the TLD -- otherwise every .coffee domain hits 'coffee' and one TLD floods the
harvest -- prefers the longest matching term so 'breakroom' explains a row
rather than 'room', and ignores terms under four characters because 'co'
matches a large share of the internet.

Ahrefs DR moved out of serverFunctions/ahrefs.ts into a lib that takes its
cache as a parameter. That file statically imports cloudflare:workers, so the
rating logic could be reached neither from a cron nor from a test. Behaviour is
unchanged and the 0-vs-null rule travels with it: a real 0 stores as 0, and
only a thrown lookup leaves a row ungraded for a later tick to retry.

Verified against the live feed: one real day yields 147 matches for this
project's vocabulary, each attributed to the term that hit.
The first real run returned five vending companies because the vocabulary was
five words taken from the project's own keywords. A client's keywords can only
ever describe the client's own vertical, so a harvest built from them returns
more of the same vertical -- by construction, not by accident.

Adjacent terms now reach the industries AROUND the business: the verticals its
customers are in, the venues it operates in, and topics it could credibly
publish about. A vending operator serves schools, so an education domain is a
legitimate target even though education is not its trade. Measured on one real
day of the feed, that takes matches from 147 to 1,234 -- 110 school, 99 hotel,
89 clinic, 86 fitness, 52 education, 31 university.

The vocabulary is cached per project for a month: deriving it costs a model
call and a business's surrounding industries do not move week to week. An empty
answer is never cached, or one transient model failure would lock the harvest
into the narrow vocabulary for thirty days.

The panel reads the stored shortlist on mount with no gate. Harvesting and DR
grading are genuinely free -- a flat subscription and a keyless endpoint -- and
gating free actions only trains people to click past the gates that do matter.
Availability is the exception and states its cost.
@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Aug 21, 2026 •

Copy link
Copy Markdown

Deploying with  Cloudflare Workers  Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

Status Name Latest Commit Updated (UTC)
✅ Deployment successful!
View logs
flyrocketseo 385b2c7 Aug 21 2026, 09:03 AM

ThinkingSpade and others added 3 commits August 20, 2026 23:51
The worst was fatal and I would have shipped it. D1 allows 100 BOUND
PARAMETERS per statement and drizzle binds every column, so the 50-row insert
chunk bound 250 and failed outright: any day with 21 or more matches saved
nothing at all. Chunk is 15 now.

Completion was inferred from matched rows, so a legitimate zero-match day left
no trace and was re-downloaded on every 15-minute tick -- 84 pulls of a 2 MB
file a day, while older backfill dates never ran. A harvest_runs table records
every processed date including empty ones, written only after all inserts
succeed so a partial write retries instead of being skipped. Writing that test
exposed another one: an insert failure escaped and aborted the whole backfill
rather than failing a single date.

Buffering the feed cost 70 ms of CPU for one day -- measured, against a
free-plan allowance an order of magnitude smaller -- and bounded memory by
nothing, since a gzip's compressed size implies no expanded size. It streams
now, matching as it reads through one compiled alternation instead of thirty
includes per domain, and cancelling the moment the cap is reached: 6 ms for the
same file, an 11x improvement.

Vocabulary was resolved BEFORE checking whether any date needed pulling, so a
fully-harvested project paid for an OpenRouter call on every tick -- up to 96 a
day with no user action. The scheduler also asked for dates before their 03:00
UTC publication, guaranteeing failures on every tick between midnight and 3am,
and the WhoisFreaks key gated keyless Ahrefs grading so ending the subscription
would have frozen the final days' rows ungraded forever.

Exclusions were only lowercased, so a project domain stored as a URL failed to
exclude its own registrable form and could consume a paid availability check.
The billed availability endpoint did not deduplicate its input.

Also adds the drizzle-pg migration for both tables, which was missing entirely
-- the Postgres deploy path would have failed on a nonexistent relation.
Fetch each feed date ONCE per tick and fan it out to every claimed project
instead of downloading the same 2 MB file per project. A project that fills
its cap goes inactive without cancelling the shared stream.

Claim (project, dropped_on) atomically before reading the feed. The unique
index is the mutex; an upsert may replace only an already-expired lease, and
replacing the row id fences the old owner so its late complete or release
cannot touch the new claim. A completed run has a null lease and can never be
reclaimed. Verified against real SQLite: an active lease returns no row, an
expired one returns the new token, a completed one returns no row.

Rank boundary hits ahead of mid-word collisions. Weak hits no longer end the
stream, because a stronger hit may still arrive. On a real 77,217-row .com day
this lifts genuine matches from ~195/300 to 300/300 for 2.6 ms, and the
full-scan worst case is unchanged (11.6 ms ranked vs 12.9 ms before).

Content-address the vocabulary cache key so changed keywords stop serving a
month-old answer, and surface availability age so an answered row can be
re-checked after 24 hours -- still only on an explicit click, still capped and
deduplicated, with the credit cost stated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Six defects an adversarial review found, each reproduced before fixing.

A failed KV write no longer discards an answer that was already paid for --
it used to throw the derived terms away, so a KV write outage billed the model
again on every 15-minute tick.

Vocabulary now resolves only AFTER a project wins its atomic claim, so two
overlapping ticks can no longer both pay for the same answer before one of them
discovers it has nothing to do. `terms` became a thunk to make that orderable.
The manual harvest likewise checks that a date needs pulling before resolving
anything, instead of billing a model call and then reporting it was already up
to date.

The vocabulary cache key hashes only what actually reaches the prompt, so two
profiles that derive an identical seed stop paying twice for a byte-identical
request.

Inserts now verify the claim is still owned. A stalled tick that resumed after
its lease expired could still write rows, and onConflictDoNothing then made its
stale matchedTerm win over the live owner's.

Five tests named behaviour they could not detect. Mutating touchesBoundary to
always return true -- which reduces the matcher to the previous one-shot regex
-- left them green; all five now fail. The overlapping-tick test implemented
the mutex inside its own fake, so it would have passed even if the repository
let both callers win; it now asserts on how the claim result is USED.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@ThinkingSpade
ThinkingSpade merged commit 4909660 into main Aug 21, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant