Repository navigation
Harvest the daily deleted-domain feed into a persistent shortlist - #59
Merged
Merged
Conversation
Migration 0041 creates harvested_domains, ALREADY APPLIED to prod D1 ahead of this code. Rows persist deliberately: the WhoisFreaks subscription is a tap that gets turned off after a month, and the shortlist it produces is the thing of value -- nothing downstream may depend on the feed still being live. The client pulls the DELETED feed, not the expiring one. Expiring domains are still in redemption or pending-delete and nobody can register them; deleted ones are registerable at registration price today, which is the whole point. TLDs are filtered server-side so a day's pull is .com rather than 1,529 TLDs of mostly junk. 403 maps to AUTH_FAILED rather than a credits code, unlike the APIVerve mapping: this is a flat subscription with no per-call balance to exhaust. domain_rating is nullable and null means UNGRADED or no answer, never 'no authority' -- a real DR of 0 stores as 0. That distinction silently broke a ranking verdict in this repo once already.
Runs off the existing 15-minute cron and is self-limiting rather than schedule-driven: datesToHarvest returns nothing once a day is stored, so real work happens about once a day per project and every other tick is two cheap reads. One date per tick, because a day's file is ~2 MB and ~240,000 rows and holding several at once would be pointless. The vocabulary matcher is the only thing standing between 84,000 .com names a day and what gets stored, so it is pure and tested. It matches the STEM, never the TLD -- otherwise every .coffee domain hits 'coffee' and one TLD floods the harvest -- prefers the longest matching term so 'breakroom' explains a row rather than 'room', and ignores terms under four characters because 'co' matches a large share of the internet. Ahrefs DR moved out of serverFunctions/ahrefs.ts into a lib that takes its cache as a parameter. That file statically imports cloudflare:workers, so the rating logic could be reached neither from a cron nor from a test. Behaviour is unchanged and the 0-vs-null rule travels with it: a real 0 stores as 0, and only a thrown lookup leaves a row ungraded for a later tick to retry. Verified against the live feed: one real day yields 147 matches for this project's vocabulary, each attributed to the term that hit.
The first real run returned five vending companies because the vocabulary was five words taken from the project's own keywords. A client's keywords can only ever describe the client's own vertical, so a harvest built from them returns more of the same vertical -- by construction, not by accident. Adjacent terms now reach the industries AROUND the business: the verticals its customers are in, the venues it operates in, and topics it could credibly publish about. A vending operator serves schools, so an education domain is a legitimate target even though education is not its trade. Measured on one real day of the feed, that takes matches from 147 to 1,234 -- 110 school, 99 hotel, 89 clinic, 86 fitness, 52 education, 31 university. The vocabulary is cached per project for a month: deriving it costs a model call and a business's surrounding industries do not move week to week. An empty answer is never cached, or one transient model failure would lock the harvest into the narrow vocabulary for thirty days. The panel reads the stored shortlist on mount with no gate. Harvesting and DR grading are genuinely free -- a flat subscription and a keyless endpoint -- and gating free actions only trains people to click past the gates that do matter. Availability is the exception and states its cost.
Deploying with
|
| Status | Name | Latest Commit | Updated (UTC) |
|---|---|---|---|
| ✅ Deployment successful! View logs |
flyrocketseo | 385b2c7 | Aug 21 2026, 09:03 AM |
The worst was fatal and I would have shipped it. D1 allows 100 BOUND PARAMETERS per statement and drizzle binds every column, so the 50-row insert chunk bound 250 and failed outright: any day with 21 or more matches saved nothing at all. Chunk is 15 now. Completion was inferred from matched rows, so a legitimate zero-match day left no trace and was re-downloaded on every 15-minute tick -- 84 pulls of a 2 MB file a day, while older backfill dates never ran. A harvest_runs table records every processed date including empty ones, written only after all inserts succeed so a partial write retries instead of being skipped. Writing that test exposed another one: an insert failure escaped and aborted the whole backfill rather than failing a single date. Buffering the feed cost 70 ms of CPU for one day -- measured, against a free-plan allowance an order of magnitude smaller -- and bounded memory by nothing, since a gzip's compressed size implies no expanded size. It streams now, matching as it reads through one compiled alternation instead of thirty includes per domain, and cancelling the moment the cap is reached: 6 ms for the same file, an 11x improvement. Vocabulary was resolved BEFORE checking whether any date needed pulling, so a fully-harvested project paid for an OpenRouter call on every tick -- up to 96 a day with no user action. The scheduler also asked for dates before their 03:00 UTC publication, guaranteeing failures on every tick between midnight and 3am, and the WhoisFreaks key gated keyless Ahrefs grading so ending the subscription would have frozen the final days' rows ungraded forever. Exclusions were only lowercased, so a project domain stored as a URL failed to exclude its own registrable form and could consume a paid availability check. The billed availability endpoint did not deduplicate its input. Also adds the drizzle-pg migration for both tables, which was missing entirely -- the Postgres deploy path would have failed on a nonexistent relation.
Fetch each feed date ONCE per tick and fan it out to every claimed project instead of downloading the same 2 MB file per project. A project that fills its cap goes inactive without cancelling the shared stream. Claim (project, dropped_on) atomically before reading the feed. The unique index is the mutex; an upsert may replace only an already-expired lease, and replacing the row id fences the old owner so its late complete or release cannot touch the new claim. A completed run has a null lease and can never be reclaimed. Verified against real SQLite: an active lease returns no row, an expired one returns the new token, a completed one returns no row. Rank boundary hits ahead of mid-word collisions. Weak hits no longer end the stream, because a stronger hit may still arrive. On a real 77,217-row .com day this lifts genuine matches from ~195/300 to 300/300 for 2.6 ms, and the full-scan worst case is unchanged (11.6 ms ranked vs 12.9 ms before). Content-address the vocabulary cache key so changed keywords stop serving a month-old answer, and surface availability age so an answered row can be re-checked after 24 hours -- still only on an explicit click, still capped and deduplicated, with the credit cost stated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Six defects an adversarial review found, each reproduced before fixing. A failed KV write no longer discards an answer that was already paid for -- it used to throw the derived terms away, so a KV write outage billed the model again on every 15-minute tick. Vocabulary now resolves only AFTER a project wins its atomic claim, so two overlapping ticks can no longer both pay for the same answer before one of them discovers it has nothing to do. `terms` became a thunk to make that orderable. The manual harvest likewise checks that a date needs pulling before resolving anything, instead of billing a model call and then reporting it was already up to date. The vocabulary cache key hashes only what actually reaches the prompt, so two profiles that derive an identical seed stop paying twice for a byte-identical request. Inserts now verify the claim is still owned. A stalled tick that resumed after its lease expired could still write rows, and onConflictDoNothing then made its stale matchedTerm win over the live owner's. Five tests named behaviour they could not detect. Mutating touchesBoundary to always return true -- which reduces the matcher to the previous one-shot regex -- left them green; all five now fail. The overlapping-tick test implemented the mutex inside its own fake, so it would have passed even if the repository let both callers win; it now asserts on how the claim result is USED. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Turns the WhoisFreaks subscription into a stored, graded shortlist of dropped
.comdomains in the project's industries — and the industries around it.Why the first run looked broken
The Expired Domains tab returned five vending companies. The vocabulary was five words taken from the project's own keywords (
breakroom, office, coffee, dallas, vending). A client's keywords can only describe the client's own vertical, so a harvest built from them returns more of that vertical — by construction.Adjacent terms now reach the industries around the business: the verticals its customers are in, the venues it operates in, and topics it could credibly publish about. A vending operator serves schools, so an education domain is a legitimate target even though education is not its trade.
Measured on one real day of the feed:
Top matched terms:
water146,logistics127,school110,hotel99,wellness99,clinic89,fitness86,education52,university31.Two endpoint corrections, found live
The documented endpoint does not exist.
whoisfreaks.com/api/v3/dropped-domainsreturns their marketing site's Next.js 404 page — HTML, identically for every date, which is why it parses as a failure only by luck. The real feed isfiles.whoisfreaks.com/v3.1/download/domainer/dropped, found by rendering their JavaScript docs. A test pins it.And the body is gzipped newline-delimited names, not JSON. Decoded with
DecompressionStream, CRLF-tolerant, TLD filtered client-side — there is no server-side TLD filter, so the download is the whole day (~2 MB, ~240k rows, ~84k of them.com).whois=truereturns 413 "Please upgrade your plans", so the API enforces the subscribed tier. Also pinned.Cost model — deliberately ungated
Harvesting and DR grading are free: a flat subscription and a keyless endpoint. There is no spend gate on either, and adding one would only train the user to click past the gates that do matter. Availability is the exception — 5 APIVerve credits per domain, behind an explicit action that states its cost and is capped at 25 per click.
Design notes
datesToHarvestreturns nothing once a day is stored, so real work happens about once a day per project. One day per tick — a 2 MB file should not be held several at a time..coffeedomain hits "coffee" and one TLD floods the harvest. Longest term wins, so "breakroom" explains a row rather than "room". Terms under four characters are ignored because "co" matches a large share of the internet.domain_ratingnull means UNGRADED, never "no authority" — a real 0 stores as 0. Only a thrown lookup leaves a row for a later tick to retry.Refactor included
resolveDomainRatingmoved out ofserverFunctions/ahrefs.tsinto a lib taking its cache as a parameter. That file statically importscloudflare:workers, so the rating logic could be reached from neither a cron nor a test. Behaviour is unchanged — same cache prefix, TTL, and 0-vs-null rule.Verification
pnpm ci:checkclean;pnpm testgreen — 315 files, 3181 tests.0041already applied to prod D1 ahead of this code; additive and nullable..com→ 1,234 matches, each attributed to the term that hit.Not verified
The panel is not browser-verified — the harvest table is empty until this deploys and a day is pulled.
🤖 Generated with Claude Code