Skip to content

Find lapsed, registerable domains in adjacent industries — no new vendor - #58

Merged
ThinkingSpade merged 1 commit into
mainfrom
feat/acquirable-domain-finder
Aug 20, 2026
Merged

ThinkingSpade merged 1 commit into
mainfrom
feat/acquirable-domain-finder

Conversation

@ThinkingSpade

Copy link
Copy Markdown
Owner

Answers the actual ask: open the tab, see expired domains you can buy that are relevant to the industry.

Why the existing sources could never do this

They surface domains already linking to a competitor or ranking for a tracked keyword. That is relevance by construction — and it is also the ceiling. A vending operator gets five vending companies, because that is the entire population those sources can reach. It also mostly returns domains that are still owned, not buyable.

The free path, and why it is ordered this way

  1. Generate names from the client's vocabulary crossed with adjacent industries — free
  2. Ask Wayback which ever hosted a site — free, no key
  3. Check availability only for those — 5 credits each

Step 2 is what makes a result an expired domain rather than an unregistered string, and putting it before step 3 is what makes this affordable. Reversed, you would bill for every generated name, most of which nobody ever registered.

On Delio's real keywords the generator produces nutritionvending.com, foodbreakroom.com, dallasnutrition.com, snackvending.com — the adjacency that was missing.

The alternatives, checked

  • DomCop — $816–$1,416/year and no API on any tier. Manual CSV export only.
  • ICANN CZDS — free and official, but ships whole TLD zone files. .com alone is 100M+ domains, multi-gigabyte, daily. A Worker cannot ingest that.
  • park.io — no public API, mostly non-.com.

A cost bug this surfaced, and the guard for it

A live probe of archive.org returned 429 Too Many Requests. That matters because inconclusive archive answers are deliberately kept — treating them as "never existed" would discard real targets over someone else's outage. But without a ceiling, a throttled archive would have sent every generated name to the paid availability check: 60 names × 5 credits = 300, silently.

So:

  • hard ceiling on billed availability checks per run (MAX_AVAILABILITY_CHECKS)
  • confirmed history is spent first; inconclusive names only use leftover budget
  • answers cached in KV for 30 days — "did this ever host a site" is near-immutable, and every hit is one less request against a free service
  • a run whose archive checks all failed reports archiveUnavailable instead of quietly returning less

Adjacent terms

One small model call per run via the already-configured OpenRouter key. The parser is strict and separately tested — a model that decides to explain itself must not turn a sentence into a domain name. No key configured, or the call fails: the search still runs on the client's own vocabulary, just narrower. Never throws; it is enrichment, not a precondition.

Verification

  • pnpm ci:check clean; pnpm test green — 311 files, 3143 tests.
  • 32 new unit tests: seed-term extraction and filler rejection, name generation and exclusion, the strict model-output parser, the Wayback client including the 429 path and its caching, and the spend guards.
  • Generator sanity-checked against Delio's five real tracked keywords.
  • e2e/expired-domains.spec.ts still passes — panel mounts idle, quotes its ceiling, zero metered requests before a click.

Not verified

No end-to-end run against live archive.org plus live availability — that bills, and archive.org is currently throttling this IP. The pieces are individually proven; the joined-up run is not.

🤖 Generated with Claude Code

Answers the actual ask: open the tab and see expired domains you can BUY that
are relevant to the industry. The graph sources cannot do this. They only
surface domains already linking to a competitor or ranking for a tracked
keyword, which by construction returns more of the same vertical -- a vending
operator got five vending companies.

This generates names from the client's own vocabulary crossed with ADJACENT
industries, then keeps the ones that had a site and are free to register today.
The cost ordering is the design: generation is free, the archive check is free,
and only names that once hosted a site reach the billed availability lookup.
Reversing those steps would bill for every generated name, most of which nobody
ever registered.

DomCop was the alternative and it is 16/yr with no API on any tier; ICANN CZDS
is free but ships whole multi-gigabyte TLD zone files, which a Worker cannot
ingest. The Wayback availability endpoint is free, needs no key, and answers the
one question that separates a dropped domain from an unregistered string.

It also rate-limits: a live probe returned 429. That exposed a real cost bug
here -- inconclusive archive answers are deliberately kept rather than treated
as "never existed", so a throttled archive would have sent every generated name
to the paid availability check. There is now a hard ceiling on billed checks per
run, confirmed history is spent first, answers are cached in KV for 30 days, and
a run whose archive checks all failed says so instead of quietly returning less.
@ThinkingSpade
ThinkingSpade merged commit 4506b5f into main Aug 20, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant