Skip to content

info providers: optionally fetch a part's data when its page is first viewed (to reduce cost/api calls) - #1581

Open
samyk wants to merge 6 commits into
Part-DB:masterfrom
samyk:provider-fetch-on-view
Open

samyk wants to merge 6 commits into
Part-DB:masterfrom
samyk:provider-fetch-on-view

Conversation

@samyk

@samyk samyk commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Looking up every part of a large inventory at an info provider up front is slow, and for paid providers such as Canopy expensive, while most of those parts are never looked at again. This adds an opt-in mode where a part's data is fetched the first time somebody opens its info page, and only once.

How it behaves:

  • The page renders as usual and then asks the server for the data itself, so opening it is not slowed down. It shows "Fetching product data from …", reloads with the new data once it is there, or says that the daily limit is reached or the lookup failed.
  • Only parts without an info provider reference are considered. Only missing data is filled in (manufacturer, MPN, notes, pictures, parameters, mass, manufacturing status, a price if the part has none); nothing existing is changed, since nobody reviews the result. The part then gets the provider reference, which marks it as done and lets the normal info provider tools update it later.
  • A failed lookup is remembered for a day, so a dead product is not requested on every view. A short-lived lock keeps two viewers of the same part from both triggering a request, and the session is released during the fetch so the user's other pages are not blocked.

How a part is matched to a provider (the check done on every page view never contacts a provider):

  1. An order detail's supplier product URL that the provider recognises (URLHandlerInfoProviderInterface). Canopy is not a URL handler, so its ASIN extraction for the configured marketplace is handled as a special case.
  2. An order detail whose supplier name equals the provider's name (ignoring case, spaces and punctuation) and that has a supplier part number: the provider is searched for that number, and a result is accepted only if its provider ID, MPN or order number equals it (ignoring case, leading zeros, a leading #, and a vendor prefix set apart by a separator or equal to the start of the provider's name). No exact result is a failed lookup; it never guesses. Several exact results are accepted only if they all carry the same MPN.

Settings:

  • Info provider general settings: "Fetch data when a part is viewed" (a multi-select of the active providers; PROVIDER_FETCH_ON_VIEW=key1,key2) and "Max. lookups per day when viewing parts" (PROVIDER_FETCH_ON_VIEW_DAILY_LIMIT, default 100, 0 = unlimited; per provider, counting parts looked up). Empty by default, which is the stock behaviour.
  • Canopy keeps its own "Fetch data when a part is viewed" switch and daily limit in its provider settings.

Implementation: ProviderOnViewMatcher (which providers are enabled, matching, exact-result selection), ProviderOnViewFetcher (lock, daily cap, failure memo, filling missing data), a CSRF-protected POST /{id}/fetch_on_view route that requires read permission on the part (the administrator opts in with the setting, and only missing data is added), a small Stimulus controller and template for the notice, translations and a docs section. 37 unit tests cover the matching.

The branch has two commits: the first adds the mode for Canopy only, the second generalises it to any provider.

Notes for review:

samyk added 4 commits October 3, 2026 14:49
Canopy bills per request, so looking up every Amazon part of a large
inventory up front is expensive, while most of them are never looked at
again. A new option in the Canopy provider settings fetches a part's data
only when somebody opens its info page, and only once.

The page is rendered as usual and then asks the server for the data
itself, so opening it is not slowed down by the request. It shows that
the data is being fetched and reloads with the new data once it is there;
if the daily limit is reached or Canopy fails, it says so instead.

This applies to parts without an info provider reference that have an
orderdetail linking to a product page of the configured Amazon
marketplace. Only missing data is filled in (manufacturer, notes,
pictures, parameters, a price if the part has none); nothing existing is
changed, since nobody reviews the result. The part then gets a Canopy
provider reference, which marks it as done and lets the normal info
provider tools update it later.

Costs are bounded by a second setting, the maximum number of such
requests per 24 hours (default 100). A failed lookup is remembered for a
day so a dead ASIN is not paid for on every view, and a short-lived lock
keeps two viewers of the same part from both triggering a request.
Fetching data on the first view of a part only existed for Amazon parts
via Canopy. The same reasoning applies to the web stores: looking up all
parts of an inventory at once means hundreds of requests to the website
of a small store, while most parts are never looked at again. So the
mechanism is no longer tied to Canopy and can be switched on per provider.

A new general info provider setting (or the PROVIDER_FETCH_ON_VIEW
environment variable, a comma separated list of provider keys) selects
the providers which do this. It is empty by default, so nothing changes
unless it is set. Canopy keeps its own switch and its own daily limit in
its settings and behaves as before; it is now just one user of the
generic code.

A part without a provider reference is recognized by its orderdetails:
first by a product URL the provider handles, which gives the product ID
directly; otherwise by a supplier named like the provider together with a
supplier part number. In the second case the provider is searched for the
number and a result is only used if its ID, part number or order number
is exactly that number (ignoring case and a vendor prefix like ADA4062 or
DEV-13975). Nobody reviews what is added, so a similar product is never
taken: no exact result counts as a failed lookup. Deciding whether a
fetch is needed never contacts a provider, as it runs on every page view.

The safety properties stay the same: only missing data is filled in, the
part gets the provider reference afterwards, a failed lookup (including
a provider refusing or pausing its requests) is remembered for a day and
reported on the page instead of raising an error, a lock prevents
parallel lookups of one part, and the lookups per provider are capped
per 24 hours (PROVIDER_FETCH_ON_VIEW_DAILY_LIMIT, default 100).

The existing orderdetail the part was recognized by is completed with
the product URL and prices instead of adding a second one for the same
store. As stores pace their requests, a fetch can take much longer than
a Canopy request: the session is released while it runs, so the user's
other pages are not blocked, and the lock and the page wait longer.
@codecov

codecov Bot commented Oct 4, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 45.83333% with 156 lines in your changes missing coverage. Please review.
✅ Project coverage is 64.13%. Comparing base (a7f2d49) to head (bf8de1a).

Files with missing lines Patch % Lines
...vices/InfoProviderSystem/ProviderOnViewFetcher.php 3.33% 145 Missing ⚠️
src/Controller/PartController.php 15.38% 11 Missing ⚠️
Additional details and impacted files
@@             Coverage Diff              @@
##             master    #1581      +/-   ##
============================================
- Coverage     64.46%   64.13%   -0.33%     
- Complexity    10210    10342     +132     
============================================
  Files           762      766       +4     
  Lines         32711    32999     +288     
============================================
+ Hits          21087    21164      +77     
- Misses        11624    11835     +211     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

samyk added 2 commits October 4, 2026 17:08
The coverage upload step failed in some of the test jobs of the last run
(the tests themselves passed); an empty commit to run them again.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant