Pravda should retrieve non-HTML resources without starting a browser unnecessarily, while retaining Chromium as a fallback for blocked or browser-dependent sources.
Add an auto retrieval mode:
- Attempt the URL through direct HTTP.
- Validate the response.
- Escalate to Chromium navigation when direct retrieval fails or returns unusable content.
- Return the first usable result and retain metadata for all attempts.
Initial resource support should include JSON, XML, CSV, spreadsheets, PDFs, and arbitrary downloads.
Behavior
- Store response bodies unchanged in content-addressed storage.
- Record requested and final URL, status, headers, media type, charset, size, retrieval mode, and capture time.
- Follow redirects.
- Enforce a configurable maximum response size.
- Support expected media-type validation.
- Escalate on exhausted transport errors, blocked responses, recognized challenge pages, or content-validation failure.
- Do not escalate genuine
404 or 410 responses.
- Record direct HTTP and Chromium attempts separately under the resulting snapshot.
- Apply the same state, failure-cause, retry, and stale-fallback semantics as browser snapshots.
The first version should support GET only. Request bodies, authentication, persistent sessions, and browser-to-HTTP cookie transfer are out of scope.
Questions
- Should
auto, http, and browser be explicit request modes?
- Which statuses and validation failures trigger Chromium escalation?
- How are multiple attempts represented without duplicating the logical snapshot?
- Should direct HTTP retrieval produce the same snapshot response model with browser-specific artifacts set to null?
- Which response headers are safe and necessary to retain?
Pravda should retrieve non-HTML resources without starting a browser unnecessarily, while retaining Chromium as a fallback for blocked or browser-dependent sources.
Add an
autoretrieval mode:Initial resource support should include JSON, XML, CSV, spreadsheets, PDFs, and arbitrary downloads.
Behavior
404or410responses.The first version should support
GETonly. Request bodies, authentication, persistent sessions, and browser-to-HTTP cookie transfer are out of scope.Questions
auto,http, andbrowserbe explicit request modes?