Repository navigation
Version packages - #6
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR was opened by the Changesets release GitHub action. When you're ready to do a release, you can merge this and the packages will be published to npm automatically. If you're not ready to do a release yet, that's fine, whenever you add more changesets to main, this PR will be updated.
Releases
@agenticschema/browser@0.3.1
Patch Changes
@agenticschema/core@0.3.1
Patch Changes
300e864: Read the ld+json type attribute the way a parser does, and stop building a DOM
for pages that have nothing to gain from one.
Two changes to the same seam, found while measuring where the time from a page
to tools actually goes. On the 177-page corpus it is 93% happy-dom parsing, 6.4%
walking the DOM for markup, and 0.6% everything else — normalizing, mapping,
guarding, all of it.
The string path missed entity-encoded types.
extractscanned for a literaltype="application/ld+json", but the attribute is not always written literally:marmiton.org serves
type="application/ld+json". The DOM path decodesit and finds the block; the string path compared raw text and found nothing, on
13 of 177 corpus pages. The only thing reported was
no-structured-dataatinfolevel, which is what a page with no markup at all looks like, sotoTools(html)returned an empty tool list and nothing said why. The typeattribute is now decoded before it is compared, and matched case-insensitively,
because HTML compares
typethat way in selectors and the two paths have toagree. They now agree on all 177 pages.
A DOM is only built when the page can repay it. The new
needsDocument(html)answers whether anything on the page needs a real tree: microdata anchors on
itemscope, RDFa ontypeof, and JSON-LD needs neither.createServeruses itto pass the HTML string straight through when no DOM is warranted, which is 71%
of the corpus — roughly 2x overall, and 19-70x on pages carrying JSON-LD alone.
Pages with real RDFa, such as MDN's, are unaffected: they genuinely need the
parse. The check errs towards building a DOM, since a false positive costs one
parse that finds nothing while a false negative would drop an entity, and it does
not try to tell a schema.org
typeoffrom another vocabulary's, because thatjudgement needs the parsed value.
Scanning is now linear. The single pattern spanning
<script ...>…</script>restarted at every opening tag that never closed: 20k of them took about four
seconds, on input fetched from whatever URL the caller passed. Positions are now
found once, by a scan that only moves forward, which puts those cases under a
millisecond. Two of the three pathological inputs predate this release and one
came from reading the type out of every script tag rather than only the matching
ones; all three are closed, and a regression test sits beside the equivalent one
for
sanitizeText.@agenticschema/profiles@0.3.1
Patch Changes
@agenticschema/server@0.3.1
Patch Changes
300e864: Read the ld+json type attribute the way a parser does, and stop building a DOM
for pages that have nothing to gain from one.
Two changes to the same seam, found while measuring where the time from a page
to tools actually goes. On the 177-page corpus it is 93% happy-dom parsing, 6.4%
walking the DOM for markup, and 0.6% everything else — normalizing, mapping,
guarding, all of it.
The string path missed entity-encoded types.
extractscanned for a literaltype="application/ld+json", but the attribute is not always written literally:marmiton.org serves
type="application/ld+json". The DOM path decodesit and finds the block; the string path compared raw text and found nothing, on
13 of 177 corpus pages. The only thing reported was
no-structured-dataatinfolevel, which is what a page with no markup at all looks like, sotoTools(html)returned an empty tool list and nothing said why. The typeattribute is now decoded before it is compared, and matched case-insensitively,
because HTML compares
typethat way in selectors and the two paths have toagree. They now agree on all 177 pages.
A DOM is only built when the page can repay it. The new
needsDocument(html)answers whether anything on the page needs a real tree: microdata anchors on
itemscope, RDFa ontypeof, and JSON-LD needs neither.createServeruses itto pass the HTML string straight through when no DOM is warranted, which is 71%
of the corpus — roughly 2x overall, and 19-70x on pages carrying JSON-LD alone.
Pages with real RDFa, such as MDN's, are unaffected: they genuinely need the
parse. The check errs towards building a DOM, since a false positive costs one
parse that finds nothing while a false negative would drop an entity, and it does
not try to tell a schema.org
typeoffrom another vocabulary's, because thatjudgement needs the parsed value.
Scanning is now linear. The single pattern spanning
<script ...>…</script>restarted at every opening tag that never closed: 20k of them took about four
seconds, on input fetched from whatever URL the caller passed. Positions are now
found once, by a scan that only moves forward, which puts those cases under a
millisecond. Two of the three pathological inputs predate this release and one
came from reading the type out of every script tag rather than only the matching
ones; all three are closed, and a regression test sits beside the equivalent one
for
sanitizeText.300e864: Parse with
linkedominstead ofhappy-dom.About four times faster on the pages that need a parser at all, and since
parsing is nearly all of the time from a page to tools, that is most of the
end-to-end difference. Together with only building a DOM when the page carries
microdata or RDFa, the corpus goes from 35.5 ms to 5.1 ms a page.
Speed is the smaller half.
happy-domemulates a browser, so it can load what apage points at — that was closed by turning six settings off and bounding the
timers, which works, but it is a defence that has to be remembered and sits one
careless option away from coming back.
linkedomis a parser and nothing else:htmlparser2,css-selectandcssomunderneath it, no HTTP client anywhere inthe tree, no script evaluation, no timers to bound. The socket-level test
asserting that parsing fetches nothing now passes because there is nothing that
could fetch, rather than because six flags are set correctly.
happy-domis nolonger a dependency of this package; it stays a dev dependency, where simulating
a browser is the point.
Checked before the swap rather than after: all 177 corpus pages were run through
toToolswith both parsers and compared on tool names, descriptions, inputschemas, annotations and diagnostic codes. All 177 agreed.
parseDocumentalso stopped throwing on input with no elements in it.linkedomleaves such a document without a root element, and
headandbodythen throwrather than being absent — an empty response, a plain-text error page or a bare
doctype was enough to do it. It now builds the shell a browser would, keeping the
text of a page that is only text.
Dropping
happy-domalso uncovered something it had been hiding. This packageuses
process,Bufferandnode:httpacrosscli.ts, and never declared@types/nodefor any of them — the types resolved only becausehappy-domdepended on
@types/nodeand npm nested a copy inside this workspace. Removingit took them away and the declaration build stopped compiling a global it had
always used.
@types/nodeis now a declared dev dependency and the tsconfignames it outright, so the build no longer rests on what a runtime dependency
happens to drag in.
The tradeoff, stated plainly:
htmlparser2is not a spec-compliant HTML5 treebuilder, so on badly broken markup it can build something a browser would not.
The corpus suite is what guards that, and
npm run test:corpusis the gate forany future change here.
Updated dependencies [300e864]