Skip to content

HTML API: add a parsing benchmark tool under tests/performance/html-api - #97

Open
sirreal wants to merge 3 commits into
trunkfrom
html-api/benchmark
Open

sirreal wants to merge 3 commits into
trunkfrom
html-api/benchmark

Conversation

@sirreal

@sirreal sirreal commented Oct 2, 2026

Copy link
Copy Markdown
Owner

Add a parsing benchmark for the HTML API under tests/performance/html-api/, so a change to WP_HTML_Tag_Processor or WP_HTML_Processor can be measured against a baseline checkout with a confidence interval instead of a pair of wall-clock numbers.

The timed loop constructs a processor and calls next_token() until it returns false, for the Tag Processor and for the HTML Processor through create_full_parser(). Nothing else is timed: no get_tag(), no attribute reads, no modifications. A document the HTML Processor cannot finish is still timed but reported separately and left out of the aggregate.

php tests/performance/html-api/bench.php --base ~/a8c/wordpress-develop/trunk --head .

How a run works:

  • bench.php starts one worker.php process per checkout. Each worker loads that checkout's src/wp-includes/html-api/ with stubs for the few WordPress functions it calls (__(), esc_url(), wp_kses_uri_attributes() and similar), so WordPress itself is not loaded and --base may point at a checkout that predates this tool.
  • Samples alternate between the two trees, and the tree that goes first alternates each sample, so machine drift lands on both sides of the ratio.
  • The report gives, per document and parser, the median per tree, the change, a 95% interval on the ratio of medians, a verdict (faster, slower, no difference), and the coefficient of variation of each tree's samples; then a geometric-mean overall line per parser with its own interval. Output is a table, markdown or JSON; --save keeps every sample, and compare.php compares two saved runs.

Documents:

  • lib/class-benchmark-synthetic.php generates 13 shapes at a target size (default 200 KB) from a seed: dense tags, long text, entities, multi-byte text, many comments, attribute-heavy tags, deep nesting, script and style bodies, whitespace runs, inline SVG and MathML, misnested formatting, tables, and a block-themed WordPress post. Output is byte-identical for a given seed on PHP 8.4 and 8.5, and every shape parses to the end on trunk's HTML Processor. The tables and formatting-adoption shapes stay inside what trunk supports, since foster parenting and the general adoption agency case still bail.
  • corpus/manifest.json names 20 public pages with sizes and hashes (the HTML standard's parsing section, the W3C HTML5 syntax chapter, RFC 9110, seven Wikipedia articles in four scripts pinned by revision, five WordPress-rendered pages, an MDN reference, a php.net manual page, Hacker News), plus the 15 MB single-page HTML standard behind --include-optional. fetch-corpus.php restores them into the gitignored corpus/real/; nothing third-party is committed. --corpus <dir> reads any directory of HTML files.

One measurement result shaped the defaults. Two PHP processes running identical code differ by a constant 0.3% to 2% for the life of the process, and the second-started one is the slower, so interleaving within a process pair cannot cancel it: a one-round A/A run reported slower at +1.8% on every HTML Processor row with CVs of 0.2%. The tool therefore repeats the schedule over --rounds fresh worker pairs and takes a t-interval over the per-round ratios; the default is 4 rounds of 10 samples because an odd count leaves the start order unbalanced (3 rounds gave 8 of 64 rows slower, all the same sign). With the defaults, an A/A run over all 32 documents reports no difference on all 64 rows, with both overall intervals centred on zero; per-row resolution is about ±1%, and a full run takes about 34 minutes on an M-series laptop.

As a first use, replacing the two length-1 strspn() set checks in the Tag Processor's parse loop with strpos() on the set measured -1.2% [-2.3%, -0.2%] on the Tag Processor and -0.6% [-1.5%, +0.4%] on the HTML Processor, every point estimate negative, with identical token output. That change is not in this PR.

Limitations: the WHATWG pages are fetched unpinned because the site's robots.txt disallows /commit-snapshots/, and the Wikipedia pages carry per-request values so their hashes are informational; fetch-corpus.php reports drift on unpinned entries rather than failing. SPEC.md is the design document the tool was built from and can be dropped before any upstream PR.

Trac ticket: none yet; this is a fork PR for review of the tool itself. trac#65967 (the HTML API fuzzing ticket) is the nearest existing ticket.

Use of AI Tools

AI assistance: Yes
Tool(s): Claude Code
Model(s): Claude Fable 5.1
Used for: Implementation, corpus selection, synthetic document generator, documentation and the validation runs, from a design settled in a question round with the author. The author has not yet reviewed the code line by line; the A/A validation and the strspn experiment were run and read by the agent.


This Pull Request is for code review only. Please keep all other discussion in the Trac ticket. Do not merge this Pull Request. See GitHub Pull Requests for Code Review in the Core Handbook for more details.

A PHP CLI that times WP_HTML_Tag_Processor and WP_HTML_Processor parsing
(next_token() to the end of the document) over a corpus of real public
pages and seeded synthetic documents, and compares two checkouts with
interleaved samples, fresh worker processes per round and confidence
intervals on the ratio of medians.

- bench.php orchestrates one worker per checkout and prints a table,
  markdown or JSON report; compare.php compares two saved runs.
- worker.php loads one checkout's HTML API with function stubs and runs
  timed jobs over a JSON line protocol.
- lib/class-benchmark-synthetic.php generates 13 document shapes (dense
  tags, long text, entities, UTF-8, comments, attributes, deep nesting,
  script and style, whitespace, foreign content, formatting, tables, a
  WordPress post), byte-identical for a given seed across PHP versions.
- corpus/manifest.json names twenty public pages with sizes and hashes;
  fetch-corpus.php restores them into the gitignored corpus/real/.

An A/A run with the defaults (4 rounds of 10 samples) reports no
difference on all 64 rows. The README documents the method and why the
round count defaults to an even number.
…e, which this repository uses instead of nested ignore files.
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the props-bot label.

Core Committers: Use this line as a base for the props when committing in SVN:

Props jonsurrell.

To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant