Embedded pieces of HTML, such as comments, the descriptions in feeds or the fields of a CMS, read into Hiccup and rendered from there.
TODO: an optional report of what parsing discarded and sanitizing removed, e.g. in metadata, so that a validator can tell the author of the HTML what a client won't show.
TODO: in the browser, an option to parse with the browser's own DOMParser and read its DOM as Hiccup, as hickory does, which would keep the tokenizer, the tree builder and the table of character references out of the bundle. It would come at a cost:
Embedded pieces of HTML, such as comments, the descriptions in feeds or the fields of a CMS, read into Hiccup and rendered from there. TODO: an optional report of what parsing discarded and sanitizing removed, e.g. in metadata, so that a validator can tell the author of the HTML what a client won't show. TODO: in the browser, an option to parse with the browser's own DOMParser and read its DOM as Hiccup, as hickory does, which would keep the tokenizer, the tree builder and the table of character references out of the bundle. It would come at a cost: - the result could differ from the JVM's and Node's, since a browser follows every rule of tree construction, e.g. it carries formatting on past a misnested end tag, moves the content of tables, and rearranges the html, head and body of a whole page - Node has no DOM, so the tokenizer would still be needed there - DOMParser reads a whole document, so a fragment would be the children of its body, and :max-depth would have to be applied while reading the DOM
The named character references of HTML, such as & and é, by the table in section 13.5 of the HTML Standard.
The table is the entities.json that WHATWG publishes with the standard, as of 2026-10-08, under CC BY 4.0.
Only the names of HTML 4.01 are kept, with ' of XML and the legacy names that HTML reads without a semicolon. The rest came from MathML, e.g. ≂̸, and stay text: they would make the table many times as large, and embedded HTML hardly ever uses them.
TODO: read the names from MathML too, by an option or a namespace of their own, for HTML that writes mathematics with them.
The named character references of HTML, such as & and é, by the table in section 13.5 of the HTML Standard. The table is the entities.json that WHATWG publishes with the standard, as of 2026-10-08, under CC BY 4.0. Only the names of HTML 4.01 are kept, with ' of XML and the legacy names that HTML reads without a semicolon. The rest came from MathML, e.g. ≂̸, and stay text: they would make the table many times as large, and embedded HTML hardly ever uses them. TODO: read the names from MathML too, by an option or a namespace of their own, for HTML that writes mathematics with them.
HTML text as tokens, by the tokenizer of the HTML Standard, section 13.2.5, as of 2026-10-08, but without parse errors.
The tokenizer is for HTML embedded in other content, so it leaves out what only whole documents and SVG or MathML need. Each part ends where the standard ends it, so no other token changes:
<? starts a bogus comment, as before the processing instructions of 13.2.5.72 to 13.2.5.76
TODO: read these too, by an option or a namespace of their own, for a parser of whole documents or of SVG and MathML.
HTML text as tokens, by the tokenizer of the HTML Standard, section 13.2.5, as of 2026-10-08, but without parse errors. The tokenizer is for HTML embedded in other content, so it leaves out what only whole documents and SVG or MathML need. Each part ends where the standard ends it, so no other token changes: - a DOCTYPE runs to the next >, without its name, identifiers or the force-quirks flag, 13.2.5.53 to 13.2.5.68 - <? starts a bogus comment, as before the processing instructions of 13.2.5.72 to 13.2.5.76 - a CDATA section is a bogus comment, as it is outside SVG and MathML, without the CDATA states of 13.2.5.69 to 13.2.5.71 TODO: read these too, by an option or a namespace of their own, for a parser of whole documents or of SVG and MathML.
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |