Embedded pieces of HTML, such as comments, the descriptions in feeds or the fields of a CMS, read into Hiccup and rendered from there.
TODO: an optional report of what parsing discarded and sanitizing removed, e.g. in metadata, so that a validator can tell the author of the HTML what a client won't show.
TODO: in the browser, an option to parse with the browser's own DOMParser and read its DOM as Hiccup, as hickory does, which would keep the tokenizer, the tree builder and the table of character references out of the bundle. It would come at a cost:
Embedded pieces of HTML, such as comments, the descriptions in feeds or the fields of a CMS, read into Hiccup and rendered from there. TODO: an optional report of what parsing discarded and sanitizing removed, e.g. in metadata, so that a validator can tell the author of the HTML what a client won't show. TODO: in the browser, an option to parse with the browser's own DOMParser and read its DOM as Hiccup, as hickory does, which would keep the tokenizer, the tree builder and the table of character references out of the bundle. It would come at a cost: - the result could differ from the JVM's and Node's, since a browser follows every rule of tree construction, e.g. it carries formatting on past a misnested end tag, moves the content of tables, and rearranges the html, head and body of a whole page - Node has no DOM, so the tokenizer would still be needed there - DOMParser reads a whole document, so a fragment would be the children of its body, and :max-depth would have to be applied while reading the DOM
The named character references of HTML, such as & and é, by the table in section 13.5 of the HTML Standard.
The table is the entities.json that WHATWG publishes with the standard, as of 2026-10-08, under CC BY 4.0.
Only the names of HTML 4.01 are kept, with ' of XML and the legacy names that HTML reads without a semicolon. The rest came from MathML, e.g. ≂̸, and stay text: they would make the table many times as large, and embedded HTML hardly ever uses them.
TODO: read the names from MathML too, by an option or a namespace of their own, for HTML that writes mathematics with them.
The named character references of HTML, such as & and é, by the table in section 13.5 of the HTML Standard. The table is the entities.json that WHATWG publishes with the standard, as of 2026-10-08, under CC BY 4.0. Only the names of HTML 4.01 are kept, with ' of XML and the legacy names that HTML reads without a semicolon. The rest came from MathML, e.g. ≂̸, and stay text: they would make the table many times as large, and embedded HTML hardly ever uses them. TODO: read the names from MathML too, by an option or a namespace of their own, for HTML that writes mathematics with them.
HTML text as tokens, by the tokenizer of the HTML Standard, section 13.2.5, as of 2026-10-08, but without parse errors.
The tokenizer is for HTML embedded in other content, so it leaves out what only whole documents and SVG or MathML need. Each part ends where the standard ends it, so no other token changes:
<? starts a bogus comment, as before the processing instructions of 13.2.5.72 to 13.2.5.76
TODO: read these too, by an option or a namespace of their own, for a parser of whole documents or of SVG and MathML.
HTML text as tokens, by the tokenizer of the HTML Standard, section 13.2.5, as of 2026-10-08, but without parse errors. The tokenizer is for HTML embedded in other content, so it leaves out what only whole documents and SVG or MathML need. Each part ends where the standard ends it, so no other token changes: - a DOCTYPE runs to the next >, without its name, identifiers or the force-quirks flag, 13.2.5.53 to 13.2.5.68 - <? starts a bogus comment, as before the processing instructions of 13.2.5.72 to 13.2.5.76 - a CDATA section is a bogus comment, as it is outside SVG and MathML, without the CDATA states of 13.2.5.69 to 13.2.5.71 TODO: read these too, by an option or a namespace of their own, for a parser of whole documents or of SVG and MathML.
URLs resolved against a base URL as a browser resolves them, and checked against a set of allowed schemes.
A URL is resolved by RFC 3986, section 5, with the changes of the WHATWG URL Standard that browsers make, as of 2026-10-10:
TODO: the host parser of the URL Standard, which maps a domain by IDNA, reads the numbers of an IPv4 address in any base, compresses an IPv6 address and rejects the forbidden code points, and the drive letters of file URLs.
URLs resolved against a base URL as a browser resolves them, and checked against a set of allowed schemes. A URL is resolved by RFC 3986, section 5, with the changes of the WHATWG URL Standard that browsers make, as of 2026-10-10: - tabs and line breaks are dropped, and so are the control characters and spaces at the ends - in a URL of a special scheme, e.g. http, a backslash is a slash, the path is at least /, the host is in lower case and a default port is left out - a URL of the scheme of its base and no authority is relative, e.g. http:g, and a URL of a special scheme reads its authority past any number of slashes - a dot segment can be percent-encoded, e.g. %2e%2e - the characters of the percent-encode sets are percent-encoded as UTF-8 TODO: the host parser of the URL Standard, which maps a domain by IDNA, reads the numbers of an IPv4 address in any base, compresses an IPv6 address and rejects the forbidden code points, and the drive letters of file URLs.
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |