Liking cljdoc? Tell your friends :D

dk.simongray.html-pieces

Embedded pieces of HTML, such as comments, the descriptions in feeds or the fields of a CMS, read into Hiccup and rendered from there.

TODO: an optional report of what parsing discarded and sanitizing removed, e.g. in metadata, so that a validator can tell the author of the HTML what a client won't show.

TODO: in the browser, an option to parse with the browser's own DOMParser and read its DOM as Hiccup, as hickory does, which would keep the tokenizer, the tree builder and the table of character references out of the bundle. It would come at a cost:

  • the result could differ from the JVM's and Node's, since a browser follows every rule of tree construction, e.g. it carries formatting on past a misnested end tag, moves the content of tables, and rearranges the html, head and body of a whole page
  • Node has no DOM, so the tokenizer would still be needed there
  • DOMParser reads a whole document, so a fragment would be the children of its body, and :max-depth would have to be applied while reading the DOM
Embedded pieces of HTML, such as comments, the descriptions in feeds or
the fields of a CMS, read into Hiccup and rendered from there.

TODO: an optional report of what parsing discarded and sanitizing
removed, e.g. in metadata, so that a validator can tell the author of
the HTML what a client won't show.

TODO: in the browser, an option to parse with the browser's own
DOMParser and read its DOM as Hiccup, as hickory does, which would keep
the tokenizer, the tree builder and the table of character references
out of the bundle. It would come at a cost:

- the result could differ from the JVM's and Node's, since a browser
  follows every rule of tree construction, e.g. it carries formatting on
  past a misnested end tag, moves the content of tables, and rearranges
  the html, head and body of a whole page
- Node has no DOM, so the tokenizer would still be needed there
- DOMParser reads a whole document, so a fragment would be the children
  of its body, and :max-depth would have to be applied while reading the
  DOM
raw docstring

dk.simongray.html-pieces.entities

The named character references of HTML, such as & and é, by the table in section 13.5 of the HTML Standard.

The table is the entities.json that WHATWG publishes with the standard, as of 2026-10-08, under CC BY 4.0.

Only the names of HTML 4.01 are kept, with ' of XML and the legacy names that HTML reads without a semicolon. The rest came from MathML, e.g. ≂̸, and stay text: they would make the table many times as large, and embedded HTML hardly ever uses them.

TODO: read the names from MathML too, by an option or a namespace of their own, for HTML that writes mathematics with them.

The named character references of HTML, such as & and é, by
the table in section 13.5 of the HTML Standard.

The table is the entities.json that WHATWG publishes with the standard,
as of 2026-10-08, under CC BY 4.0.

Only the names of HTML 4.01 are kept, with ' of XML and the legacy
names that HTML reads without a semicolon. The rest came from MathML,
e.g. ≂̸, and stay text: they would make the table many
times as large, and embedded HTML hardly ever uses them.

TODO: read the names from MathML too, by an option or a namespace of
their own, for HTML that writes mathematics with them.
raw docstring

dk.simongray.html-pieces.tokenizer

HTML text as tokens, by the tokenizer of the HTML Standard, section 13.2.5, as of 2026-10-08, but without parse errors.

The tokenizer is for HTML embedded in other content, so it leaves out what only whole documents and SVG or MathML need. Each part ends where the standard ends it, so no other token changes:

  • a DOCTYPE runs to the next >, without its name, identifiers or the force-quirks flag, 13.2.5.53 to 13.2.5.68
  • <? starts a bogus comment, as before the processing instructions of 13.2.5.72 to 13.2.5.76

  • a CDATA section is a bogus comment, as it is outside SVG and MathML, without the CDATA states of 13.2.5.69 to 13.2.5.71

TODO: read these too, by an option or a namespace of their own, for a parser of whole documents or of SVG and MathML.

HTML text as tokens, by the tokenizer of the HTML Standard, section
13.2.5, as of 2026-10-08, but without parse errors.

The tokenizer is for HTML embedded in other content, so it leaves out
what only whole documents and SVG or MathML need. Each part ends where
the standard ends it, so no other token changes:

- a DOCTYPE runs to the next >, without its name, identifiers or the
  force-quirks flag, 13.2.5.53 to 13.2.5.68
- <? starts a bogus comment, as before the processing instructions of
  13.2.5.72 to 13.2.5.76
- a CDATA section is a bogus comment, as it is outside SVG and MathML,
  without the CDATA states of 13.2.5.69 to 13.2.5.71

TODO: read these too, by an option or a namespace of their own, for a
parser of whole documents or of SVG and MathML.
raw docstring

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
←Move to previous article
→Move to next article
Ctrl+/Jump to the search field
× close