This is a Clojure and ClojureScript library for working with embedded pieces of HTML, e.g. comments, the descriptions in RSS feeds, or the fields of a CMS.
The same code runs on the JVM, in Node and in the browser, with the same results on each.
This library was spun out of podcast-clj, a library for making podcast software, which uses it to read show notes and comments. Like podcast-clj, it was developed with assistance from frontier LLMs.
It requires Clojure 1.11+ and Java 11+. For the latest release, add it
from Clojars to the
:deps in your deps.edn:
dk.simongray/html-pieces {:mvn/version "0.4.0"}
For changes that aren't released yet, use the SHA of the latest commit on
master instead:
dk.simongray/html-pieces
{:git/url "https://github.com/simongray/html-pieces"
:git/sha "…"}
For ClojureScript, shadow-cljs only reads Git dependencies from
deps.edn, so also set :deps true in your shadow-cljs.edn.
Then try it on a piece of HTML:
(require '[dk.simongray.html-pieces :as html])
(def piece
(str "<h2>Hi!</h2>"
"<p>Nice <a href='https://example.com/' onclick='steal()'>post</a>,"
"<br>thanks."
"<ul><li>one<li>two</ul>"
"<script>track()</script>"))
(html/hiccup piece)
;; => ([:h2 {} "Hi!"]
;; [:p {} "Nice " [:a {:href "https://example.com/"} "post"] ","
;; [:br {}] "thanks."]
;; [:ul {} [:li {} "one"] [:li {} "two"]])
(html/text piece)
;; => "Hi!\n\nNice post,\nthanks.\n\n- one\n- two"
(html/markdown piece)
;; => "## Hi!\n\nNice [post](https://example.com/),\\\nthanks.\n\n- one\n- two"
(html/sanitize piece)
;; => "<h2>Hi!</h2><p>Nice <a href=\"https://example.com/\">post</a>,<br>thanks.</p><ul><li>one</li><li>two</li></ul>"
Both hiccup and sanitize parse HTML and keep only what's safe, so the
<script> and the onclick are gone, and the unclosed <p> and <li>
are closed where a browser closes them. Plain text works too: it becomes
paragraphs and line breaks, so you can give them any field without
checking what's in it first. Text with a tag in it is read as HTML,
though, so for a field that's always plain text, e.g. a summary that can
hold a <, pass {:plain-text? true}:
(html/hiccup "Use <b> for bold." {:plain-text? true})
;; => ([:p {} "Use <b> for bold."])
The result of hiccup goes straight into a Hiccup renderer. Its text is
decoded, e.g. < becomes <, so use a renderer that escapes text
again, such as Replicant, Reagent or hiccup2.core/html (the older
hiccup.core/html doesn't escape text).
Every function but markup? also takes Hiccup that you've built or
changed yourself, and reads it as every Hiccup renderer does:
for or map, is spliced in:p#intro.lead is a p with an :id and a :classConventions that only some renderers have, such as event handlers, style maps and a class as a collection, aren't supported.
Every function but markup? takes an optional map of options as its last
argument. The functions share one set of options, so you can pass the
same map to all of them. The var html/default-options lists every
option with its default. For example, :links? makes text show the URL
of each link:
(html/text piece {:links? true})
;; => "Hi!\n\nNice post (https://example.com/),\nthanks.\n\n- one\n- two"
Relative URLs are left out, since a browser would resolve them against
your page rather than the page the HTML came from. If you know that page,
pass its URL as :base-url:
(html/hiccup "<a href='/about'>About</a>"
{:base-url "https://example.com/blog/"})
;; => ([:a {:href "https://example.com/about"} "About"])
For a URL outside HTML, e.g. the link of a feed, resolve in the
dk.simongray.html-pieces.url namespace does the same.
To change what hiccup keeps, pass your own tables. An image loads from
its own server, which can see who views it. To keep images from loading,
take :src out of the allowed attributes of :img:
(def attributes (update html/allowed-attributes :img disj :src))
(html/hiccup "<img src='https://example.com/cat.png' alt='A cat'>"
{:allowed-attributes attributes})
;; => ([:img {:alt "A cat"}])
A link keeps its rel, but not the link types that would speak for your
page, which html/dropped-rels lists: me, author, license, and the
endpoints of Webmention, IndieAuth and the like. RelMeAuth and Mastodon
take a rel="me" anywhere on a page as another profile of its owner, so
a comment with one could sign its writer in as you:
(html/hiccup "<a rel='me nofollow' href='https://mastodon.example/@mallory'>Hi</a>")
;; => ([:a {:rel "nofollow" :href "https://mastodon.example/@mallory"} "Hi"])
For HTML that you trust, e.g. the links of your own profile, pass
:dropped-rels #{}.
To change an element rather than keep it or leave it out, pass an
:element-fn. It's given each element that's kept, once its attributes
and children are safe, and returns the Hiccup to put in its place. For
example, to keep the headings of a reply from another site out of your
page's outline, and to mark its links as user content:
(defn reply-element
"The `element` of a reply as your page shows it."
[[tag attrs & children :as element]]
(case tag
(:h1 :h2 :h3) (into [:p {}] children)
:a (assoc-in element [1 :rel] "nofollow ugc")
element))
(html/sanitize "<h1>Hi</h1><p>See <a href='https://example.com/'>this</a>"
{:element-fn reply-element})
;; => "<p>Hi</p><p>See <a href=\"https://example.com/\" rel=\"nofollow ugc\">this</a></p>"
What it returns is sanitized again, so it can't bring in an element or
an attribute that isn't allowed. To keep it as it is, e.g. with a class
that isn't allowed, also pass :trusted-element-fn? true. For the text
or Markdown of the result, give text or markdown what hiccup
returns.
To read HTML without sanitizing it, e.g. to find the microformats of a
page, use parse, and to write it, use serialize. Neither is safe for
HTML that you don't trust, so render what hiccup or sanitize gives
instead.
In the tree that parse gives, each element is a vector of the tag, an
attribute map and the children, and text is a string:
(html/parse "<p class=x>See <a href=/about>this</a><script>track()</script>")
;; => ([:p {:class "x"} "See " [:a {:href "/about"} "this"]
;; [:script {} "track()"]])
For Hiccup that you've built, parse gives the same shape, so that you
can walk any Hiccup the same way.
Given HTML or Hiccup, serialize writes it as a browser writes the
innerHTML of an element:
(html/serialize "<P class=x>Hi<br/>there<script>track()</script>")
;; => "<p class=\"x\">Hi<br>there<script>track()</script></p>"
A tag in HTML can hold a dot or a #, e.g. that of the custom element
<x-a.b>, but in Hiccup they're shorthand for a class or an id. When you
give the Hiccup of parse to serialize or another function, pass
{:shorthand? false}, so that each tag is read as it's written:
(html/serialize (html/parse "<x-a.b>t</x-a.b>") {:shorthand? false})
;; => "<x-a.b>t</x-a.b>"
To find what a whole page links to, e.g. the endpoint of Webmention, use
links. It resolves each href against the page's base URL, which
base-url gives, so pass the page's URL as :base-url:
(html/links "<link rel=webmention href=/mention><p>Hi <a href=/about>me</a>"
{:base-url "https://example.com/post" :rel "webmention"})
;; => ([:link {:rel "webmention" :href "https://example.com/mention"}])
The :rel option takes the links of one link type, in any case. The
:link-tags option gives the elements that count, a, area and link
by default. Pass :body? false to take only the links in the head of the
page, e.g. for IndieAuth, since a comment on the page can add a link to
its body.
For more of a page, head gives the elements in its head, also of a page
without a head tag, and elements gives all its elements in document
order.
{:quirks? false}.hiccup and sanitize keep only the tags,
attributes and URLs that a page can render safely. Both text and
markdown leave out the control characters that a terminal would obey,
and markdown also escapes what a Markdown renderer would read as
markup. Only parse and serialize keep HTML as it is..cljc. Its only dependency
is data.json, which reads the table of character references when the
code compiles.See the benchmarks for how html-pieces compares with libraries that do the same jobs, on the JVM, in Node and in the browser.
clojure -X:test # the tests on the JVM
npm install # once, for the Node tests and benchmarks
clojure -M:cljs compile test # the tests in Node
clojure -X:bench # the benchmarks against other libraries
The benchmarks take a few minutes. They print a report, which is kept in
target/bench/ with the tables of doc/benchmarks.md in
benchmarks.md. To run some of them, pass e.g. :only '#{:jvm}',
:areas '#{:parse}' or :timing :quick.
The tests run the tokenizer tests of
html5lib-tests when they're
in dev-resources/html5lib-tests/, which Git ignores. CI clones them into
that folder. To do the same yourself, at the version that the tokenizer
passes:
git clone https://github.com/html5lib/html5lib-tests dev-resources/html5lib-tests
git -C dev-resources/html5lib-tests checkout c777c408b61078ea2eb4acefc2535f54dbc8b28a
The html-pieces project is licensed under the MIT licence. The
table of character references in resources/ is WHATWG's
entities.json, under CC BY 4.0.
Can you improve this documentation?Edit on GitHub
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |