This is a Clojure and ClojureScript library for working with embedded pieces of HTML, e.g. comments, the descriptions in RSS feeds, or the fields of a CMS.
The same code runs on the JVM, in Node and in the browser, with the same results on each.
This library was spun out of podcast-clj, a library for making podcast software, which uses it to read show notes and comments. Like podcast-clj, it was developed with assistance from frontier LLMs.
It requires Clojure 1.11+ and Java 11+. For the latest release, add it
from Clojars to the
:deps in your deps.edn:
dk.simongray/html-pieces {:mvn/version "0.1.0"}
For changes that aren't released yet, use the SHA of the latest commit on
master instead:
dk.simongray/html-pieces
{:git/url "https://github.com/simongray/html-pieces"
:git/sha "…"}
For ClojureScript, shadow-cljs only reads Git dependencies from
deps.edn, so also set :deps true in your shadow-cljs.edn.
Then try it on a piece of HTML:
(require '[dk.simongray.html-pieces :as html])
(def piece
(str "<h2>Hi!</h2>"
"<p>Nice <a href='https://example.com/' onclick='steal()'>post</a>,"
"<br>thanks."
"<ul><li>one<li>two</ul>"
"<script>track()</script>"))
(html/hiccup piece)
;; => ([:h2 {} "Hi!"]
;; [:p {} "Nice " [:a {:href "https://example.com/"} "post"] ","
;; [:br {}] "thanks."]
;; [:ul {} [:li {} "one"] [:li {} "two"]])
(html/text piece)
;; => "Hi!\n\nNice post,\nthanks.\n\n- one\n- two"
(html/markdown piece)
;; => "## Hi!\n\nNice [post](https://example.com/),\\\nthanks.\n\n- one\n- two"
(html/sanitize piece)
;; => "<h2>Hi!</h2><p>Nice <a href=\"https://example.com/\">post</a>,<br>thanks.</p><ul><li>one</li><li>two</li></ul>"
Both hiccup and sanitize parse HTML and keep only what's safe, so the
<script> and the onclick are gone, and the unclosed <p> and <li>
are closed where a browser closes them. Plain text works too: it becomes
paragraphs and line breaks, so you can give them any field without
checking what's in it first. They also take Hiccup that you've built or
changed yourself.
The result of hiccup goes straight into a Hiccup renderer. Its text is
decoded, e.g. < becomes <, so use a renderer that escapes text
again, such as Replicant, Reagent or hiccup2.core/html. The older
hiccup.core/html doesn't escape text.
NOTE: Relative URLs are left out, since a browser would resolve them
against your page rather than the page the HTML came from. If you know
that page, pass a :url-fn that makes them absolute:
(def page (java.net.URI. "https://example.com/blog/"))
(html/hiccup "<a href='/about'>About</a>" {:url-fn #(str (.resolve page %))})
;; => ([:a {:href "https://example.com/about"} "About"])
Every function but markup? takes an optional map of options as its last
argument. The functions share one set of options, so you can pass the
same map to all of them. The var html/default-options lists every
option with its default. For example, :links? makes text show the URL
of each link:
(html/text piece {:links? true})
;; => "Hi!\n\nNice post (https://example.com/),\nthanks.\n\n- one\n- two"
To change what hiccup keeps, pass your own tables. An image loads from
its own server, which can see who views it. To keep images from loading,
take :src out of the allowed attributes of :img:
(def attributes (update html/allowed-attributes :img disj :src))
(html/hiccup "<img src='https://example.com/cat.png' alt='A cat'>"
{:allowed-attributes attributes})
;; => ([:img {:alt "A cat"}])
{:quirks? false}.hiccup and sanitize keep only the tags,
attributes and URLs that a page can render safely. Both text and
markdown leave out the control characters that a terminal would obey,
and markdown also escapes what a Markdown renderer would read as
markup..cljc. Its only dependency
is data.json, which reads the table of character references when the
code compiles.See the benchmarks for how html-pieces compares with libraries that do the same jobs, on the JVM, in Node and in the browser.
clojure -X:test # the tests on the JVM
npm install # once, for the Node tests and benchmarks
clojure -M:cljs compile test # the tests in Node
clojure -X:bench # the benchmarks against other libraries
The benchmarks take a few minutes. They print a report, which is kept in
target/bench/ with the tables of doc/benchmarks.md in
benchmarks.md. To run some of them, pass e.g. :only '#{:jvm}',
:areas '#{:parse}' or :timing :quick.
The tests run the tokenizer tests of
html5lib-tests when they're
in dev-resources/html5lib-tests/, which Git ignores. CI clones them into
that folder. To do the same yourself, at the version that the tokenizer
passes:
git clone https://github.com/html5lib/html5lib-tests dev-resources/html5lib-tests
git -C dev-resources/html5lib-tests checkout c777c408b61078ea2eb4acefc2535f54dbc8b28a
The html-pieces project is licensed under the MIT licence. The
table of character references in resources/ is WHATWG's
entities.json, under CC BY 4.0.
Can you improve this documentation?Edit on GitHub
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |