PDF extraction and inspection for Clojure, built on Apache PDFBox.
The Clojure counterpart to Python's pdfplumber:
pull text, tables, and geometry out of digitally generated PDFs as plain,
EDN/JSON-friendly data.
Try it live - a hosted table extractor over this library: upload a PDF to see its detected table regions drawn on each page, tune the detection strategy, and download the rows as CSV or JSON. Files are processed in memory and never stored.
Stable (1.0.0). Data shapes are settled and follow semantic versioning:
breaking changes to the returned maps require a major bump.
Covers the full Python pdfplumber extraction surface (text, words, chars, objects, tables, crop), plus:
Parity is measured, not asserted. The parity workflow runs weekly: it fetches
the upstream jsvine/pdfplumber test corpus, extracts every PDF with Python
pdfplumber to build a baseline, and compares this library against it on page
count, text similarity, and word count. The corpus is not committed, so run it
locally with dev/fetch-corpus.sh then dev/gen_golden.py.
deps.edn
net.clojars.savya/pdfplumber-clj {:mvn/version "1.0.0"}
Leiningen
[net.clojars.savya/pdfplumber-clj "1.0.0"]
Requires JDK 17+.
(require '[pdfplumber.core :as pdf])
(pdf/with-pdf [doc "statement.pdf"]
(pdf/text doc {:page 1})) ; => "Account statement\n..."
;; Supply :password for password-protected PDFs.
(pdf/with-pdf [doc "statement.pdf" {:password "hunter2"}]
(pdf/text doc {:page 1}))
(pdf/with-pdf [doc "statement.pdf"]
(pdf/words doc {:page 1})) ; => [{:text "Account" :x0 .. :top .. :x1 .. :bottom ..} ...]
(pdf/with-pdf [doc "invoice.pdf"]
(pdf/extract-table doc {:page 1 :strategy :lines}))
The reducible-* functions on pdfplumber.core return IReduceInit streams
that extract one page at a time. Transducers terminate early without extracting
later pages. They are re-exported from pdfplumber.reducible.
(into []
(comp (filter #(> (:size %) 10)) (take 100))
(pdf/reducible-chars doc))
(transduce (take 20) conj [] (pdf/reducible-words doc))
extract-tables returns every independent table region on a page, ordered
top-to-bottom then left-to-right. extract-table returns only the first region.
Configure each axis independently with :vertical-strategy and
:horizontal-strategy; each accepts :lines, :lines-strict, :text, or
:explicit. The legacy :strategy option sets both axes.
(pdf/extract-tables doc
{:page 1
:vertical-strategy :explicit
:horizontal-strategy :lines
:explicit-vertical-lines [70 170 260]
:snap-tolerance 3.0
:join-tolerance 3.0
:edge-min-length 3.0
:intersection-tolerance 3.0
:min-words-vertical 3
:min-words-horizontal 1})
Use :explicit-horizontal-lines with :horizontal-strategy :explicit.
Explicit lines may be coordinates or maps with bounded line coordinates.
table->maps converts an extracted table or raw rows to a sequence of
header-keyed maps.
(-> (pdf/extract-table doc {:page 1})
pdf/table->maps)
The result feeds tech.v3.dataset/->dataset directly with no added dependency.
Use :keywordize? true for keyword keys. Set :header to :first (the
default), an explicit key vector, or false for integer keys.
pdfplumber.core/to-image renders a page through PDFBox and returns a
PageImage; pdfplumber.core/page-image? identifies one. Overlay, reset/copy,
save, and display verbs live in pdfplumber.image.
(require '[pdfplumber.image :as image])
(pdf/with-pdf [doc "invoice.pdf"]
(-> (pdf/to-image doc {:page 1 :resolution 144})
(image/outline-words)
(image/draw-rect [72 72 240 160])
(image/save "debug.png")))
(pdf/with-pdf [doc "invoice.pdf"]
(-> (pdf/to-image doc {:page 1})
(image/debug-tablefinder {:vertical-strategy :lines})
(image/save "tables.png")))
More verbs in pdfplumber.image:
draw-line, draw-vline, draw-hline, draw-rect, draw-rects, draw-circle, draw-circlesoutline-words, outline-charsreset, copy, save, showstructure-tree returns a tagged PDF's nested logical structure.
page-structure-tree restricts it to a 1-based page. Untagged PDFs return [].
(pdf/structure-tree doc)
(pdf/page-structure-tree doc 1)
form-fields returns terminal AcroForm field maps with values, constraints,
options, and first-widget geometry. field-values returns the name-to-value
map.
(pdf/form-fields doc)
;; => [{:name "customer.email", :type :text, :value "ada@example.com",
;; :required? true, :read-only? false, :page-number 1,
;; :bbox [72.0 120.0 288.0 140.0]}]
(pdf/field-values doc)
;; => {"customer.email" "ada@example.com"}
Widget annotations from annots also carry :field-name, :field-value, and
:field-type.
outline returns nested bookmarks with resolved 1-based page numbers.
(pdf/outline doc)
;; => [{:title "Introduction", :page-number 1, :children []}]
attachments returns embedded-file metadata. Set :include-data? true to add
decoded :bytes.
(pdf/attachments doc)
;; => [{:name "data.csv", :size 128, :mime-type "text/csv"}]
(pdf/attachments doc {:include-data? true})
permissions reports encryption state and effective access flags.
(pdf/permissions doc)
;; => {:encrypted? true, :can-print? true, :can-modify? false, ...}
signatures surfaces signature metadata plus a :covers-whole-document?
integrity signal. signed? reports signature-dictionary presence. These APIs do
not validate cryptographic signatures, certificates, or trust.
(pdf/signatures doc)
;; => [{:name "Ada Lovelace", :byte-range [0 1024 2048 512],
;; :covers-whole-document? true}]
(pdf/signed? doc)
;; => true
Dump selected PDF objects as CSV or JSON:
clojure -M -m pdfplumber.cli statement.pdf \
--format json --pages 1,2 --types char,line,rect,curve,image,annot \
--precision 2 --indent 2
images returns drawn image objects; they also appear as :image entries in
objects. Each carries:
:bbox, plus pixel :width and :height:colorspace and :bits:mask? and :smask?Decoded PNG bytes are omitted by default.
(pdf/images doc {:page 1})
(pdf/images doc {:page 1 :include-image-data? true}) ; adds :bytes
Public coordinates use a top-left origin (matching pdfplumber), with bounding
boxes as [x0 top x1 bottom] in PDF user-space points. PDFBox's native bottom-left
coordinates are converted internally.
In:
Not in scope (same as Python pdfplumber): PDF generation, OCR, scanned/image PDFs, and layout ML.
Two caveats:
:text strategy is heuristic, for digitally generated PDFs.Copyright © 2026 Savyasachi.
Distributed under the Eclipse Public License 2.0.
Can you improve this documentation? These fine people already did:
Savyasachi & Savyasachi JagadeeshanEdit on GitHub
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |