All notable changes to this project are documented here. Format follows Keep a Changelog; this project adheres to Semantic Versioning.
:keep-separator option with :start, :end, and false modes.chunk.core/split-with-offsets for source character offsets.:length-fn calls for already-measured joined candidates are cached per split.:length-fn so token-budget mode no longer exceeds the budget on non-additive (BPE) tokenizers.chunk.core/language-separators - separator presets for :markdown,
:python, :clojure, :javascript, :typescript, :java, :go, :rust,
:html, and :latex, so chunks land on structural boundaries (headings,
function definitions, tags) before the paragraph/line/word/character tail.chunk.core/separators-for - preset lookup; unknown language throws with the
known set in ex-data.:language option on split - sugar for :separators (separators-for lang);
passing both :language and :separators throws.chunk.core/split - recursive text splitter for RAG / LLM pipelines. Splits text
into overlapping chunks no larger than a target size, trying ordered separators
(paragraph -> line -> word -> character) so chunks land on natural boundaries.:chunk-size, :overlap, :separators, and a pluggable :length-fn
(default count = characters) so you can chunk by tokens instead - e.g. with
tokenizers-clj's count-tokens.Can you improve this documentation?Edit on GitHub
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |