Liking cljdoc? Tell your friends :D

dk.simongray.drop-in-search.analysis

Text analysis for search on both platforms: folding, words, and the terms of a text by position.

A word is a run of letters, marks and digits, and anything else splits words, e.g. "Simon's" is the words simon and s. A run in Korean, or in a script without spaces between words, such as Chinese, Japanese or Thai, has its pairs of characters, bigrams, as terms.

The comments name the source of each part, and where it departs from it:

Text analysis for search on both platforms: folding, words, and the
terms of a text by position.

A word is a run of letters, marks and digits, and anything else splits
words, e.g. "Simon's" is the words simon and s. A run in Korean, or in
a script without spaces between words, such as Chinese, Japanese or
Thai, has its pairs of characters, bigrams, as terms.

The comments name the source of each part, and where it departs from it:

- Unicode Standard Annex #15, Unicode Normalization Forms,
  https://www.unicode.org/reports/tr15/
- Unicode Standard Annex #29, Unicode Text Segmentation, section 4,
  https://www.unicode.org/reports/tr29/
- CaseFolding.txt of the Unicode Character Database,
  https://www.unicode.org/Public/UCD/latest/ucd/CaseFolding.txt
- Lucene's ASCIIFoldingFilter and CJKBigramFilter,
  https://lucene.apache.org/core/9_11_0/analysis/common/
- Groonga's TokenBigram,
  https://groonga.org/docs/reference/tokenizers/token_bigram.html
- Xapian's TermGenerator with FLAG_NGRAMS,
  https://xapian.org/docs/apidoc/html/classXapian_1_1TermGenerator.html
raw docstring

foldclj/s

(fold s)

The string s lower-cased, with its diacritics stripped and the forms of each letter made one, e.g. "Café Ø Pod" becomes "cafe o pod".

The string `s` lower-cased, with its diacritics stripped and the forms
of each letter made one, e.g. "Café Ø Pod" becomes "cafe o pod".
sourceraw docstring

spansclj/s

(spans s)

The terms of s with where each is in s, for highlighting, as maps of :term, :start, :end and :position. The terms at one position have the same :position, e.g. the last pair of characters in a run of Chinese and its last character.

The terms of `s` with where each is in `s`, for highlighting, as maps
of :term, :start, :end and :position. The terms at one position have
the same :position, e.g. the last pair of characters in a run of
Chinese and its last character.
sourceraw docstring

tokensclj/s

(tokens s)

The terms of s by position, for an index or a query.

Each is a folded word, or a vector of the terms at one position. A run in a script without spaces between words has a pair of characters at each position, and its last character shares the position of the last pair, e.g.

(tokens "Café 中秋节")
;; => ["cafe" "中秋" ["秋节" "节"]]
The terms of `s` by position, for an index or a query.

Each is a folded word, or a vector of the terms at one position. A run
in a script without spaces between words has a pair of characters at
each position, and its last character shares the position of the last
pair, e.g.

    (tokens "Café 中秋节")
    ;; => ["cafe" "中秋" ["秋节" "节"]]
sourceraw docstring

unspaced?clj/s

(unspaced? s)

Whether s starts with a letter of Korean or of a script without spaces between words, such as Chinese, Japanese or Thai.

Whether `s` starts with a letter of Korean or of a script without
spaces between words, such as Chinese, Japanese or Thai.
sourceraw docstring

wordsclj/s

(words s)

The words of s as they are written in it.

A word in a script without spaces between words is the whole run, and a change to or from such a script splits a word, e.g. "AIメモリ" is two.

The words of `s` as they are written in it.

A word in a script without spaces between words is the whole run, and a
change to or from such a script splits a word, e.g. "AIメモリ" is two.
sourceraw docstring

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
←Move to previous article
→Move to next article
Ctrl+/Jump to the search field
× close