Liking cljdoc? Tell your friends :D

dk.simongray.drop-in-search

A full-text index of documents, and queries of it ranked by BM25F, on both platforms.

A document is a map of its :id, its :fields, and what a result gives back, under :stored. A field holds text, or anything whose text counts, such as a number or a set of keywords, but a vector holds terms to take as they are.

The comments name the source of each part, and where it departs from it:

A full-text index of documents, and queries of it ranked by BM25F, on
both platforms.

A document is a map of its :id, its :fields, and what a result gives
back, under :stored. A field holds text, or anything whose text counts,
such as a number or a set of keywords, but a vector holds terms to take
as they are.

The comments name the source of each part, and where it departs from it:

- Robertson and Zaragoza, The Probabilistic Relevance Framework: BM25
  and Beyond, Foundations and Trends in Information Retrieval 3(4),
  2009, https://www.staff.city.ac.uk/~sbrp622/papers/foundations_bm25_review.pdf
- Lucene's BM25Similarity and BooleanQuery, and the segments of its
  index, https://lucene.apache.org/core/9_11_1/core/
raw docstring

addclj/s

(add index)
(add index doc & docs)

The index with the document doc, and the docs after it, each a map of :id, :fields and :stored. A document replaces any earlier one with the same id.

The `index` with the document `doc`, and the `docs` after it, each a map
of :id, :fields and :stored. A document replaces any earlier one with
the same id.
sourceraw docstring

default-limitsclj/s

The limits that keep the cost of a query that anyone can type within bounds, when its options don't set them:

  • :max-completions, the terms that a prefix stands for, those of the most documents, as in Xapian
  • :max-expansions, the terms that a word with typos stands for, the closest, as in Lucene's FuzzyQuery
The limits that keep the cost of a query that anyone can type within
bounds, when its options don't set them:

- :max-completions, the terms that a prefix stands for, those of the
  most documents, as in Xapian
- :max-expansions, the terms that a word with typos stands for, the
  closest, as in Lucene's FuzzyQuery
sourceraw docstring

idsclj/s

(ids index)

The ids of the documents in index.

The ids of the documents in `index`.
sourceraw docstring

indexclj/s

(index)
(index docs)
(index docs opts)

An index of the documents docs, or an empty one, with opts as the defaults of its queries, e.g. the :boosts of its fields. The opts print with the index, so they must be data.

An index of the documents `docs`, or an empty one, with `opts` as the
defaults of its queries, e.g. the :boosts of its fields. The `opts`
print with the index, so they must be data.
sourceraw docstring

queryclj/s

(query index q)
(query index q opts)

The documents of index that match the query q, best first, as maps of :id, :score and :stored, with the opts below.

The query is a string, the same query as data, or a vector of terms taken as they are. The opts are those of search.queries/parse, and these, over the defaults of the index:

  • :limit, the most results, and :offset, how many to skip first
  • :fuzzy, which words match words with typos too: by default those that no document holds, true for all, a number for all with at most that many edits, up to 2, and false for none but those marked with ~
  • :typo-lengths, the shortest words with one edit and with two, [3 6] by default, as Elasticsearch's AUTO
  • :boosts, the weight of a match in each field, 1 by default, where 0 leaves a field out unless the query names it
  • :fields, the set of fields to search, all by default, which gives the other fields the weight 0
  • :k1 and :b, the saturation and length normalization of BM25F
  • :filter-pred, a predicate of a result to keep it
  • :rank-fn, a function of a result to order by, ascending, instead of the score
  • :max-completions and :max-expansions, the most terms that a prefix and a word with typos stand for

Equal scores are ordered by id, numbers ascending and other ids by hash, so the order is stable on both platforms. When words matched others with typos, the results have :typos in their metadata, a map of each such word to those it matched, the closest first, and so does each result that holds a word only with typos:

(meta (query idx "intervew"))
;; => {:typos {"intervew" ["interview" "interviews"]}}
The documents of `index` that match the query `q`, best first, as maps
of :id, :score and :stored, with the `opts` below.

The query is a string, the same query as data, or a vector of terms
taken as they are. The `opts` are those of search.queries/parse, and
these, over the defaults of the index:

- :limit, the most results, and :offset, how many to skip first
- :fuzzy, which words match words with typos too: by default those that
  no document holds, true for all, a number for all with at most that
  many edits, up to 2, and false for none but those marked with ~
- :typo-lengths, the shortest words with one edit and with two, [3 6]
  by default, as Elasticsearch's AUTO
- :boosts, the weight of a match in each field, 1 by default, where 0
  leaves a field out unless the query names it
- :fields, the set of fields to search, all by default, which gives the
  other fields the weight 0
- :k1 and :b, the saturation and length normalization of BM25F
- :filter-pred, a predicate of a result to keep it
- :rank-fn, a function of a result to order by, ascending, instead of
  the score
- :max-completions and :max-expansions, the most terms that a prefix
  and a word with typos stand for

Equal scores are ordered by id, numbers ascending and other ids by
hash, so the order is stable on both platforms.
When words matched others with typos, the results have :typos in their
metadata, a map of each such word to those it matched, the closest
first, and so does each result that holds a word only with typos:

    (meta (query idx "intervew"))
    ;; => {:typos {"intervew" ["interview" "interviews"]}}
sourceraw docstring

query-optsclj/s

(query-opts index)
(query-opts index opts)

The opts over the defaults of index, to read a query as the index does: with the names of its fields among the :aliases, and a :known-pred that tells the terms its documents hold.

The `opts` over the defaults of `index`, to read a query as the index
does: with the names of its fields among the :aliases, and a
:known-pred that tells the terms its documents hold.
sourceraw docstring

removeclj/s

(remove index)
(remove index id & ids)

The index without the document id, and the ids after it.

The `index` without the document `id`, and the `ids` after it.
sourceraw docstring

restoreclj/s

(restore index)

The index read back from EDN with its arrays made again, which a query of it otherwise does each time, or nil for a nil index.

The `index` read back from EDN with its arrays made again, which a query
of it otherwise does each time, or nil for a nil `index`.
sourceraw docstring

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
←Move to previous article
→Move to next article
Ctrl+/Jump to the search field
× close