A full-text index of documents, and queries of it ranked by BM25F, on both platforms.
A document is a map of its :id, its :fields, and what a result gives back, under :stored. A field holds text, or anything whose text counts, such as a number or a set of keywords, but a vector holds terms to take as they are.
The comments name the source of each part, and where it departs from it:
A full-text index of documents, and queries of it ranked by BM25F, on both platforms. A document is a map of its :id, its :fields, and what a result gives back, under :stored. A field holds text, or anything whose text counts, such as a number or a set of keywords, but a vector holds terms to take as they are. The comments name the source of each part, and where it departs from it: - Robertson and Zaragoza, The Probabilistic Relevance Framework: BM25 and Beyond, Foundations and Trends in Information Retrieval 3(4), 2009, https://www.staff.city.ac.uk/~sbrp622/papers/foundations_bm25_review.pdf - Lucene's BM25Similarity and BooleanQuery, and the segments of its index, https://lucene.apache.org/core/9_11_1/core/
(add index)(add index doc & docs)The index with the document doc, and the docs after it, each a map
of :id, :fields and :stored. A document replaces any earlier one with
the same id.
The `index` with the document `doc`, and the `docs` after it, each a map of :id, :fields and :stored. A document replaces any earlier one with the same id.
The limits that keep the cost of a query that anyone can type within bounds, when its options don't set them:
The limits that keep the cost of a query that anyone can type within bounds, when its options don't set them: - :max-completions, the terms that a prefix stands for, those of the most documents, as in Xapian - :max-expansions, the terms that a word with typos stands for, the closest, as in Lucene's FuzzyQuery
(ids index)The ids of the documents in index.
The ids of the documents in `index`.
(index)(index docs)(index docs opts)An index of the documents docs, or an empty one, with opts as the
defaults of its queries, e.g. the :boosts of its fields. The opts
print with the index, so they must be data.
An index of the documents `docs`, or an empty one, with `opts` as the defaults of its queries, e.g. the :boosts of its fields. The `opts` print with the index, so they must be data.
(query index q)(query index q opts)The documents of index that match the query q, best first, as maps
of :id, :score and :stored, with the opts below.
The query is a string, the same query as data, or a vector of terms
taken as they are. The opts are those of search.queries/parse, and
these, over the defaults of the index:
Equal scores are ordered by id, numbers ascending and other ids by hash, so the order is stable on both platforms. When words matched others with typos, the results have :typos in their metadata, a map of each such word to those it matched, the closest first, and so does each result that holds a word only with typos:
(meta (query idx "intervew"))
;; => {:typos {"intervew" ["interview" "interviews"]}}
The documents of `index` that match the query `q`, best first, as maps
of :id, :score and :stored, with the `opts` below.
The query is a string, the same query as data, or a vector of terms
taken as they are. The `opts` are those of search.queries/parse, and
these, over the defaults of the index:
- :limit, the most results, and :offset, how many to skip first
- :fuzzy, which words match words with typos too: by default those that
no document holds, true for all, a number for all with at most that
many edits, up to 2, and false for none but those marked with ~
- :typo-lengths, the shortest words with one edit and with two, [3 6]
by default, as Elasticsearch's AUTO
- :boosts, the weight of a match in each field, 1 by default, where 0
leaves a field out unless the query names it
- :fields, the set of fields to search, all by default, which gives the
other fields the weight 0
- :k1 and :b, the saturation and length normalization of BM25F
- :filter-pred, a predicate of a result to keep it
- :rank-fn, a function of a result to order by, ascending, instead of
the score
- :max-completions and :max-expansions, the most terms that a prefix
and a word with typos stand for
Equal scores are ordered by id, numbers ascending and other ids by
hash, so the order is stable on both platforms.
When words matched others with typos, the results have :typos in their
metadata, a map of each such word to those it matched, the closest
first, and so does each result that holds a word only with typos:
(meta (query idx "intervew"))
;; => {:typos {"intervew" ["interview" "interviews"]}}(query-opts index)(query-opts index opts)The opts over the defaults of index, to read a query as the index
does: with the names of its fields among the :aliases, and a
:known-pred that tells the terms its documents hold.
The `opts` over the defaults of `index`, to read a query as the index does: with the names of its fields among the :aliases, and a :known-pred that tells the terms its documents hold.
(remove index)(remove index id & ids)The index without the document id, and the ids after it.
The `index` without the document `id`, and the `ids` after it.
(restore index)The index read back from EDN with its arrays made again, which a query
of it otherwise does each time, or nil for a nil index.
The `index` read back from EDN with its arrays made again, which a query of it otherwise does each time, or nil for a nil `index`.
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |