Liking cljdoc? Tell your friends :D

dk.simongray.drop-in-search

A full-text index of documents, and queries of it ranked by BM25F, on both platforms.

A document is a map of its :id, its :fields, and what a result gives back, under :stored. A field holds text, or anything whose text counts, such as a number or a set of keywords, but a vector holds terms to take as they are. The text of an instant is its date and time in UTC, e.g. 2024-03-01 09:30:00, and that of anything else is what str gives.

Most of the code behind these functions is in these namespaces:

  • [[dk.simongray.drop-in-search.indexing]]: adding and removing documents
  • dk.simongray.drop-in-search.queries: the query language
  • [[dk.simongray.drop-in-search.terms]]: the terms that the words of a query stand for
  • [[dk.simongray.drop-in-search.matching]]: the documents that match a query
  • [[dk.simongray.drop-in-search.scoring]]: the BM25F score of a document that matches
  • [[dk.simongray.drop-in-search.hits]]: the documents that match in a segment, and their scores, as arrays
  • [[dk.simongray.drop-in-search.ranking]]: the order of the results
A full-text index of documents, and queries of it ranked by BM25F, on
both platforms.

A document is a map of its :id, its :fields, and what a result gives
back, under :stored. A field holds text, or anything whose text counts,
such as a number or a set of keywords, but a vector holds terms to take
as they are. The text of an instant is its date and time in UTC, e.g.
2024-03-01 09:30:00, and that of anything else is what str gives.

Most of the code behind these functions is in these namespaces:

- [[dk.simongray.drop-in-search.indexing]]: adding and removing
  documents
- [[dk.simongray.drop-in-search.queries]]: the query language
- [[dk.simongray.drop-in-search.terms]]: the terms that the words of a
  query stand for
- [[dk.simongray.drop-in-search.matching]]: the documents that match a
  query
- [[dk.simongray.drop-in-search.scoring]]: the BM25F score of a
  document that matches
- [[dk.simongray.drop-in-search.hits]]: the documents that match in a
  segment, and their scores, as arrays
- [[dk.simongray.drop-in-search.ranking]]: the order of the results
raw docstring

addclj/s

(add index)
(add index doc & docs)

The index with the document doc, and the docs after it, each a map of :id, :fields and :stored. A document replaces any earlier one with the same id.

The words are mapped by the :term-fn of the index. An index read back from EDN or CIFF has none until restore or ciff/read-files gives it one, and without it the words of a new document stay as they are, unlike those already in the index.

The `index` with the document `doc`, and the `docs` after it, each a map
of :id, :fields and :stored. A document replaces any earlier one with
the same id.

The words are mapped by the :term-fn of the index. An index read back
from EDN or CIFF has none until restore or ciff/read-files gives it
one, and without it the words of a new document stay as they are,
unlike those already in the index.
sourceraw docstring

default-limitsclj/s

The limits that keep the cost of a query that anyone can type within bounds, when its options don't set them:

  • :max-completions, the terms that a prefix stands for, those of the most documents, as in Xapian
  • :max-expansions, the terms that a word with typos stands for, the closest, as in Lucene's FuzzyQuery
The limits that keep the cost of a query that anyone can type within
bounds, when its options don't set them:

- :max-completions, the terms that a prefix stands for, those of the
  most documents, as in Xapian
- :max-expansions, the terms that a word with typos stands for, the
  closest, as in Lucene's FuzzyQuery
sourceraw docstring

idsclj/s

(ids index)

The ids of the documents in index.

The ids of the documents in `index`.
sourceraw docstring

indexclj/s

(index)
(index docs)
(index docs {:keys [term-fn] :as opts})

An index of the documents docs, or an empty one, with opts as the defaults of its queries, e.g. the :boosts of its fields, and its :term-fn.

The :term-fn is a function of each term of a text that gives the term to index and to look for in its place, e.g. a stemmer's, so that posts finds post. The other opts print with the index, so they must be data, but a function can't, so give the :term-fn to restore and ciff/read-files again.

An index of the documents `docs`, or an empty one, with `opts` as the
defaults of its queries, e.g. the :boosts of its fields, and its
:term-fn.

The :term-fn is a function of each term of a text that gives the term
to index and to look for in its place, e.g. a stemmer's, so that posts
finds post. The other `opts` print with the index, so they must be data,
but a function can't, so give the :term-fn to restore and
ciff/read-files again.
sourceraw docstring

queryclj/s

(query index q)
(query index q opts)

The documents of index that match the query q, best first, as maps of :id, :score and :stored, with the opts below.

The query is a string, the same query as data, or a vector of terms taken as they are. The opts are those of search.queries/parse, and these, over the defaults of the index:

  • :limit, the most results, and :offset, how many to skip first
  • :fuzzy, which words match words with typos too: by default those that no document holds, true for all, a number for all with at most that many edits, up to 2, and false for none but those marked with ~
  • :typo-lengths, the shortest words with one edit and with two, [3 6] by default, as Elasticsearch's AUTO
  • :boosts, the weight of a match in each field, 1 by default, where 0 leaves a field out unless the query names it
  • :fields, the set of fields to search, all by default, which gives the other fields the weight 0
  • :k1 and :b, the saturation and length normalization of BM25F
  • :filter-pred, a predicate of a result to keep it
  • :rank-fn, a function of a result to order by, ascending, instead of the score
  • :max-completions and :max-expansions, the most terms that a prefix and a word with typos stand for

Equal scores are ordered by id, numbers first and ascending, and other ids by hash, so the order is stable on both platforms.

The results have the :total of the documents that match in their metadata, e.g. for "1 to 10 of 243", unless a :filter-pred had no need to read them all. When words matched others with typos, the metadata also has :typos, a map of each such word to those it matched, the closest first, and so does each result that holds a word only with typos:

(meta (query idx "intervew"))
;; => {:total 2 :typos {"intervew" ["interview" "interviews"]}}
The documents of `index` that match the query `q`, best first, as maps
of :id, :score and :stored, with the `opts` below.

The query is a string, the same query as data, or a vector of terms
taken as they are. The `opts` are those of search.queries/parse, and
these, over the defaults of the index:

- :limit, the most results, and :offset, how many to skip first
- :fuzzy, which words match words with typos too: by default those that
  no document holds, true for all, a number for all with at most that
  many edits, up to 2, and false for none but those marked with ~
- :typo-lengths, the shortest words with one edit and with two, [3 6]
  by default, as Elasticsearch's AUTO
- :boosts, the weight of a match in each field, 1 by default, where 0
  leaves a field out unless the query names it
- :fields, the set of fields to search, all by default, which gives the
  other fields the weight 0
- :k1 and :b, the saturation and length normalization of BM25F
- :filter-pred, a predicate of a result to keep it
- :rank-fn, a function of a result to order by, ascending, instead of
  the score
- :max-completions and :max-expansions, the most terms that a prefix
  and a word with typos stand for

Equal scores are ordered by id, numbers first and ascending, and other
ids by hash, so the order is stable on both platforms.

The results have the :total of the documents that match in their
metadata, e.g. for "1 to 10 of 243", unless a :filter-pred had no need
to read them all. When words matched others with typos, the metadata
also has :typos, a map of each such word to those it matched, the
closest first, and so does each result that holds a word only with
typos:

    (meta (query idx "intervew"))
    ;; => {:total 2 :typos {"intervew" ["interview" "interviews"]}}
sourceraw docstring

query-optsclj/s

(query-opts index)
(query-opts index opts)

The opts over the defaults of index, to read a query as the index does: with the names of its fields among the :aliases, its :term-fn, and a :known-pred that tells the terms its documents hold.

The `opts` over the defaults of `index`, to read a query as the index
does: with the names of its fields among the :aliases, its :term-fn,
and a :known-pred that tells the terms its documents hold.
sourceraw docstring

removeclj/s

(remove index)
(remove index id & ids)

The index without the document id, and the ids after it.

The `index` without the document `id`, and the `ids` after it.
sourceraw docstring

restoreclj/s

(restore index)
(restore index {:keys [term-fn] :as opts})

The index read back from EDN with its arrays made again, which a query of it otherwise does each time, and the :term-fn of opts that it was built with, or nil for a nil index.

The `index` read back from EDN with its arrays made again, which a query
of it otherwise does each time, and the :term-fn of `opts` that it was
built with, or nil for a nil `index`.
sourceraw docstring

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
←Move to previous article
→Move to next article
Ctrl+/Jump to the search field
× close