Liking cljdoc? Tell your friends :D

dk.simongray.drop-in-search

A full-text index of documents, and queries of it ranked by BM25F, on both platforms.

A document is a map of its :id, its :fields, and what a result gives back, under :stored. A field holds text, or anything whose text counts, such as a number or a set of keywords, but a vector holds terms to take as they are. The text of an instant is its date and time in UTC, e.g. 2024-03-01 09:30:00, and that of anything else is what str gives.

Most of the code behind these functions is in these namespaces:

  • [[dk.simongray.drop-in-search.indexing]]: adding and removing documents
  • dk.simongray.drop-in-search.queries: the query language
  • [[dk.simongray.drop-in-search.terms]]: the terms that the words of a query stand for
  • [[dk.simongray.drop-in-search.matching]]: the documents that match a query
  • [[dk.simongray.drop-in-search.scoring]]: the BM25F score of a document that matches
  • [[dk.simongray.drop-in-search.hits]]: the documents that match in a segment, and their scores, as arrays
  • [[dk.simongray.drop-in-search.ranking]]: the order of the results
A full-text index of documents, and queries of it ranked by BM25F, on
both platforms.

A document is a map of its :id, its :fields, and what a result gives
back, under :stored. A field holds text, or anything whose text counts,
such as a number or a set of keywords, but a vector holds terms to take
as they are. The text of an instant is its date and time in UTC, e.g.
2024-03-01 09:30:00, and that of anything else is what str gives.

Most of the code behind these functions is in these namespaces:

- [[dk.simongray.drop-in-search.indexing]]: adding and removing
  documents
- [[dk.simongray.drop-in-search.queries]]: the query language
- [[dk.simongray.drop-in-search.terms]]: the terms that the words of a
  query stand for
- [[dk.simongray.drop-in-search.matching]]: the documents that match a
  query
- [[dk.simongray.drop-in-search.scoring]]: the BM25F score of a
  document that matches
- [[dk.simongray.drop-in-search.hits]]: the documents that match in a
  segment, and their scores, as arrays
- [[dk.simongray.drop-in-search.ranking]]: the order of the results
raw docstring

dk.simongray.drop-in-search.analysis

Text analysis for search on both platforms: folding, words, and the terms of a text by position.

A word is a run of letters, marks and digits, and anything else splits words, e.g. "Simon's" is the words simon and s. A run in Korean, or in a script without spaces between words, such as Chinese, Japanese or Thai, has its pairs of characters, bigrams, as terms.

The comments name the source of each part, and where it departs from it:

Text analysis for search on both platforms: folding, words, and the
terms of a text by position.

A word is a run of letters, marks and digits, and anything else splits
words, e.g. "Simon's" is the words simon and s. A run in Korean, or in
a script without spaces between words, such as Chinese, Japanese or
Thai, has its pairs of characters, bigrams, as terms.

The comments name the source of each part, and where it departs from it:

- Unicode Standard Annex #15, Unicode Normalization Forms,
  https://www.unicode.org/reports/tr15/
- Unicode Standard Annex #29, Unicode Text Segmentation, section 4,
  https://www.unicode.org/reports/tr29/
- CaseFolding.txt of the Unicode Character Database,
  https://www.unicode.org/Public/UCD/latest/ucd/CaseFolding.txt
- Lucene's ASCIIFoldingFilter and CJKBigramFilter,
  https://lucene.apache.org/core/9_11_0/analysis/common/
- Groonga's TokenBigram,
  https://groonga.org/docs/reference/tokenizers/token_bigram.html
- Xapian's TermGenerator with FLAG_NGRAMS,
  https://xapian.org/docs/apidoc/html/classXapian_1_1TermGenerator.html
raw docstring

dk.simongray.drop-in-search.ciff

An index kept as files in a store: a manifest in EDN, and a file for each segment in the Common Index File Format, CIFF, which other search engines can read.

CIFF is that of https://github.com/osirrc/ciff, its README and its CommonIndexFileFormat.proto of 2020-03. What CIFF lacks has fields of its own, which a reader of CIFF skips, as the encoding of protobuf has it, https://protobuf.dev/programming-guides/encoding/

An index kept as files in a store: a manifest in EDN, and a file for
each segment in the Common Index File Format, CIFF, which other search
engines can read.

CIFF is that of https://github.com/osirrc/ciff, its README and its
CommonIndexFileFormat.proto of 2020-03. What CIFF lacks has fields of
its own, which a reader of CIFF skips, as the encoding of protobuf has
it, https://protobuf.dev/programming-guides/encoding/
raw docstring

dk.simongray.drop-in-search.hits

The hits of a query in a segment, and how the hits of its parts combine.

Hits are a map of :docs, the numbers of the documents that match, in order, and :scores, the score of each, both arrays.

The hits of a query in a segment, and how the hits of its parts
combine.

Hits are a map of :docs, the numbers of the documents that match, in
order, and :scores, the score of each, both arrays.
raw docstring

No vars found in this namespace.

dk.simongray.drop-in-search.indexing

The segments of an index, and how adding and removing documents changes them, as in Lucene's index, https://lucene.apache.org/core/9_11_1/core/

An index is a map of:

  • :segments, a vector of its segments
  • :docs, a map of each id to the index of its segment and its number there
  • :fields, a map of each field to its number
  • :defaults, the options of its queries, if any

Its :term-fn, if any, is in its metadata, since a function can't be printed with it.

The segments of an index, and how adding and removing documents changes
them, as in Lucene's index, https://lucene.apache.org/core/9_11_1/core/

An index is a map of:

- :segments, a vector of its segments
- :docs, a map of each id to the index of its segment and its number
  there
- :fields, a map of each field to its number
- :defaults, the options of its queries, if any

Its :term-fn, if any, is in its metadata, since a function can't be
printed with it.
raw docstring

No vars found in this namespace.

dk.simongray.drop-in-search.matching

The documents of each segment of an index that match a prepared query, and their scores, its clauses combined as in Lucene's BooleanQuery, https://lucene.apache.org/core/9_11_1/core/

The documents of each segment of an index that match a prepared query,
and their scores, its clauses combined as in Lucene's BooleanQuery,
https://lucene.apache.org/core/9_11_1/core/
raw docstring

No vars found in this namespace.

dk.simongray.drop-in-search.queries

The query language of full-text search, and a query matched against one text without an index, on both platforms.

The comments name the source of each part, and where it departs from it:

The query language of full-text search, and a query matched against one
text without an index, on both platforms.

The comments name the source of each part, and where it departs from it:

- Lucene's simple query parser and FuzzyQuery,
  https://lucene.apache.org/core/9_11_1/core/
- Elasticsearch's simple_query_string,
  https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-simple-query-string-query
- SQLite FTS5's query syntax and its snippet function,
  https://www.sqlite.org/fts5.html
raw docstring

dk.simongray.drop-in-search.ranking

The best results of a query: the highest scores first, and equal scores by id, unless a function of a result ranks them.

The best results of a query: the highest scores first, and equal scores
by id, unless a function of a result ranks them.
raw docstring

No vars found in this namespace.

dk.simongray.drop-in-search.scoring

BM25F's scores of the documents of a segment that hold a term, from the weights and lengths of their fields.

The comments name the source of each part, and where it departs from it:

BM25F's scores of the documents of a segment that hold a term, from the
weights and lengths of their fields.

The comments name the source of each part, and where it departs from it:

- Robertson and Zaragoza, The Probabilistic Relevance Framework: BM25
  and Beyond, Foundations and Trends in Information Retrieval 3(4),
  2009, https://www.staff.city.ac.uk/~sbrp622/papers/foundations_bm25_review.pdf
- Lucene's BM25Similarity, https://lucene.apache.org/core/9_11_1/core/
raw docstring

No vars found in this namespace.

dk.simongray.drop-in-search.segment

The segments of an index, on both platforms: the postings of some of its documents in sorted arrays of ints, as a segment of Lucene's index holds them, https://lucene.apache.org/core/9_11_1/core/

The segments of an index, on both platforms: the postings of some of
its documents in sorted arrays of ints, as a segment of Lucene's index
holds them, https://lucene.apache.org/core/9_11_1/core/
raw docstring

No vars found in this namespace.

dk.simongray.drop-in-search.terms

The terms of an index that the words of a query stand for, and what scoring them takes.

A prefix stands for the terms that it starts, and a word with typos for the terms within its edits. A query is prepared with them once for all the segments, and the words that a result's document holds only with typos go under its :typos.

The terms of an index that the words of a query stand for, and what
scoring them takes.

A prefix stands for the terms that it starts, and a word with typos for
the terms within its edits. A query is prepared with them once for all
the segments, and the words that a result's document holds only with
typos go under its :typos.
raw docstring

No vars found in this namespace.

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
←Move to previous article
→Move to next article
Ctrl+/Jump to the search field
× close