Liking cljdoc? Tell your friends :D

dk.simongray.drop-in-search

A full-text index of documents, and queries of it ranked by BM25F, on both platforms.

A document is a map of its :id, its :fields, and what a result gives back, under :stored. A field holds text, or anything whose text counts, such as a number or a set of keywords, but a vector holds terms to take as they are.

The comments name the source of each part, and where it departs from it:

A full-text index of documents, and queries of it ranked by BM25F, on
both platforms.

A document is a map of its :id, its :fields, and what a result gives
back, under :stored. A field holds text, or anything whose text counts,
such as a number or a set of keywords, but a vector holds terms to take
as they are.

The comments name the source of each part, and where it departs from it:

- Robertson and Zaragoza, The Probabilistic Relevance Framework: BM25
  and Beyond, Foundations and Trends in Information Retrieval 3(4),
  2009, https://www.staff.city.ac.uk/~sbrp622/papers/foundations_bm25_review.pdf
- Lucene's BM25Similarity and BooleanQuery, and the segments of its
  index, https://lucene.apache.org/core/9_11_1/core/
raw docstring

dk.simongray.drop-in-search.analysis

Text analysis for search on both platforms: folding, words, and the terms of a text by position.

A word is a run of letters, marks and digits, and anything else splits words, e.g. "Simon's" is the words simon and s. A run in Korean, or in a script without spaces between words, such as Chinese, Japanese or Thai, has its pairs of characters, bigrams, as terms.

The comments name the source of each part, and where it departs from it:

Text analysis for search on both platforms: folding, words, and the
terms of a text by position.

A word is a run of letters, marks and digits, and anything else splits
words, e.g. "Simon's" is the words simon and s. A run in Korean, or in
a script without spaces between words, such as Chinese, Japanese or
Thai, has its pairs of characters, bigrams, as terms.

The comments name the source of each part, and where it departs from it:

- Unicode Standard Annex #15, Unicode Normalization Forms,
  https://www.unicode.org/reports/tr15/
- Unicode Standard Annex #29, Unicode Text Segmentation, section 4,
  https://www.unicode.org/reports/tr29/
- CaseFolding.txt of the Unicode Character Database,
  https://www.unicode.org/Public/UCD/latest/ucd/CaseFolding.txt
- Lucene's ASCIIFoldingFilter and CJKBigramFilter,
  https://lucene.apache.org/core/9_11_0/analysis/common/
- Groonga's TokenBigram,
  https://groonga.org/docs/reference/tokenizers/token_bigram.html
- Xapian's TermGenerator with FLAG_NGRAMS,
  https://xapian.org/docs/apidoc/html/classXapian_1_1TermGenerator.html
raw docstring

dk.simongray.drop-in-search.ciff

An index kept as files in a store: a manifest in EDN, and a file for each segment in the Common Index File Format, CIFF, which other search engines can read.

CIFF is that of https://github.com/osirrc/ciff, its README and its CommonIndexFileFormat.proto of 2020-03. What CIFF lacks has fields of its own, which a reader of CIFF skips, as the encoding of protobuf has it, https://protobuf.dev/programming-guides/encoding/

An index kept as files in a store: a manifest in EDN, and a file for
each segment in the Common Index File Format, CIFF, which other search
engines can read.

CIFF is that of https://github.com/osirrc/ciff, its README and its
CommonIndexFileFormat.proto of 2020-03. What CIFF lacks has fields of
its own, which a reader of CIFF skips, as the encoding of protobuf has
it, https://protobuf.dev/programming-guides/encoding/
raw docstring

dk.simongray.drop-in-search.queries

The query language of full-text search, and a query matched against one text without an index, on both platforms.

The comments name the source of each part, and where it departs from it:

The query language of full-text search, and a query matched against one
text without an index, on both platforms.

The comments name the source of each part, and where it departs from it:

- Lucene's simple query parser and FuzzyQuery,
  https://lucene.apache.org/core/9_11_1/core/
- Elasticsearch's simple_query_string,
  https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-simple-query-string-query
- SQLite FTS5's query syntax and its snippet function,
  https://www.sqlite.org/fts5.html
raw docstring

dk.simongray.drop-in-search.segment

The segments of an index, on both platforms: the postings of some of its documents in sorted arrays of ints, as a segment of Lucene's index holds them, https://lucene.apache.org/core/9_11_1/core/

The segments of an index, on both platforms: the postings of some of
its documents in sorted arrays of ints, as a segment of Lucene's index
holds them, https://lucene.apache.org/core/9_11_1/core/
raw docstring

No vars found in this namespace.

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
←Move to previous article
→Move to next article
Ctrl+/Jump to the search field
× close