A full-text index of documents, and queries of it ranked by BM25F, on both platforms.
A document is a map of its :id, its :fields, and what a result gives back, under :stored. A field holds text, or anything whose text counts, such as a number or a set of keywords, but a vector holds terms to take as they are.
The comments name the source of each part, and where it departs from it:
A full-text index of documents, and queries of it ranked by BM25F, on both platforms. A document is a map of its :id, its :fields, and what a result gives back, under :stored. A field holds text, or anything whose text counts, such as a number or a set of keywords, but a vector holds terms to take as they are. The comments name the source of each part, and where it departs from it: - Robertson and Zaragoza, The Probabilistic Relevance Framework: BM25 and Beyond, Foundations and Trends in Information Retrieval 3(4), 2009, https://www.staff.city.ac.uk/~sbrp622/papers/foundations_bm25_review.pdf - Lucene's BM25Similarity and BooleanQuery, and the segments of its index, https://lucene.apache.org/core/9_11_1/core/
Text analysis for search on both platforms: folding, words, and the terms of a text by position.
A word is a run of letters, marks and digits, and anything else splits words, e.g. "Simon's" is the words simon and s. A run in Korean, or in a script without spaces between words, such as Chinese, Japanese or Thai, has its pairs of characters, bigrams, as terms.
The comments name the source of each part, and where it departs from it:
Text analysis for search on both platforms: folding, words, and the terms of a text by position. A word is a run of letters, marks and digits, and anything else splits words, e.g. "Simon's" is the words simon and s. A run in Korean, or in a script without spaces between words, such as Chinese, Japanese or Thai, has its pairs of characters, bigrams, as terms. The comments name the source of each part, and where it departs from it: - Unicode Standard Annex #15, Unicode Normalization Forms, https://www.unicode.org/reports/tr15/ - Unicode Standard Annex #29, Unicode Text Segmentation, section 4, https://www.unicode.org/reports/tr29/ - CaseFolding.txt of the Unicode Character Database, https://www.unicode.org/Public/UCD/latest/ucd/CaseFolding.txt - Lucene's ASCIIFoldingFilter and CJKBigramFilter, https://lucene.apache.org/core/9_11_0/analysis/common/ - Groonga's TokenBigram, https://groonga.org/docs/reference/tokenizers/token_bigram.html - Xapian's TermGenerator with FLAG_NGRAMS, https://xapian.org/docs/apidoc/html/classXapian_1_1TermGenerator.html
An index kept as files in a store: a manifest in EDN, and a file for each segment in the Common Index File Format, CIFF, which other search engines can read.
CIFF is that of https://github.com/osirrc/ciff, its README and its CommonIndexFileFormat.proto of 2020-03. What CIFF lacks has fields of its own, which a reader of CIFF skips, as the encoding of protobuf has it, https://protobuf.dev/programming-guides/encoding/
An index kept as files in a store: a manifest in EDN, and a file for each segment in the Common Index File Format, CIFF, which other search engines can read. CIFF is that of https://github.com/osirrc/ciff, its README and its CommonIndexFileFormat.proto of 2020-03. What CIFF lacks has fields of its own, which a reader of CIFF skips, as the encoding of protobuf has it, https://protobuf.dev/programming-guides/encoding/
The query language of full-text search, and a query matched against one text without an index, on both platforms.
The comments name the source of each part, and where it departs from it:
The query language of full-text search, and a query matched against one text without an index, on both platforms. The comments name the source of each part, and where it departs from it: - Lucene's simple query parser and FuzzyQuery, https://lucene.apache.org/core/9_11_1/core/ - Elasticsearch's simple_query_string, https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-simple-query-string-query - SQLite FTS5's query syntax and its snippet function, https://www.sqlite.org/fts5.html
The segments of an index, on both platforms: the postings of some of its documents in sorted arrays of ints, as a segment of Lucene's index holds them, https://lucene.apache.org/core/9_11_1/core/
The segments of an index, on both platforms: the postings of some of its documents in sorted arrays of ints, as a segment of Lucene's index holds them, https://lucene.apache.org/core/9_11_1/core/
No vars found in this namespace.
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |