A full-text index of documents, and queries of it ranked by BM25F, on both platforms.
A document is a map of its :id, its :fields, and what a result gives back, under :stored. A field holds text, or anything whose text counts, such as a number or a set of keywords, but a vector holds terms to take as they are. The text of an instant is its date and time in UTC, e.g. 2024-03-01 09:30:00, and that of anything else is what str gives.
Most of the code behind these functions is in these namespaces:
dk.simongray.drop-in-search.queries: the query languageA full-text index of documents, and queries of it ranked by BM25F, on both platforms. A document is a map of its :id, its :fields, and what a result gives back, under :stored. A field holds text, or anything whose text counts, such as a number or a set of keywords, but a vector holds terms to take as they are. The text of an instant is its date and time in UTC, e.g. 2024-03-01 09:30:00, and that of anything else is what str gives. Most of the code behind these functions is in these namespaces: - [[dk.simongray.drop-in-search.indexing]]: adding and removing documents - [[dk.simongray.drop-in-search.queries]]: the query language - [[dk.simongray.drop-in-search.terms]]: the terms that the words of a query stand for - [[dk.simongray.drop-in-search.matching]]: the documents that match a query - [[dk.simongray.drop-in-search.scoring]]: the BM25F score of a document that matches - [[dk.simongray.drop-in-search.hits]]: the documents that match in a segment, and their scores, as arrays - [[dk.simongray.drop-in-search.ranking]]: the order of the results
Text analysis for search on both platforms: folding, words, and the terms of a text by position.
A word is a run of letters, marks and digits, and anything else splits words, e.g. "Simon's" is the words simon and s. A run in Korean, or in a script without spaces between words, such as Chinese, Japanese or Thai, has its pairs of characters, bigrams, as terms.
The comments name the source of each part, and where it departs from it:
Text analysis for search on both platforms: folding, words, and the terms of a text by position. A word is a run of letters, marks and digits, and anything else splits words, e.g. "Simon's" is the words simon and s. A run in Korean, or in a script without spaces between words, such as Chinese, Japanese or Thai, has its pairs of characters, bigrams, as terms. The comments name the source of each part, and where it departs from it: - Unicode Standard Annex #15, Unicode Normalization Forms, https://www.unicode.org/reports/tr15/ - Unicode Standard Annex #29, Unicode Text Segmentation, section 4, https://www.unicode.org/reports/tr29/ - CaseFolding.txt of the Unicode Character Database, https://www.unicode.org/Public/UCD/latest/ucd/CaseFolding.txt - Lucene's ASCIIFoldingFilter and CJKBigramFilter, https://lucene.apache.org/core/9_11_0/analysis/common/ - Groonga's TokenBigram, https://groonga.org/docs/reference/tokenizers/token_bigram.html - Xapian's TermGenerator with FLAG_NGRAMS, https://xapian.org/docs/apidoc/html/classXapian_1_1TermGenerator.html
An index kept as files in a store: a manifest in EDN, and a file for each segment in the Common Index File Format, CIFF, which other search engines can read.
CIFF is that of https://github.com/osirrc/ciff, its README and its CommonIndexFileFormat.proto of 2020-03. What CIFF lacks has fields of its own, which a reader of CIFF skips, as the encoding of protobuf has it, https://protobuf.dev/programming-guides/encoding/
An index kept as files in a store: a manifest in EDN, and a file for each segment in the Common Index File Format, CIFF, which other search engines can read. CIFF is that of https://github.com/osirrc/ciff, its README and its CommonIndexFileFormat.proto of 2020-03. What CIFF lacks has fields of its own, which a reader of CIFF skips, as the encoding of protobuf has it, https://protobuf.dev/programming-guides/encoding/
The hits of a query in a segment, and how the hits of its parts combine.
Hits are a map of :docs, the numbers of the documents that match, in order, and :scores, the score of each, both arrays.
The hits of a query in a segment, and how the hits of its parts combine. Hits are a map of :docs, the numbers of the documents that match, in order, and :scores, the score of each, both arrays.
No vars found in this namespace.
The segments of an index, and how adding and removing documents changes them, as in Lucene's index, https://lucene.apache.org/core/9_11_1/core/
An index is a map of:
Its :term-fn, if any, is in its metadata, since a function can't be printed with it.
The segments of an index, and how adding and removing documents changes them, as in Lucene's index, https://lucene.apache.org/core/9_11_1/core/ An index is a map of: - :segments, a vector of its segments - :docs, a map of each id to the index of its segment and its number there - :fields, a map of each field to its number - :defaults, the options of its queries, if any Its :term-fn, if any, is in its metadata, since a function can't be printed with it.
No vars found in this namespace.
The documents of each segment of an index that match a prepared query, and their scores, its clauses combined as in Lucene's BooleanQuery, https://lucene.apache.org/core/9_11_1/core/
The documents of each segment of an index that match a prepared query, and their scores, its clauses combined as in Lucene's BooleanQuery, https://lucene.apache.org/core/9_11_1/core/
No vars found in this namespace.
The query language of full-text search, and a query matched against one text without an index, on both platforms.
The comments name the source of each part, and where it departs from it:
The query language of full-text search, and a query matched against one text without an index, on both platforms. The comments name the source of each part, and where it departs from it: - Lucene's simple query parser and FuzzyQuery, https://lucene.apache.org/core/9_11_1/core/ - Elasticsearch's simple_query_string, https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-simple-query-string-query - SQLite FTS5's query syntax and its snippet function, https://www.sqlite.org/fts5.html
The best results of a query: the highest scores first, and equal scores by id, unless a function of a result ranks them.
The best results of a query: the highest scores first, and equal scores by id, unless a function of a result ranks them.
No vars found in this namespace.
BM25F's scores of the documents of a segment that hold a term, from the weights and lengths of their fields.
The comments name the source of each part, and where it departs from it:
BM25F's scores of the documents of a segment that hold a term, from the weights and lengths of their fields. The comments name the source of each part, and where it departs from it: - Robertson and Zaragoza, The Probabilistic Relevance Framework: BM25 and Beyond, Foundations and Trends in Information Retrieval 3(4), 2009, https://www.staff.city.ac.uk/~sbrp622/papers/foundations_bm25_review.pdf - Lucene's BM25Similarity, https://lucene.apache.org/core/9_11_1/core/
No vars found in this namespace.
The segments of an index, on both platforms: the postings of some of its documents in sorted arrays of ints, as a segment of Lucene's index holds them, https://lucene.apache.org/core/9_11_1/core/
The segments of an index, on both platforms: the postings of some of its documents in sorted arrays of ints, as a segment of Lucene's index holds them, https://lucene.apache.org/core/9_11_1/core/
No vars found in this namespace.
The terms of an index that the words of a query stand for, and what scoring them takes.
A prefix stands for the terms that it starts, and a word with typos for the terms within its edits. A query is prepared with them once for all the segments, and the words that a result's document holds only with typos go under its :typos.
The terms of an index that the words of a query stand for, and what scoring them takes. A prefix stands for the terms that it starts, and a word with typos for the terms within its edits. A query is prepared with them once for all the segments, and the words that a result's document holds only with typos go under its :typos.
No vars found in this namespace.
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |