TextSearch API

Base.sum — Method
Base.sum(cluster::AbstractVector{<:SparseVector})

SparseVector counterpart of sum(::AbstractVector{<:Dict}): concatenate every (index, value) pair from every input vector, sort once by index, then combine consecutive equal indices in a single linear pass. Beat every alternative tried (naive +-folding, pairwise tree merging, a k-way heap merge) at every cluster size benchmarked, and — unlike a dense-accumulator approach — its cost does not depend on the vectors' dimension, only on their total number of stored entries. See sadit/TextSearch.jl#25.

All vectors in cluster must have the same dimension (length).

source
Distances.evaluate — Method
evaluate(::NormCosine, a::SparseVector, b::SparseVector)::Float64
evaluate(::Cosine, a::SparseVector, b::SparseVector)::Float64
evaluate(::NormAngle, a::SparseVector, b::SparseVector)::Float64
evaluate(::Angle, a::SparseVector, b::SparseVector)::Float64

SparseVector counterparts of the Dict-based distance functions above, using sparsedot instead of the plain-merge dot for the underlying inner product. NormCosine/NormAngle assume a/b are already normalized; Cosine/Angle do not.

Example

julia> using SparseArrays

julia> evaluate(NormCosine(), sparsevec(UInt32[1, 2], Float32[0.6, 0.8], 10), sparsevec(UInt32[2], Float32[1.0], 10))
0.19999998807907104
source
SparseArrays.sparsevec — Method
sparsevec(vec::Dict{Ti,Tv}, m=0) where {Ti<:Integer,Tv<:Number}

Creates a sparse vector from a Dict-based sparse vector

Example

julia> sparsevec(Dict{UInt32,Float32}(1 => 0.5, 3 => 0.2))
  [1]  =  0.5
  [3]  =  0.2
source
TextSearch.BinaryInvertedFile — Function
BinaryInvertedFile(vocsize::Integer, dist=Dist.Sets.Jaccard())

Creates an empty InvertedFile indexed by a set distance (e.g. Dist.Sets.Jaccard(), Dist.Sets.Dice(), Dist.Sets.Intersection(), Dist.Sets.CosineSet()), suitable for set/token-membership objects (sets or sorted vectors of integer ids).

source
TextSearch._absorb! — Method
_absorb!(a::_VocabularyBatch, b::_VocabularyBatch) -> a

Adds b's counts into a. Tokens new to a are appended in b's order, so when b covers the documents right after a's, the result lists tokens in order of first appearance over both.

source
TextSearch._blended_numtokens — Method
_blended_numtokens(avgdoclen, voc, voc_sample) -> Int64

Resolves blend_vocabularies' avgdoclen option into the numtokens to store: :blend sums the surviving occurrences, :sample matches the sample's average document length, and a positive number is used as that average directly.

source
TextSearch._bow_sizehint — Method
_bow_sizehint(voc::Vocabulary)

Estimated final size (number of unique tokens) of a BOW computed under voc, used to sizehint! it up front and avoid rehashing while it's filled. Uses voc's own avgdoclen, falling back to a small default before voc has seen any training documents.

avgdoclen is a mean document length, so it over-estimates the number of distinct tokens a document holds – measured on Spanish Wikipedia, by about 3.4x for whole articles and 1.6x for paragraphs. That is the harmless direction for a sizehint!: too large wastes a little memory per document, too small brings back the rehashing this exists to avoid. A tighter estimate would need a distinct-tokens-per-document statistic, which a Vocabulary does not keep.

source
TextSearch._candidate_group — Function
_candidate_group(voc, tok, variants, norm, edits=nothing) -> Vector{Tuple{String,Symbol,Int}}

The vocabulary spellings tok could be searched as besides itself, each with the reason it was reached and its document count. Computed spellings come first, then stored ones, then guessed ones, so the order a caller sees follows how much had to be assumed.

The guess is a genuine last resort, gated twice over. It runs only when tok is absent from the vocabulary – unlike an accent, where presence is weak evidence (musica at 9 documents against música's 4,404), an arbitrary edit is well evidenced against by the corpus simply holding the word. And it runs only when the folds above found nothing: if leon already reaches León, guessing at a second, differently-spelled word adds risk to an answer that is already had.

source
TextSearch._check_refit_textconfig — Method
_check_refit_textconfig(expected::TextConfig, got::TextConfig)

Errors unless got is the config a refit requires, comparing field by field via the same predicates a merge uses – == on these structs is unreliable (see the note atop mergeprofiles.jl).

Worth checking loudly: a sample tokenized under a different config produces tokens that do not correspond to the base's, and the blend would then quietly interpolate unrelated counters instead of failing.

source
TextSearch._dequantize_u8 — Method
_dequantize_u8(q, lo, hi) -> Vector{Float32}

Inverse of _quantize_u8, to Float32 because that is what every consumer of these arrays wants and storing more precision than u8 carried would be a lie about the content.

source
TextSearch._derivable_forms — Method
_derivable_forms(folded) -> Tuple

The corpus spellings a folded token can be computed back into: the token itself, its first letter capitalized, and its upper-case form. These need no storage, which is most of what a variant map would otherwise hold – measured on 272,466 Spanish paragraphs, 46,200 of 68,693 folded-to-token pairs are of this shape, so leaving them out cuts the stored map by 70%.

Accents are the opposite case and are why the map exists at all: from practico there is no way to compute whether the corpus writes práctico or practicó, and both are real words.

source
TextSearch._entry — Method
_entry(e, key) -> value
_entry(e, key, default) -> value

Reads key out of a manifest entry whichever way it is keyed.

_save_array builds a Dict{String,Any} and a round trip through JSON3 hands the same entry back keyed by Symbol, so a reader that assumed one of the two worked in production and failed in a test – or would have failed the first time anything decoded an entry it had just built. Accepting both is one line here against a conversion at every call site.

source
TextSearch._extend_lemmas_from_sample — Method
_extend_lemmas_from_sample(base, sample_voc, lemmamap; kwargs...) -> Dict{String,String}

Finds lemma entries for the tokens sample_voc brings that base's vocabulary never had.

The grouping runs over the base and sample vocabularies merged, not over the sample alone, and that is the point: a new inflected form usually belongs to a family whose lemma the base already knows, so "audifonos" must be able to elect the base's "audifono". The merged counts also make :most_frequent prefer the established form over the newcomer. Only the new tokens get entries (see extend_lemmas_morphological), so the base's own clustering decisions are never overruled.

source
TextSearch._fit_min_ndocs — Method
_fit_min_ndocs(p::TextProfile, default::Integer) -> Int

The min_ndocs the fit that produced p was run at, read from its lineage, or default when it is not recorded – profiles fitted before it was recorded, and bases assembled in memory rather than fitted.

This is what lets a refit adopt the bar the profile was actually built at instead of a constant of its own. Note that adopting it can never delete anything: _fit_vocabulary already pruned below it, so every token the base holds clears it by construction. That is the property worth having in a default – it is named rather than magic, and it cannot surprise.

source
TextSearch._fit_vocabulary — Method
_fit_vocabulary(tc, corpus, min_ndocs; label="", verbose=true) -> Vocabulary

Builds a vocabulary under tc and prunes it to tokens in at least min_ndocs documents.

The pruning is not cosmetic: the expansion network is an all-pairs kNN over the vocabulary, so this cuts the most expensive stage of a fit quadratically. Pruning to nothing is an error rather than an empty model, because it is always a mistake in the threshold and silently returning nothing wastes whatever comes after.

source
TextSearch._fold — Method
_fold(tok; lc::Bool, diac::Bool) -> String

One folded spelling of tok: case folded when lc, marks stripped when diac.

Mirrors what the normalization stage would have produced for this word had the profile been configured that way – lowercase first and then Unicode.normalize, in that order and for the reason recorded in _preprocessing: they disagree on the Turkish dotted capital I, where lowercase gives i and case folding alone gives i plus a combining dot, and the second is a token nobody can type.

source
TextSearch._fold_batches! — Method
_fold_batches!(batches) -> _VocabularyBatch

Folds consecutive batches into one, pairwise and in parallel: each round absorbs batch 2k into batch 2k-1, halving the list while keeping it in corpus order. Folding them one by one into the vocabulary instead was the slow half of the build (2.4s of 5.3s with word bigrams over 185k documents and 16 batches), because it is sequential.

source
TextSearch._hash_canonical — Method
_hash_canonical(ctx, v)

Folds v into ctx in an order that does not depend on how a Dict happens to iterate – which in Julia is not a stable order, and would otherwise make an id differ between two runs over the same profile.

source
TextSearch._hash_le — Method
_hash_le(ctx, A)

Folds a numeric array into ctx in little-endian order, for the same reason arraystore.jl writes its members that way: a profile is published and read back by whoever, and an id that depended on the host's byte order would not survive the trip.

source
TextSearch._impute_removed_stopwords — Method
_impute_removed_stopwords(profiles, vocs, voc, pol, doc_freq_threshold) -> Vocabulary

Restores what per-batch stopword removal destroyed, so the merged counters can be read at corpus scale.

fit applies stopwords by tokenizing the batch with them in the pipeline, so a flagged token never enters that batch's vocabulary and its counts are simply gone. When every input flagged it there is nothing to do – it is absent from the merge and stays a stopword. The hard case is a token some inputs flagged and others did not: the merged counters then hold only the batches that kept it, which is a fraction of the truth. Measured on Portuguese Wikipedia, 18 of 35 merged stopwords were in that state, como among them at df=0.049 against a real corpus df above 0.5 – an idf near 3.0 where 0.5 is right.

Neither of the obvious rules works. Dropping such a token deletes content words: on English Wikipedia 35 of 89 were flagged by at most 2 of 48 batches, american, united, states, family and history among them, each made locally ubiquitous by one run of stub articles. Keeping it with the partial counts is the inflated-idf bug.

So the missing counts are imputed instead. A batch that flagged a token recorded no number, but it did record a fact: the token's document frequency there exceeded that batch's threshold. threshold * trainsize is therefore a real lower bound, and the tightest one available. Only batches whose vocabulary genuinely lacks the token are imputed for: a profile may list a stopword it never applied, and its counts are then already exact. Occurrences are scaled by the occs-per-document ratio the batches that kept it observed. Every token then carries its best available estimate and the corpus-scale threshold decides: como lands at 0.399 and stays as a normal token, american at 0.055, the was flagged everywhere and remains a stopword.

The estimate is a lower bound, so a token near the threshold can be judged a normal token when the truth is just above it. That is the safe direction: idf already drives a high-document-frequency token's weight toward zero, while deleting a content word is unrecoverable.

source
TextSearch._input_threshold — Method
_input_threshold(p::TextProfile, default::Real) -> Float64

The doc_freq_threshold fit used on p, read from its lineage, or default when it is not recorded – profiles fitted before the threshold was recorded, and merges of merges, whose summarized lineage drops per-batch params.

source
TextSearch._leader_groups — Method
_leader_groups(items, order, close) -> Vector{Vector{T}}

Groups items around seeds: walking them in order, the first unassigned item becomes a seed and every still-unassigned item that is close to that seed joins it.

This deliberately replaces single-linkage for morphology. Single linkage chains – A~B and B~C merge even when A and C are unrelated – and on a real vocabulary the chains swallow everything sharing a prefix: measured on 143k Spanish tokens it produced a 292-member "family" spanning concentra...cons, and merged cara with caracas and caracalla. Requiring closeness to the seed instead bounds every group by one radius around its lemma, which is also exactly the shape "a lemma plus its variants" should have.

Visiting in the selector's own order (see _selector_key) makes the seed the token the selector would have elected anyway.

source
TextSearch._link_subclusters — Method
_link_subclusters(items, close) -> Vector{Vector{T}}

Single-linkage grouping of items under the predicate close(i, j) (indices into items), by union-find. O(length(items)^2) predicate calls, so callers must keep the input small (by blocking, or by having partitioned already).

source
TextSearch._load_array — Method
_load_array(read_bytes, entry) -> Array

Reads back what _save_array wrote. read_bytes(name) fetches a member's raw bytes – the closure _profile_reader returns, which serves a directory and a zip alike.

A quantized member comes back as Float32, dequantized; an unquantized one comes back at its stored type. Either way the shape recorded in the manifest is restored, so a matrix is a matrix.

source
TextSearch._load_expansion — Method
_load_expansion(read_bytes, entry, voc) -> (query_expansion, distances)

Rebuilds the network from the members _save_expansion wrote, turning ids back into the tokens voc holds. distances is nothing when the profile carries only the ranking.

source
TextSearch._morphology_metric — Method
_morphology_metric(morphology, qgram) -> (prepare, distance)

Builds the pair of functions the subclustering needs: prepare(token_string) computes whatever representation the metric compares, and distance(a, b) scores two prepared representations on [0, 1] (0 = identical surface form).

  • :jaccard: Jaccard distance over character qgram-gram sets, via Dist.Sets.Jaccard. Insensitive to where the difference falls, so it handles prefixal, suffixal and infixal variation alike.
  • :levenshtein: edit distance via Dist.Seqs.Levenshtein, normalized by the longer token so a single threshold means the same thing for short and long words.

Both are normalized deliberately: an absolute edit distance of 2 is negligible between long words and total between short ones, so a raw threshold would behave inconsistently across a real vocabulary.

source
TextSearch._prefix_blocks — Method
_prefix_blocks(voc, ids, prefix_len) -> Vector{Vector{UInt32}}

Buckets ids by their tokens' first prefix_len characters. When linking requires a shared prefix of that length, this blocking is exact – two tokens in different buckets can never link – and it is what makes morphology-first clustering affordable: comparing the whole vocabulary pairwise is O(vocsize^2) (10^10 pairs at 143k tokens), while the sum over buckets is smaller by orders of magnitude.

prefix_len <= 0 cannot block, so everything lands in a single bucket – which also means min_common_prefix = 0 gives up the blocking speedup entirely.

Requiring a shared prefix is not only an optimization: character n-gram similarity is position-blind, so without it abioticos/bioticos and abandonadas/donadas link on sharing nearly every gram despite being different words. It encodes that the target language inflects by suffix, so set it to 0 for languages where that does not hold.

source
TextSearch._profile_reader — Method
_profile_reader(path::AbstractString) -> read_bytes::Function

Returns a read_bytes(name::AbstractString) -> Vector{UInt8} closure that fetches a named member of the profile at path – a plain directory if isdir(path), otherwise a .zip archive (opened once and re-read from memory for every subsequent call). This is what lets load_profile not care which of the two forms it was handed.

It hands back bytes rather than parsed JSON so that one closure serves both kinds of member: load_profile wraps it in JSON3.read for the text ones and passes it straight to _load_array for the binary ones. Parsing here would have meant a second, parallel reader for the binary path – and two readers that must agree about a directory-versus-zip distinction is exactly the kind of duplicated knowledge this format has been bitten by before.

source
TextSearch._qgram_ids — Method
_qgram_ids(s, q, vocab) -> Vector{Int32}

Sorted, deduplicated ids of s's character q-grams, interning grams through vocab so distinct grams never collide. This is the representation SimilaritySearch.Dist.Sets.Jaccard expects (a sorted set as a vector). Tokens shorter than q are represented by themselves, so they can only match identical tokens.

source
TextSearch._quantize_columns — Method
_quantize_columns(P::AbstractMatrix{Float32}) -> (codes, mins, scales)

Quantizes each column over its own extrema, the same way SimilaritySearch's SQu8 does, so the stored codes can be read back either as a QuantizedProjection or as an SQu8Database without being re-quantized into a different thing.

source
TextSearch._quantize_u8 — Method
_quantize_u8(x) -> (q::Vector{UInt8}, lo, hi)

Linearly quantizes x onto 0:255 over its own observed range.

The range is taken from the data rather than fixed, because that is what makes the step size follow what is actually stored: the expansion network's cosine distances occupy [0, 0.938] rather than the [0, 2] cosine allows, and assuming the theoretical range would throw away a bit for nothing.

A constant array (hi == lo) quantizes to all zeros and dequantizes back to that constant, which is exact – so the degenerate case needs no special handling by callers.

source
TextSearch._read_le — Method
_read_le(bytes, T, n) -> Vector{T}

Inverse of _write_le: n elements of type T out of bytes.

The length check is not a formality. A truncated or mismatched member would otherwise reinterpret into plausible-looking garbage, and a profile is meant to be verified.

source
TextSearch._remap_expansion_to_lemmas — Method
_remap_expansion_to_lemmas(network, distances, lemmas) -> (; query_expansion, distances)

Rewrites an expansion network's keys and values through lemmas.

Needed because the network is derived from the unlemmatized vocabulary – it has to be, since the lemma map itself comes from embeddings over that vocabulary – while a profile that applies lemmas no longer has those tokens. Left alone, every inflected entry would be dropped at query time in silence.

Entries that collapse onto the same lemma are merged, keeping each candidate's best rank or distance, and a lemma pointing at itself is dropped. Distances are kept for a token only if every one of its candidates has one, so the two lists can never fall out of alignment.

source
TextSearch._restrict_query_expansion — Method
_restrict_query_expansion(query_expansion, distances, voc) -> (query_expansion, distances)

Drops every network entry naming a token absent from voc, keeping rank order and the parallel distances aligned.

Necessary because the refit prunes: an entry left pointing at a dropped token would be discarded at query time by expand_query! without a word (token2id returning 0), so it would cost file size and tell a reader the network is richer than it is.

source
TextSearch._save_array — Method
_save_array(dir, name, A; quantize=false) -> Dict

Writes A as a binary member name of the profile at dir, and returns the manifest entry that describes it: file name, dtype tag, shape, and the quantization range when quantize is set.

quantize turns a real-valued array into u8 over its own range. Use it where the consumer's tolerance is known to exceed the step size – which is a thing to measure, not to assume; see the note at the top of this file for the measurement that justified it for expansion distances.

The entry is returned rather than written because the caller owns the manifest's layout: an array belongs to some artifact, and only the caller knows where under artifacts it hangs.

source
TextSearch._save_expansion — Method
_save_expansion(dir, p::TextProfile) -> Dict{String,Any}

Writes the expansion network under dir as three binary members over vocabulary ids – column counts, concatenated neighbour ids in rank order, and optionally their quantized distances – and returns the manifest entry describing them.

Errors if the network names a token the vocabulary lacks; see the note above for why that is a refusal rather than a silent drop.

source
TextSearch._selector_key — Method
_selector_key(selector) -> (voc, tid) -> key

Sort key matching _lemma_pick's preference, so a leader pass can visit candidates in the order the selector would elect them and its seed is already the lemma.

source
TextSearch._semantic_clustering — Method
_semantic_clustering(algorithm, dist, wordvecs, num_clusters, m)

Runs the requested SimilaritySearch clustering over the token embeddings, defaulting num_clusters to ceil(sqrt(m)).

source
TextSearch._vote_lemmas — Method
_vote_lemmas(profiles, voc) -> Dict{String,String}

Merges the per-profile token => lemma maps by plurality vote (ties broken by the canonical-token rule: most frequent, then shortest, then lexicographic), keeping only tokens and lemmas that survive in the merged vocabulary.

Independent votes can disagree in ways a single clustering never does – a => b in some profiles and b => a in others – so the winning edges are then followed to a fixed point so that a whole chain collapses onto one canonical token, and any cycle is resolved by electing its most frequent member. Without that pass the merged map could contain cycles, which would make naive lemma lookup non-terminating.

source
TextSearch._with_textconfig — Method
_with_textconfig(model::VectorModel, tc::TextConfig) -> VectorModel

Copies model with tc as its vocabulary's config, sharing the counter arrays rather than copying them – only the config field differs.

source
TextSearch._write_le — Method
_write_le(io, A)

Writes A's elements in little-endian order, whatever the host's own order is.

Fixing the byte order rather than inheriting it is what makes a profile portable: these files are published as release attachments and downloaded by whoever, so a big-endian reader must get the same numbers. On the overwhelmingly common little-endian host this is a single bulk write and costs nothing; elsewhere it converts element by element.

source
TextSearch.bagofwords! — Method
bagofwords!(bow::BOW, voc::Vocabulary, messages)

Computes a bag of words from a multi-field document (a list of texts), accumulating the result into bow. See bagofwords for the non-mutating version.

Example

julia> voc = Vocabulary(TextConfig(), ["hello world"]; verbose=false);

julia> bow = BOW();

julia> TextSearch.bagofwords!(bow, voc, "hello hello world");

julia> bow
Dict{UInt32, Int32}(0x00000002 => 1, 0x00000001 => 2)
source
TextSearch.bagofwords! — Method
bagofwords!(bow::BOW, voc::Vocabulary, tokenlist::TokenizedText)

Accumulates the tokens in tokenlist into bow (a BOW), looking up each token's id in voc; out-of-vocabulary tokens are skipped. Returns bow.

Example

julia> voc = Vocabulary(TextConfig(), ["hello world"]; verbose=false);

julia> TextSearch.bagofwords!(BOW(), voc, tokenize(TextConfig(), "hello hello"))
Dict{UInt32, Int32}(0x00000001 => 2)
source
TextSearch.bagofwords — Method
bagofwords(voc::Vocabulary, messages; isnormalized::Bool=false)

Tokenizes messages (a string or a list of strings) under voc's TextConfig and returns its bag of words (BOW): a token id => occurrence count mapping. An already-computed BOW is returned unchanged.

Example

julia> voc = Vocabulary(TextConfig(), ["hello world"]; verbose=false);

julia> bagofwords(voc, "hello hello world")
Dict{UInt32, Int32}(0x00000002 => 1, 0x00000001 => 2)
source
TextSearch.bagofwords_corpus — Method
bagofwords_corpus(voc::Vocabulary, corpus::AbstractVector; isnormalized::Bool=false, verbose=true)

Computes a list of bag of words (BOWs) from a corpus, one per document, in parallel across threads.

Example

julia> voc = Vocabulary(TextConfig(), ["hello world", "hello there"]; verbose=false);

julia> bagofwords_corpus(voc, ["hello world", "hello there"]; verbose=false)[1]
Dict{UInt32, Int32}(0x00000002 => 1, 0x00000001 => 1)
source
TextSearch.blend_vocabularies — Method
blend_vocabularies(voc_base, voc_sample; kappa=nothing, min_ndocs::Integer=1,
                   avgdoclen=:blend)
    -> Vocabulary

Interpolates two vocabularies into one, treating voc_base as a prior worth kappa documents and voc_sample as observed evidence.

The blend

Read kappa as "the base is worth this many documents". Both counters are then scaled the same way – by the base's average per document – and added to what the sample observed:

base_doc_rate(t) = ndocs_base(t) / N_base       # fraction of base documents holding t
base_occ_rate(t) = occs_base(t)  / N_base       # occurrences of t per base document

ndocs(t)  = ndocs_sample(t) + round(kappa * base_doc_rate(t))
occs(t)   = occs_sample(t)  + round(kappa * base_occ_rate(t))
trainsize = N_sample + kappa
numtokens = sum(occs)                           # recomputed from the survivors

kappa = nothing (the default) means N_sample, which weights the two sides equally; halve it for 1/3 base, double it for 2/3. Expressing the base's authority in documents rather than as a fraction is what makes the knob mean something concrete – and it is the only spelling, since the fraction w is just kappa = N_sample * w / (1 - w) and one knob with two units is one knob too many.

The knob is weak, which is worth knowing before reaching for it. Swept against 1,000 known-item queries at three sample sizes, an 18x range of the base's effective weight (kappa / (N_sample + kappa), from 0.048 to 0.926) moved recall@10 by at most 1.5 points, and never against the base: more prior was mildly better at every sample size, including one 20x larger than the base's own influence would suggest. min_ndocs moves 29 points on the same measurement. Set kappa when you have a reason; the default is not costing you anything measurable.

kappa sets weight, and only weight. It used to decide membership as well – a base-only token whose round(kappa * base_doc_rate) came out zero simply vanished – so the vocabulary shrank hardest at small samples, which is exactly where the base is the only evidence there is. What the vocabulary keeps is the gate's decision now.

Using the same per-document denominator for both counters is what keeps the result a possible corpus. Scaling occs by the base's share of total tokens instead (occs_b/T_b) looks equally reasonable and is not: the two counters then round against different denominators, and a token carried from the base lands with ndocs >= 1 but occs == 0 – present in documents yet never occurring. Sharing the denominator preserves each token's occurrences-per-document ratio, so occs >= ndocs holds by construction.

avgdoclen

By default (avgdoclen = :blend) numtokens is the sum of the surviving occs, so avgdoclen comes out as a weighted mean of the two corpora's average document lengths. That is the honest reading of the blend – the pseudo-documents the prior contributes are base documents, and they are as long as base documents are. But it moves BM25's length normalization toward the base, and when the two corpora's documents are nothing alike the effect is large: Wikipedia-es against 400 product reviews lands at 141 tokens/document at kappa = N_sample and 56 at kappa = N_sample/4, against the sample's own ~21.

avgdoclen = :sample instead sets numtokens so the average matches the sample's, and a positive number sets it to that average directly. This deliberately decouples numtokens from sum(occs), which is safe because that field has exactly one consumer: avgdoclen, and through it BM25Scorer's length normalization. (TpWeighting also divides by a "numtokens", but that one is the document's in-vocabulary token count computed per call in vectorize!, not this.) Use it when the profile will index documents shaped like the sample – which is the usual reason to refit at all – and leave it on :blend when the base's documents are representative of what you will index.

The gate

A token absent from the sample is kept when the base saw it in enough documents:

keep(t) = ndocs_sample(t) > 0 || ndocs_base(t) >= min_ndocs

min_ndocs is the same knob fit_profile applies to its own corpus, in the same unit and with the same meaning: how many documents of evidence a token needs to be in a vocabulary. There is deliberately only one, and it is a document count rather than a rate, because that is the unit the evidence arrives in – "seen in at least 12 base documents" is something a caller can reason about, where a rate cannot be read at all without knowing the base's size. It is also what keeps a token seen in one or two documents of a huge corpus – a typo, an ID – out of the result.

refit_profile defaults it to the bar the base's own fit was run at, read from its lineage, so everything the base has is kept: the fit already decided what counts as attested, and a refit silently re-imposing a different bar would overrule that with a number the caller never chose. Raising it is how a caller asks for a smaller profile; lowering it below the fit's bar does nothing, because the tokens it would admit were never in the base.

Whatever the gate keeps is then representable: ndocs is floored at one document. That floor is what makes min_ndocs a control rather than a suggestion, because without it kappa decides membership, and it decides it backwards.

Measured on a 6,068-token base (16,640 documents) refitted against samples of a different corpus, scoring 1,000 known-item queries against a 10,000-document index:

sampledecided byvocsizerecall@10recall@1archive
100 documentsrounding (old)6630.6320.37618 KB
100 documentsthe gate6,2840.9210.799227 KB
2,000 documentseither9,8240.9490.842257 KB

(Archive sizes are deflated zips; a stored one is about 3x each, and the ratios between them are what the knob trades, not the absolute figures.)

The two agree from about kappa > N_base / (2 * min ndocs_base) upward – 1,664 for that base – since above it the rounding was keeping everything anyway. Below it, the gap is the difference between a profile that answers and one that cannot.

The cost is real, and trading it is what min_ndocs is for: a refitted profile is no longer automatically sample-sized, and at kappa = 100 above it costs 12x the bytes. The curve is steep at the cheap end, on the same measurement – min_ndocs = 12 keeps 84% of the recall gain for 29% of the bytes, min_ndocs = 9 keeps 91% for 41% – so a caller who needs a small profile raises it and reads what it costs. That is a decision; letting kappa make it silently was not.

Note what needs no rule: a token the base did consider important but the sample never shows keeps only its kappa-weighted share, so it survives with reduced weight automatically. Lowering importance is arithmetic; dropping is the only part that needs a decision.

source
TextSearch.decode — Method
decode(voc::Vocabulary, bow::Dict)

Converts a Dict sparse vector indexed by token id (e.g., a BOW) into a Dict indexed by the corresponding token string.

Example

julia> voc = Vocabulary(TextConfig(), ["hello world"]; verbose=false);

julia> decode(voc, bagofwords(voc, "hello hello"))
Dict{String, Int32}("hello" => 2)
source
TextSearch.derive_edits — Method
derive_edits(voc::Vocabulary; minlength=4, min_ndocs=1, verbose=false) -> EditIndex

Indexes voc's tokens for distance-1 Damerau-Levenshtein lookup, so resolve_query_tokens can correct a mistyped query token.

The BK-tree is keyed by DamerauLevenshtein itself, with checkmetric=false. That makes the index approximate, and knowingly so: the restricted (OSA) variant violates the triangle inequality that BKT's pruning relies on, so a true neighbour can be missed. Measured over 3,000 synthetic typos against a 13,871-token Spanish vocabulary, the true source was among the returned candidates 99.2% of the time. The exact alternative – key by Levenshtein, search at radius 2, filter by Damerau-Levenshtein – costs 33% of an exhaustive scan against 7.9% for this route, per BKT's own measurements, which is not a good trade for 0.8%.

min_ndocs

A floor on what may be corrected to. 1 indexes every token, which is usually right because a fitted vocabulary is already pruned (fit_profile's own min_ndocs). Raise it to keep a correction from landing on a token the corpus barely holds.

minlength

A floor on the typed token, and it is a cost gate rather than a quality one – which is worth saying plainly, because it was built expecting the opposite. The uniqueness rule in edit_candidates already subsumes it: measured by typed length over 3,860 synthetic typos, precision given a unique candidate is flat at ~0.999 across every band, including the short ones, because a short token essentially never has a unique distance-1 neighbourhood in the first place:

typed length345678-910+
fraction with a unique candidate0.010.240.510.720.850.910.96
precision when it fires1.001.001.000.9981.000.9991.00

So raising the floor only buys cost: minlength=7 drops coverage from 0.684 to 0.442 and leaves precision at 0.999. The default of 4 skips the band where a lookup is nearly always wasted – at length 1 every single-character token (emoji, punctuation, single letters) is one substitution from every other, ~909 of them in a 60,636-token vocabulary, so uniqueness can never hold.

See also edit_candidates, derive_variants.

source
TextSearch.derive_variants — Method
derive_variants(voc::Vocabulary; min_ndocs=1, maxforms=8) -> Dict{String,Vector{String}}

Builds the query-side variant map for voc: for a folded spelling, the vocabulary tokens it should reach that cannot be computed from it.

This is what lets a profile keep case and diacritics without becoming unsearchable. Preserving them separates senses that folding destroys – measured on Spanish Wikipedia, granada folded returns only the city, while unfolded it also returns the heraldic charge (gules azur bordura), and likewise for cuba (the barrel), concepción (the concept) and león (the animal). The cost is that a query typed leon matches nothing, since the corpus writes León. This map is that bridge, and it is query-only: applying it while indexing would blur the very distinctions it exists to make searchable.

It cannot lean on the query-expansion network, and that is measured rather than assumed: across 16,376 case-twinned Spanish tokens the twin appears in its counterpart's expansion list only 13.6% of the time and at rank 1 in 6.3%. Where it does appear, the pair are function words whose capital is merely sentence-initial; where it does not, the two forms have genuinely different senses and the network is right to keep them apart.

This is computed, never stored

A profile does not carry a variant map: it is a pure function of the vocabulary the profile already holds, so storing one is a second copy of the same information – and a copy that goes wrong, because per-part maps cannot be combined into the map the combined vocabulary yields. Measured on 9 parts of Portuguese Wikipedia, unioning them gave 21,646 keys against the 30,968 the merged vocabulary itself produces: a strict subset missing 30%, tropecar -> tropeçar among them at 132 documents corpus-wide and about 15 per part, under any per-part floor. Deriving from the merged counters instead costs 0.24s over 479,245 tokens.

What it leaves out

Derivable capitalization, per _derivable_forms: madrid -> Madrid is not included because it is computed at query time from the folded form. Two thirds of the pairs are of that shape, and the fraction falls with frequency – 67% of pairs at 5 documents against 53% at 100 – because rare tokens are disproportionately proper nouns whose only variation is a capital, while frequent ones carry real accent alternatives.

Anything below min_ndocs, which defaults to no filtering at all. The floor was worth having while the map was an artifact on disk; now that it is transient, its only remaining job is cost, and there is little to buy: on the 479,245-token Portuguese vocabulary a floor of 1 gives 61,925 keys in 0.62s and 10.3 MB against 30,968 in 0.27s and 4.1 MB at a floor of 20. The 62k map has twice the coverage of exactly the long tail a person is most likely to mistype and least likely to find otherwise. Note also that a vocabulary pruned at fit time already imposes its own floor: these profiles use min_ndocs=5 there, so 1, 2 and 5 here produce identical maps.

Query-time quality is not this function's job either. A bridged spelling is admitted only if it is not negligible beside the commonest spelling of its group – see negligible_ratio in QueryPolicy – which is a relative test and a better one than any absolute count.

Values are ordered by document frequency, most frequent first, and capped at maxforms.

source
TextSearch.download_lsi — Method
download_lsi(nickname_or_url::AbstractString;
             repo::AbstractString="sadit/TextSearch.jl",
             tag::AbstractString=PROFILES_RELEASE_TAG,
             dest::Union{Nothing,AbstractString}=nothing,
             url::Union{Nothing,AbstractString}=nothing,
             force::Bool=false) -> String

Downloads the LSI projection published beside profile nickname (the release asset <nickname>-lsi.zip) and returns its local path: ~/.textsearch/lsi/<nickname>.zip by default (or under $TEXTSEARCH_HOME), or dest. Kept apart from profiles/ because it is not a profile, and it is several times larger than one, so it is only ever fetched on request.

Nothing is checked here. What binds the file to its profile is profile_id, and load_lsi refuses it against any other profile; lsi_summary reads which profile it names without loading the projection.

p = load_profile(download_profile("es"))
lsi = load_lsi(download_lsi("es"), p; outdim=64)
source
TextSearch.download_profile — Method
download_profile(nickname_or_url::AbstractString;
                 repo::AbstractString="sadit/TextSearch.jl",
                 tag::AbstractString=PROFILES_RELEASE_TAG,
                 dest::Union{Nothing,AbstractString}=nothing,
                 url::Union{Nothing,AbstractString}=nothing,
                 force::Bool=false) -> String

Downloads a pre-computed linguistic profile (<nickname>.zip) from a GitHub release or direct URL and saves it locally. By default, installs under ~/.textsearch/profiles/<nickname>.zip (or $TEXTSEARCH_HOME/profiles/<nickname>.zip), or into dest if explicitly specified. Returns the file path of the downloaded archive. Its LSI projection, if the release has one, is fetched separately with download_lsi.

source
TextSearch.dvec — Method
dvec(x::AbstractSparseVector)

Converts an sparse vector into a dict-based sparse vector

Example

julia> dvec(sparsevec(Dict{UInt32,Float32}(1 => 0.5, 3 => 0.2)))
Dict{UInt32, Float32}(0x00000003 => 0.2, 0x00000001 => 0.5)
source
TextSearch.edit_candidates — Method
edit_candidates(ei::EditIndex, tok) -> Vector{UInt32}

The vocabulary ids within Damerau-Levenshtein distance 1 of tok, ascending. Empty when tok is shorter than ei.minlength, so a caller pays nothing for the band where a lookup cannot produce a usable answer.

This returns the whole neighbourhood; deciding whether it is safe to act on belongs to resolve_query_tokens, which acts only when there is exactly one candidate. That rule is measured rather than chosen for tidiness – over 3,000 synthetic typos against a 13,871-token Spanish vocabulary:

rulecoverageprecision when it fires
act on the most frequent candidate, always1.0000.868
length >= 80.3280.968
dominant beats runner-up by negligible_ratio (50x)0.7100.990
exactly one candidate0.7010.999

A unique candidate has a dominance ratio of Inf, so the unique set is a strict subset of the dominant one: those extra 0.9 points of coverage cost an order of magnitude of precision. There is therefore no ratio knob here, and QueryPolicy gains no field.

Why it works at all: an out-of-vocabulary string sits in a far sparser region than a real token. Measured on a 60,636-token vocabulary, a typo string has 1.91 distance-1 neighbours on average (median 1, p90 4) against 18.9 for a token the vocabulary holds.

source
TextSearch.encode — Method
encode(voc::Vocabulary, bow::Dict)

Converts a Dict sparse vector indexed by token string into a Dict indexed by token id, the inverse of decode.

Example

julia> voc = Vocabulary(TextConfig(), ["hello world"]; verbose=false);

julia> encode(voc, Dict("hello" => 2))
Dict{UInt32, Int64}(0x00000001 => 2)
source
TextSearch.expand_query! — Method
expand_query!(vec::SparseVector, voc::Vocabulary, query_expansion;
                  distances=nothing, weight_fn=nothing, normalize::Bool=true) -> vec

Expands a query's sparse tf-idf vector IN PLACE with weighted contributions from each present token's queryexpansion (`queryexpansion, e.g. as produced by [LSI.queryexpansion](@ref)). This mutatesvec-- pass an unnormalized, disposable query vector (vectorize(model, query; normalize=false)`); never a vector you still need afterwards, and never a document vector (documents are never expanded, only queries). Normalizing before calling this would also make the original-vs-queryexpansion weight ratio depend on how many tokens the query had, not on the intended per-query_expansion weighting – that's why normalize (default true) happens here, as the final step.

query_expansion maps a token to its neighbor tokens in rank order (nearest first). For each of vec's original nonzero (tokenID, weight) pairs (captured once, before any appending), looks up its string via gettoken(voc, tokenID); if it's a key of query_expansion, appends weight * weight_fn(...) at token2id(voc, query_expansion) for every neighbor (an OOV queryexpansion – token2id returning 0 – is silently skipped, matching bagofwords!/vectorize!'s existing convention). The appended entries are then merged into vec's existing nonzeros: the combined (nzind, nzval) arrays are heap-sorted by id (reusing SimilaritySearch.heapify!/heapsort!, the same coupled-array sort used to build a SparseVector out of a KnnQueue), duplicate ids (a queryexpansion that was also already present, or reached via two different original tokens) are combined in a single two-pointer reduction pass, and the backing arrays are resize!d down to the final count – an in-place O(n log n) merge, no new allocation for the index/value storage itself.

Weighting modes

There are two, chosen by whether distances is given:

  • rank (distances === nothing, the default): weight_fn receives the neighbor's 1-based rank, and defaults to 1/rank. This is the normal mode. A network's ranking is what transfers between models – distances live in whichever embedding space produced them, and a merged or refitted network's distances are no longer distances in any single space at all.
  • distance: pass distances, a parallel mapping token => Vector{Float32} aligned with query_expansion[token]; weight_fn then receives the distance and defaults to exp(-d) (1.0 at distance 0, decaying smoothly). Pass e.g. d -> d < 0.3 ? 0.5 : 0.0 for a hard cutoff. A token missing from distances, or a short distance list, falls back to rank weighting for the neighbors it does not cover, so a partially-populated distances is safe rather than an error.
source
TextSearch.expand_query! — Method
expand_query!(bow::AbstractDict{<:Integer,<:Real}, voc::Vocabulary, query_expansion) -> bow

Expands a query's bag-of-words IN PLACE by adding every present token's queryexpansion as extra keys – the BM25InvertedFile counterpart of the SparseVector method above. There is no `weightfn/normalize/distanceshere: BM25 scoring (bm25score) never reads the query side's frequencies, only which token ids are present ("query's own frequencies are not used"), so an injected query_expansion only needs to make its id present inbow` – any positive count works, and an id already present (e.g. the query_expansion also appears literally in the query) is left untouched rather than overwritten. This is why a network's distances are not needed on the normal path at all.

As with the SparseVector method, bow's original keys are snapshotted once (via collect) before any insertion, so newly-added queryexpansion ids are never themselves expanded. An OOV queryexpansion (token2id returning 0) is silently skipped, matching bagofwords!'s existing convention.

source
TextSearch.expansion_sources — Method
expansion_sources(r::QueryResolution) -> Vector{String}

The tokens a consumer should look up in a query-expansion network: one per typed token, its group's commonest spelling.

Not every token that was searched, and this is measured rather than stylistic. Expansion over the whole bridged set mixes senses, because bridging deliberately reaches spellings the corpus barely holds and their neighbour lists come from a handful of documents: on Spanish Wikipedia paragraphs SOL (5 documents) gives digitalizada máx chip flash SDRAM, Ano (5) gives the Annobón islands, rio (12, the verb reír) gives llorar tiró Nazgûl, and musica (9, Italian-language paragraphs) gives libreto Puccini Verdi Semiramide. A search for musica de leon returned Antonio Vivaldi and a chess article.

Expanding only what the user typed is not the fix either – it fails in exactly those cases, since the typed form is the rare one. The dominant spelling is: sol bridged gives Sol (1,659 documents) and afelio perihelio eclipses eclíptica, rio gives río and afluente confluencia cauce desemboca. The rarer spellings stay in the search set as matching terms, where a wrong one costs a handful of false positives instead of eight high-idf junk terms.

When nothing was bridged, the group is the typed token alone and this is exactly the token list – so an unbridged query expands as it always did.

source
TextSearch.explain — Method
explain(r::QueryResolution) -> Vector{String}

One human-readable line per typed token that gained something, for a consumer that wants to tell the user what was actually searched.

source
TextSearch.extend_lemmas_morphological — Method
extend_lemmas_morphological(voc::Vocabulary, lemmas::AbstractDict;
                            candidates=nothing,
                            morphology::Symbol=:jaccard, morphology_threshold::Real=0.3,
                            qgram::Integer=2, min_common_prefix::Integer=3,
                            selector::Symbol=:most_frequent) -> Dict{String,String}

Derives additional token => lemma entries for voc from surface similarity alone, and returns only the new ones (merge them into lemmas yourself).

This exists because morphology is the signal that actually groups an inflection family: lemma_clusters uses embeddings only to split a family whose members mean different things, never to form one. So a family can be recovered without fitting any embedding – which is what makes it usable on a vocabulary that arrived after the model was trained, e.g. the tokens a refit's sample brings that its base profile never saw (refit_profile). Nothing here needs wordvecs, an LSI, or a second pass over a corpus. The tradeoff is that no semantic check can veto a grouping, so two look-alike words with unrelated meanings will merge where full lemma_clusters would have kept them apart.

candidates bounds both the cost and the scope. Only prefix blocks containing at least one candidate token are examined – the reason this stays cheap when voc is a whole base vocabulary and only a handful of tokens are new – and only candidate tokens get entries. That restriction is deliberate: a family may well contain two tokens the base's own clustering saw and chose not to link, and silently overruling that decision is not this function's business. Pass nothing to consider everything.

Tokens already keyed in lemmas are skipped, so no chain token -> lemma -> other lemma can be created. Note that under an applied lemma stage they are not vocabulary tokens to begin with.

source
TextSearch.filter_tokens! — Method
filter_tokens!(voc::Vocabulary, text::TokenizedText)

Removes tokens from a given tokenized text based using the valid vocabulary

Example

julia> voc = Vocabulary(TextConfig(), ["hello world"]; verbose=false);

julia> tks = tokenize(TextConfig(), "hello unknownword world");

julia> filter_tokens!(voc, tks); collect(tks)
["hello", "world"]
source
TextSearch.filter_tokens — Method
filter_tokens(pred::Function, model::VectorModel)

Returns a copy of model reduced to the tokens for which pred(t) is true, where t is a (; id, occs, ndocs, weight, token) named tuple (see also filter_tokens(pred, voc::Vocabulary)).

Example

julia> corpus = ["hello world", "hello there", "the cat sat"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> model = VectorModel(IdfWeighting(), TfWeighting(), voc);

julia> vocsize(filter_tokens(t -> t.ndocs >= 2, model))
1
source
TextSearch.filter_tokens — Method
filter_tokens(pred::Function, voc::Vocabulary)

Returns a copy of reduced vocabulary based on evaluating pred function for each entry in voc

Example

julia> voc = Vocabulary(TextConfig(), ["hello world", "hello there"]; verbose=false);

julia> voc2 = filter_tokens(t -> t.ndocs >= 2, voc);

julia> vocsize(voc2)
1
source
TextSearch.fit_profile — Method
fit_profile(textconfig::TextConfig, corpus; kwargs...) -> TextProfile

Distills corpus into a portable TextProfile: vocabulary and counters, weights, a stopword set, an expansion network and a lemma map, with the lineage recording how.

This exists because the ordering is not obvious and getting it wrong is silent. Three passes over the corpus, and each one has to be where it is:

  1. Stopwords before the vocabulary. They are detected from an unfiltered pass whose only product is the list of tokens above the threshold, and then removed while building the vocabulary the encoder trains on – so a filtered token never enters the counters, the factorization, or the network. It is also the single largest stage of a fit: measured on 272,466 Spanish Wikipedia paragraphs, 36.5s of 123.2s.
  2. The encoder, then the lemmas. Lemma families are found by clustering token embeddings, so the embeddings have to exist first.
  3. The vocabulary again, under the lemma map, when lemmas.apply is set. A lemma is a normalization, so it belongs in the TextConfig where every consumer applies it to documents and queries alike and the idf counts an inflection family together instead of splitting it across forms. This pass cannot be folded into an earlier one – the map is derived from embeddings over the vocabulary it rewrites. LSI is deliberately not redone afterwards: the embeddings' job was to find the families and they did.

Keywords, grouped as the concerns they belong to

  • min_ndocs = 1 – drop tokens in fewer documents than this, before the encoder runs.
  • stopwords = (; doc_freq_threshold=0.0, reuse=nothing) – 0 disables detection. reuse takes a set another batch already detected, which is how batches of one corpus end up with identical sets and therefore an exact merge (nothing to impute).
  • encoder = (; outdim=256, scaling=:none, factorization=:auto, wordvectors=nothing) – LSI unless wordvectors hands over external embeddings, in which case they are used as they are. Reading them from a file is the caller's business; this takes vectors.
  • expansion = (; k=8, head_df=0.0, max_target_ratio=50.0, approx=:auto, construction_recall=0.97, search_recall=0.9) – see query_expansion.
  • lemmas = (; apply=false, algorithm=:fft, ...) – see lemma_clusters. apply=false is the default because a base profile computes the map and leaves the choice to whoever tunes from it.
  • verbose = true.

Example

julia> p = fit_profile(TextConfig(), corpus; min_ndocs=5, stopwords=(; doc_freq_threshold=0.5));

julia> isbase(p)
true
source
TextSearch.fold_lemmas — Method
fold_lemmas(voc::Vocabulary, lemmas) -> (; voc, folded, capped, dropped)

Rewrites voc's tokens through lemmas, merging each inflection family's counters into its lemma. Used to bring a base vocabulary that was built without a lemma step onto the same footing as a sample tokenized with one.

The two counters fold differently, and only one is exact:

  • occs is exact. Occurrences are additive, so a family's total occurrence count is the sum of its forms'.
  • ndocs overestimates. A document containing both "casa" and "casas" counts once for each, but once folded it should count once for "casa" – and a vocabulary carries no co-occurrence information to correct with.

That overestimate is why every ndocs is capped at trainsize. The cap is a correctness requirement, not tidiness: ndocs > trainsize makes idf negative (log2((0.5+trainsize)/(0.5+ndocs))) and drives BM25's numerator (trainsize - ndocs + 0.5) below zero. capped reports how often it bit, so the approximation stays visible instead of assumed harmless.

A token whose lemma is absent from voc is dropped rather than reintroduced: that happens when the lemma was itself filtered out at fit time (a stopword, or pruned as rare), and resurrecting it here would smuggle back a token the pipeline deliberately excludes. folded counts remapped tokens, dropped the discarded ones.

source
TextSearch.getndocs — Method
getndocs(voc::Vocabulary, tokenID::Integer)
getndocs(voc::Vocabulary)

Number of documents containing the token tokenID (0 is out-of-vocabulary and yields 0 instead of erroring), or the whole per-token vector when called without a tokenID.

Example

julia> voc = Vocabulary(TextConfig(), ["hello world", "hello there"]; verbose=false);

julia> getndocs(voc, token2id(voc, "hello"))
2
source
TextSearch.getoccs — Method
getoccs(voc::Vocabulary, tokenID::Integer)
getoccs(voc::Vocabulary)

Total occurrences of the token tokenID across the corpus (0 is out-of-vocabulary and yields 0 instead of erroring), or the whole per-token vector when called without a tokenID.

Example

julia> voc = Vocabulary(TextConfig(), ["hello world", "hello there"]; verbose=false);

julia> getoccs(voc, token2id(voc, "hello"))
2
source
TextSearch.getpolicy — Method
getpolicy(p::TextProfile) -> TextConfig
getpolicy(tc::TextConfig) -> TextConfig

The corpus-independent half of a text configuration: normalization and tokenization, with no transformation. This is what two profiles must share exactly to be merged, and what a user can write by hand without any data.

source
TextSearch.gettextconfig — Method
gettextconfig(p::TextProfile) -> TextConfig

The TextConfig this profile tokenizes with: its policy plus the artifacts it applies.

The stage order lives in TokenPipeline and not here, which is the point: lemmas run before the stopword filter, because with the filter first a form that is not itself a stopword survives it and is only then rewritten into one ("las" → "la"), smuggling the stopword back into the vocabulary. This function only decides which artifacts are applied.

source
TextSearch.gettoken — Method
gettoken(voc::Vocabulary, tokenID::Integer)
gettoken(voc::Vocabulary)

The token string for tokenID (0 is out-of-vocabulary and yields "" instead of erroring), or the whole token vector when called without a tokenID.

Example

julia> voc = Vocabulary(TextConfig(), ["hello world", "hello there"]; verbose=false);

julia> gettoken(voc, token2id(voc, "hello"))
"hello"
source
TextSearch.isbase — Method
isbase(p::TextProfile) -> Bool
istuned(p::TextProfile) -> Bool

Whether p is a bootstrap model or one adapted to a dataset, derived from its lineage: a profile with no :refit step is a base, one with a refit is tuned. Reading it off the lineage rather than storing a label means it cannot contradict what actually happened, and a refit of an already-tuned profile stays tuned without a rule for it.

source
TextSearch.istuned — Function
isbase(p::TextProfile) -> Bool
istuned(p::TextProfile) -> Bool

Whether p is a bootstrap model or one adapted to a dataset, derived from its lineage: a profile with no :refit step is a base, one with a refit is tuned. Reading it off the lineage rather than storing a label means it cannot contradict what actually happened, and a refit of an already-tuned profile stays tuned without a rule for it.

source
TextSearch.lemma_clusters — Method
lemma_clusters(voc::Vocabulary, wordvecs::AbstractDatabase;
               algorithm::Symbol=:fft, num_clusters::Integer=0,
               selector::Symbol=:most_frequent, dist=Dist.Cosine(),
               morphology::Symbol=:jaccard, morphology_threshold::Real=0.3,
               qgram::Integer=2, min_common_prefix::Integer=3,
               order::Symbol=:morphology_first, semantic_threshold::Real=1.0) -> Dict{String,String}

Derives a token => lemma map by combining two signals:

  1. Semantic clustering of voc's tokens by their embeddings in wordvecs (column t = embedding of gettoken(voc, t), e.g. from LSI.wordvectors), via one of SimilaritySearch's fft/dnet/randsel/multirandsel. num_clusters = 0 defaults to ceil(sqrt(vocsize(voc))).
  2. Morphological subclustering inside each semantic cluster (morphology, morphology_threshold, qgram – see _morphology_metric), so only tokens that also look alike end up sharing a lemma.

Then one canonical token is elected per group (selector: :most_frequent by default, :shortest, or :shortest_then_most_frequent) and every other member maps to it. The selector also decides seeding order, so it is more consequential than a tie-break: :shortest lets a short misspelling win, and a junk seed fragments the family around it – measured on 143k Spanish tokens, the typo guera seeded a group that swallowed guerra and left guerras stranded. :most_frequent seeds on the form the corpus actually uses, which recovered guerras -> guerra, jugadores -> jugador and concentraciones -> concentracion in the same run. Subclusters of one are left alone. Returns only non-identity entries – a lookup miss means the token is its own lemma.

order decides which signal partitions first:

  • :morphology_first (default): surface-similar families over the whole vocabulary (made affordable by blocking on the required shared prefix), then semantic_threshold splits a family whose members are far apart in embedding space. Whole conjugations collapse correctly this way (abandona, abandonado, abandonar, abandone, ... -> abandono).
  • :semantic_first: the original order – cluster by embedding, then split each cluster by surface similarity. Retained because it is the only order that respects a caller-supplied algorithm/num_clusters, but it fragments inflection families across clusters.

semantic_threshold is a distance under dist, so with the default cosine it lives on [0, 2]; the default 1.0 was picked by measurement rather than taste. Tightening it does not buy precision – it mostly deletes correct inflections (at 0.9 only 4 of 10 probed inflections survive, against 9 of 10 at 1.0), while loosening it past ~1.05 stops catching anything (the artifacts it legitimately removes are cross-language and truncation pairs such as academic/academia and abstracta/abstract).

Step 2 is what makes the result lemma-shaped rather than topic-shaped: embeddings alone put "guerra" next to "belico" rather than next to "guerras" (measured ~2% morphological pairs on Spanish Wikipedia), while surface similarity alone would happily merge "casa" with "caso". Pass morphology=:none to recover the purely semantic behaviour, which elects one representative per semantic cluster and is better described as topic representatives than as lemmas.

Example

lemmas = lemma_clusters(voc, wordvectors(lsi))
lemmas["casas"]   # "casa"
source
TextSearch.lineage_summary — Method
lineage_summary(p::TextProfile) -> String

One line reading how p was produced, e.g. "fit(trainsize=20000) -> merge(n_sources=16) -> refit(kappa=400.0)". Used by textsearch info.

source
TextSearch.list_remote_lsi — Method
list_remote_lsi(; repo::AbstractString="sadit/TextSearch.jl",
                  tag::AbstractString=PROFILES_RELEASE_TAG,
                  url::Union{Nothing,AbstractString}=nothing) -> Vector{NamedTuple}

The LSI projections published beside the profiles of a release, as <nickname>-lsi.zip, in the same shape list_remote_profiles returns: name is the nickname of the profile each one belongs to ("es" for es-lsi.zip), which is what download_lsi takes.

source
TextSearch.list_remote_profiles — Method
list_remote_profiles(; repo::AbstractString="sadit/TextSearch.jl",
                       tag::AbstractString=PROFILES_RELEASE_TAG,
                       url::Union{Nothing,AbstractString}=nothing) -> Vector{NamedTuple}

Queries and returns available pre-computed linguistic profiles from GitHub releases or a custom URL. Returns a vector of (name=nickname, filename=name, size=size_in_bytes, url=download_url, tag=tag). LSI artifacts published beside the profiles (<nickname>-lsi.zip) are not listed; see list_remote_lsi.

The GitHub API allows 60 unauthenticated requests an hour per IP. If GITHUB_TOKEN (or GH_TOKEN) is set, it is sent to api.github.com – and only there, never to a custom url – which raises that limit to the token's.

source
TextSearch.load_lsi — Method
load_lsi(path, profile::TextProfile; outdim=nothing) -> LatentSemanticIndexing

Reads the artifact at path (a directory or a zip) and rebuilds an LSI over profile.

profile is not fetched: a projection is bound to one profile and no other, so if the wrong one is handed in this errors naming the one the artifact was fitted against. Getting that profile is the caller's job – see download_profile – because downloading hundreds of megabytes as a side effect of opening a file is not something this should decide.

outdim truncates: it is the number of LSI components to keep, the same quantity outdim reports and maxoutdim requests – not the neighbour count that query_expansion calls k, which is a different number entirely. Truncating a truncated SVD is exact, since the singular values come out ordered and the quantization is per column, so the coordinates read out of a larger artifact are exactly the first ones it stores.

source
TextSearch.load_profile — Method
load_profile(path::AbstractString) -> TextProfile

Reads back a profile written by save_profile. path may be the directory it produced or a .zip archive of it (see zip_profile); this is auto-detected via isdir(path), and a .zip is read directly from memory with no extraction.

The returned TextProfile rebuilds its own TextConfig from the stored policy and the artifacts marked applied, so what it tokenizes with always matches what it carries.

A profile written by an older format version is refused by name rather than half-parsed: there is no compatibility path, since carrying two layouts is what let the applied and saved copies of an artifact drift apart in the first place.

source
TextSearch.lsi_summary — Method
lsi_summary(path) -> NamedTuple

What an LSI artifact (a directory or a zip written by save_lsi) says about itself, read from its manifest without loading the projection: (profile_id, name, repo, tag, outdim, maxoutdim, scaling).

profile_id is the profile_id of the profile it was fitted against – the same check load_lsi makes, at the cost of a manifest read rather than of dequantizing hundreds of megabytes, which is what a listing or an install check wants:

lsi_summary(download_lsi("es")).profile_id == profile_id(p)
source
TextSearch.merge_profiles — Method
merge_profiles(profiles; doc_freq_threshold=0.5, query_expansion_k=0, rrf_k=60) -> TextProfile

Merges several TextProfiles of one corpus into a single corpus-wide profile:

p = merge_profiles(load_profile.(paths))
save_profile(dir, p)

This is what makes fit's batching usable: batching a large corpus produces one independent profile per batch, and merging folds them back into the single corpus-wide profile.

What is exact, and what is not

  • Vocabulary counts and weights are exact. occs/ndocs/trainsize/numtokens are additive across disjoint document batches, and the weighting scheme is recomputed from the merged counters – so the merged IDF is the true corpus-wide IDF, not an average of per-batch ones. This is the main reason to merge rather than to pick one batch.
  • Query expansion are a rank-fusion consensus, not a recomputation – each input's distances come from its own embedding space (see _fuse_query_expansion). Recomputing them exactly would need the corpus, or a persisted projection, neither of which a profile carries. Scores are summed across inputs, i.e. a consensus count – see _fuse_query_expansion for why normalizing by the inputs that could have voted, though it looks fairer, was measured and rejected. No merge can repair a missing embedding either: a token only one input kept has a neighbour list resting on that one input's opinion.
  • Lemmas are a plurality vote over the inputs' clusterings (see _vote_lemmas).
  • Stopwords are recomputed from the merged counters at doc_freq_threshold, then unioned with the inputs' own sets – a token every input already removed is absent from the merged vocabulary and could not be re-derived, but is still a stopword. A token only some inputs removed is then dropped from the merged vocabulary: the inputs that removed it never recorded its counts, so what survives is a fraction of the truth (measured: como at df=0.049 against a real corpus df above 0.5), and no merge can reconstruct the rest. What the merged counters newly flag is only reported, never dropped – those counts are exact, and keeping them is the reason to merge at all. An artifact counts as applied in the merge if any input applied it.
  • Lineage keeps the inputs' stages, one entry per distinct stage with the number of inputs that contributed it, followed by the :merge step. Merging tuned profiles therefore yields a tuned profile; per-batch params are dropped, since they describe batches the merged profile no longer has.

Inputs must share their policy – normalization and tokenization – and their weighting scheme. Nothing about their artifacts has to match: differing stopword sets union, differing lemma maps vote, differing networks fuse. That asymmetry is the reason policy and artifacts are separate concepts. EntropyWeighting cannot be merged, since recomputing it needs the labeled corpus.

query_expansion_k = 0 keeps as many neighbors per token as the richest input had.

source
TextSearch.merge_voc — Method
merge_voc(voc1::Vocabulary, voc2::Vocabulary[, ...])
merge_voc(pred::Function, voc1::Vocabulary, voc2::Vocabulary[, ...])

Merges two or more vocabularies into a new one. A predicate function can be used to filter token entries.

Note: All vocabularies should had been created with a compatible TextConfig to be able to work on them.

Example

julia> cfg = TextConfig();

julia> voc1 = Vocabulary(cfg, ["hello world"]; verbose=false);

julia> voc2 = Vocabulary(cfg, ["hello there"]; verbose=false);

julia> vocsize(merge_voc(voc1, voc2))
3
source
TextSearch.profile_id — Method
profile_id(p::TextProfile) -> String

A 16-hex-character identifier for what p does to text: its policy, the artifacts it applies, its token sequence, its counters and its weighting. Two profiles share an id exactly when they turn any text into the same vector, so an artifact fitted against one – a stored dense projection, say – can record the id and refuse anything else.

Derived rather than stored, so it cannot disagree with the profile it describes. save_profile also writes it to the manifest as id, for tools that want it without loading anything, and that copy is a convenience: it is not what a check should read.

source
TextSearch.push_token! — Method
push_token!(voc::Vocabulary, token, occs::Integer, ndocs::Integer)
push_token!(voc::Vocabulary, token; occs::Integer=0, ndocs::Integer=0)

Registers token in voc if not already present (assigning it a new id), or accumulates occs/ndocs into its existing entry otherwise. Returns the token's id.

Example

julia> voc = Vocabulary(TextConfig(), 0, 0);

julia> TextSearch.push_token!(voc, "cat"; occs=1, ndocs=1)
0x00000001
source
TextSearch.quantized_wordvectors — Method
quantized_wordvectors(lsi::LatentSemanticIndexing) -> SQu8Database

The LSI embedding of every vocabulary token, as a quantized database SimilaritySearch can search directly.

Search it with ScalarQuant.Cosine(), which reconstructs a true cosine from the codes. Do not use SQu8.NormCosine(): that assumes unit vectors and LSI columns are not unit vectors, so it ranks by length rather than direction – wrongly, and plausibly enough to go unnoticed (a top-10 overlap of 0.25 against the dense answer, with a token its own nearest neighbour 7% of the time).

From a projection that came from load_lsi this reuses the stored codes untouched, so it costs nothing but the wrapper. wordvectors is the dense counterpart, and returns normalized columns as a Float32 matrix.

source
TextSearch.query_tokens — Function
query_tokens(voc::Vocabulary, query, qp::QueryPipeline=QueryPipeline(); policy=qp.policy) -> ResolvedQuery

The query pipeline: turns what a person typed into the terms to search for, and records why.

policy overrides qp.policy for this call and nothing else. The point is that a QueryPolicy is a property of the query while the rest of a QueryPipeline is a property of the corpus: the variant map and the expansion network are derived once, cost real time to derive (0.24s over a half-million-token vocabulary), and belong to the index that holds them. The policy is four scalars. So the policy is what travels, and the maps are reused exactly as they are – which is what lets one index answer both the corrected query and the literal one, the "showing results for … / search instead for …" pair QueryPolicy exists to make possible. Building a whole pipeline per call would rederive nothing but would still be a second place where the pipeline gets assembled.

Reusing the maps under any policy is correct rather than merely cheap: correction never reads variants or edits when policy.correction === :off (see resolve_query_tokens), and expansion is gated on policy.expansion here, so handing over a map or a network that this call has been told not to use changes nothing.

query is raw text, a TokenizedText, or an already-tokenized vector of strings. Text is tokenized under voc's own TextConfig – the same one the documents went through, which it must be, since the vocabulary's ids and counts came from it.

Then, in order:

  1. Correction. resolve_query_tokens replaces spellings the evidence says are wrong and leaves the rest alone – by folding and the variant map, and, for a token the vocabulary does not hold at all, by the edit index. Every spelling it produces weighs 1.
  2. Expansion. For each typed token, the network is looked up under one spelling – the commonest of its corrected group, per expansion_sources – and its neighbours are added with a weight: exp(-d) when qp.distances covers them, 1/rank otherwise. Both are the weightings expand_query! used, kept so the numbers do not move.

A neighbour reachable from two query tokens appears twice, and one the person also typed appears alongside the typed term: those are contributions, and it is the representation that decides what to do with them – queryvector adds them up, querybow and querytokenset collapse them. Neighbours absent from voc are dropped, matching what the vector-level path did with an id of 0.

source
TextSearch.querybow — Method
querybow(voc::Vocabulary, q::ResolvedQuery) -> BOW

The terms as a BOW of vocabulary ids, presence only: every term gets a count of 1 and the weights are discarded.

That is not a shortcut. BM25 scoring never reads the query side's frequencies – only which ids are present – so a weight there would be carried through the whole search and then ignored, and BOW's counts are Int32 anyway. See bm25score.

source
TextSearch.querytokenset — Method
querytokenset(q::ResolvedQuery) -> Set{String}

The terms as a plain set, for a consumer that matches by token intersection and has no use for weights – textsearch search's grep-like matching, for one.

source
TextSearch.queryvector — Method
queryvector(model::VectorModel, q::ResolvedQuery; normalize=true) -> SparseVector

The terms as a weighted sparse vector under model: each term is weighted by the model as usual and then scaled by its QueryTerm weight, so expansion neighbours enter attenuated by rank or distance while typed and corrected spellings enter at full strength.

normalize (default true) is the last step, as it was in expand_query! – a cosine index needs it and doing it before scaling would undo the attenuation.

source
TextSearch.refit_profile — Method
refit_profile(base, sample_voc::Vocabulary; kwargs...) -> NamedTuple
refit_profile(base, sample_docs; kwargs...) -> NamedTuple

Adapts the bootstrap profile base to a dataset, given a sample of it, and returns a new self-contained profile: nothing in the result refers back to base, so it can be saved with save_profile and used on its own.

base is anything with the fields load_profile returns (model, query_expansion, query_expansion_distances, lemmas, stopword_candidates, encoder) – a loaded profile, or one assembled in memory. The return value has that same shape, as merge_profiles's does.

The first form is the core, and takes a Vocabulary the caller built however it liked – streamed, accumulated across runs with push_token!/update_voc!, or from a source that is not a document list at all. It must be built under refit_textconfig(base; apply_lemmas), and is checked against it. The second form is a convenience that tokenizes sample_docs for you.

What is adjusted, and what is not

  • Counters are interpolated by blend_vocabularies, which also decides what the vocabulary keeps – min_ndocs, the same document count fit_profile uses, is that control; the weight vector is then recomputed, which is what makes the tf-idf and BM25 paths tuned by one operation rather than only the former.
  • Lemmas are reused rather than re-derived: the base already paid for them. With apply_lemmas, they enter the TextConfig and the base's counters are folded through the same map (fold_lemmas) so both sides stay comparable. extend_lemmas (corpus form only, since it needs to retokenize) additionally recovers families for tokens the base never saw, from surface similarity alone – see extend_lemmas_morphological. Without it those tokens stay unmerged, which is the price of not fitting an embedding.
  • avgdoclen is :blend by default and can be pinned to the sample's with :sample; see blend_vocabularies for why that choice matters to BM25.
  • Query expansion are inherited, restricted to tokens that survived. No embedding is fit here – that is exactly what makes a refit cheap next to a fit, and the point of bootstrapping.
  • Stopwords are the base's, unchanged. A stopword belongs to a language, not to a dataset, and the base is what models the language; a word that clears doc_freq_threshold here and is not a stopword of the language is a peculiarity of this dataset. Those are reported under verbose, for the operator to act on or ignore, and never filtered. Adding one would also be unsound: the blended counters were collected under the base's set, so the token would keep its counters, its weight and its share of numtokens while the tokenizer could no longer produce it.

EntropyWeighting is rejected, as it is for a merge: its weights are supervised and cannot be re-derived from a profile's contents.

Set verbose to see the vocabulary sizes, how much of the result the base accounts for, and the fold/cap counts from any lemma folding.

source
TextSearch.refit_textconfig — Method
refit_textconfig(base; apply_lemmas::Bool=true, lemmas=nothing) -> TextConfig

The TextConfig a refit of base runs under, and the one a caller building its own sample Vocabulary must tokenize with.

This is public because it is an invariant, not an implementation detail: the blend interpolates two vocabularies token by token, so both sides have to be produced by the same normalization, tokenization, stopword set and lemma step. Tokenizing a sample under anything else silently compares tokens that do not correspond, and the resulting numbers mean nothing.

Everything is inherited from base unchanged, with one deliberate exception: when apply_lemmas is set and base carries a lemma map it did not itself apply, that map enters the config's TokenPipeline, whose lemma stage runs before its stopword stage – the reverse order silently readmits stopwords, since "las" is not in a set holding "la" until after it is rewritten. That is the point of a base profile keeping its lemmas unapplied – whether to lemmatize belongs to the refit, and a tuned model that declines it simply does not carry the map. When lemmas are added here, refit_profile folds the base's own counts through the same map so both sides stay comparable.

lemmas overrides which map is applied, defaulting to base.lemmas. That is what lets a caller lemmatize under a map extended beyond the base's – see extend_lemmas_morphological – while keeping everything else about the config identical.

See also refit_profile, fold_lemmas.

source
TextSearch.resolve_query_tokens — Function
resolve_query_tokens(voc::Vocabulary, tokens, variants=nothing,
                     policy::QueryPolicy=QueryPolicy()) -> QueryResolution

Turns the tokens of a query into the tokens to search with, recording why.

Each typed token defines a group: the vocabulary spellings it could be searched as – itself, the spellings computed from its folded form per _derivable_forms, and whatever the stored variants map holds for that folded form. The group's commonest spelling is its dominant, and a spelling holding less than 1 / policy.negligible_ratio of the dominant's documents is negligible. policy.correction decides what to do with that; see QueryPolicy for the three modes.

Correcting replaces; enriching adds

A typed spelling is dropped from the result exactly when something was bridged for it and the evidence says it was wrong: absent from the vocabulary, or negligible. That is what a correction is, and it is why :off has to exist – a consumer that corrects by default owes the person the same query answered literally, the way a commercial engine offers "search instead for …". explain phrases the two cases differently so a consumer can render that offer.

Where no evidence says otherwise the typed spelling stays and bridging only adds: under :always a healthy sol reaches Sol while remaining itself.

Presence is not evidence of intent

min_ndocs=5 on the vocabulary means unaccented misspellings and foreign-language fragments are tokens, so "it exists, therefore they meant it" fails. On 272,466 Spanish Wikipedia paragraphs ingles holds 7 documents against inglés's 5,188, dia 10 against día's 7,093, musica 9 against música's 4,404 – and of the 518 map keys that are themselves vocabulary tokens, 111 have a spelling ten times commoner and 36 have one fifty times commoner. End to end, search musica returned 0 paragraphs while stopping at the typed form and 314 after correcting it.

The ratio is also what keeps a bridge from dragging in spellings the corpus barely holds, closing a gap where resolution admitted any spelling merely present while derive_variants applied min_ndocs when building the map: sol does not reach SOL (5 documents against 1,659), whose neighbours were digitalizada máx chip flash SDRAM.

On ambiguity

practico typed without an accent is genuinely ambiguous between an adjective and a conjugated verb, and under :always it reaches both. That is the mirror image of what del_diac=false buys on the document side, and the split is the point: the corpus keeps the distinction, so idf and embeddings stay per-sense, while the query bridges it.

source
TextSearch.save_lsi — Method
save_lsi(dir, lsi::LatentSemanticIndexing, profile::TextProfile; name="", repo="", tag="")

Writes lsi's projection into dir as its own artifact, bound to profile.

Only the projection is written: the singular values, the quantized matrix, and enough of the manifest to find and verify the profile. Everything else an LSI needs – the vocabulary, the weights, the tokenizer – is the profile's, and duplicating it is what this layout exists to avoid.

name/repo/tag record where the profile can be fetched from, for a person or a tool to act on; nothing here fetches anything. What makes the binding safe is not that reference but profile_id, recorded beside it and checked by load_lsi.

Store one artifact at the largest outdim worth keeping: truncating a truncated SVD is exact, so load_lsi(...; outdim) serves every smaller one from the same file.

source
TextSearch.save_profile — Method
save_profile(dir::AbstractString, p::TextProfile) -> dir

Serializes a TextProfile into dir (created if missing) as a small directory of plain, human-readable JSON files: one per "large" piece – vocabulary.json, weights.json, and stopwords.json/lemmas.json for whichever artifacts are non-empty – tied together by a manifest.json holding everything else.

The exception is the expansion network, which is three binary members over vocabulary ids (query_expansion_counts.bin, query_expansion_neighbors.bin and, when the profile carries them, u8-quantized query_expansion_distances.bin). Both departures from text are measured rather than assumed. On the published Spanish profile the network and its distances were 141,088,462 bytes of JSON – 86% of the whole file – against 28,680,775 as ids, because every neighbour string was already in vocabulary.json and was being stored again in every list that named it. The quantization is retrieval-identical (arraystore.jl has those numbers). Every other file stays text, and a reader who wants tokens joins against the vocabulary, which is the same join the loader does.

The manifest keeps policy and artifacts apart, which is the point of the layout:

id:         "914ca66f4ddd5767"          # see `profile_id`
policy:     { normalization: {...}, tokenization: {...} }
artifacts:  { stopwords: {file, applied}, lemmas: {file, applied},
              query_expansion: {applied, layout: "csc",
                                counts:    {file, dtype, shape},
                                neighbors: {file, dtype, shape},
                                distances: {file, dtype, shape, quant, columns}} }
lineage:    [ {stage, params}, ... ]

Each artifact is named once, with the marker saying whether the profile applies it. The token transformation is not serialized at all: it is derived from these on load, so the applied lemma map cannot differ from the saved one.

Deliberately NOT a generic object-graph dump (unlike e.g. JLD2): every field is encoded by hand into a small, versioned schema, so every file is fully inspectable/diffable/portable and there is nothing pointer- or code-shaped to accidentally serialize.

Load it back with load_profile, or package it for distribution with zip_profile. Of the extra generators a TokenizationConfig can carry, only QgramGenerators are saved (in format "1.2", which older builds refuse rather than misread); any other generator errors clearly rather than silently mis-saving.

source
TextSearch.sparse_coo — Method
sparse(cols::AbstractVector{<:Dict}, m=0; minweight=1e-9) 
sparse_coo(cols::AbstractVector{<:Dict}, minweight=1e-9)

Creates a sparse matrix from an array of Dict sparse vectors.

Example

julia> cols = [Dict{UInt32,Float32}(1 => 0.5), Dict{UInt32,Float32}(2 => 0.8)];

julia> Matrix(sparse(cols))
2×2 Matrix{Float32}:
 0.5  0.0
 0.0  0.8
source
TextSearch.sparsedot — Method
sparsedot(a::SparseVector, b::SparseVector; small_threshold::Int=30, ratio_threshold::Float64=3.0)

Adaptive dot product between two SparseVectors:

  • both sides have fewer than small_threshold stored entries, or their sizes are within ratio_threshold of each other: a plain linear merge (the same algorithm LinearAlgebra.dot already uses for SparseVector).
  • otherwise (one side much larger than the other — e.g. a short query against a long document): a Hwang-Lin/galloping merge — for each stored entry of the smaller side, an exponential ("galloping") search with memory of the last found position locates its match in the larger side in O(log gap) instead of a full linear scan.

This is deliberately not a method of LinearAlgebra.dot(::SparseVector,::SparseVector): SparseArrays already owns that method (a plain merge), and shadowing it package-wide would be a much more aggressive form of type piracy than TextSearch's existing Dict overloads of dot/normalize! (neither SparseVector nor dot belong to TextSearch, whereas extending dot for Dict doesn't collide with any other package's definitions). evaluate for SparseVector uses sparsedot internally.

Example

julia> using SparseArrays

julia> sparsedot(sparsevec(UInt32[1, 2], Float32[0.6, 0.8], 10), sparsevec(UInt32[2], Float32[1.0], 10))
0.8f0
source
TextSearch.stopword_candidates — Function
stopword_candidates(voc::Vocabulary, threshold::Real=0.5) -> Vector{String}
stopword_candidates(model::VectorModel, threshold::Real=0.5) -> Vector{String}

Flags tokens whose document-frequency ratio getndocs(voc, id) / gettrainsize(voc) exceeds threshold as stopword candidates, sorted by decreasing ratio (most extreme first). A frequency heuristic only – it does not inspect token semantics – so results should be reviewed before being wired into a TokenPipeline's stopwords stage.

Detection is per spelling; removal is per word. Under a profile that keeps case (lc=false) a function word is several vocabulary tokens, and the threshold measures each separately – so it sees the fraction of documents containing a spelling, not a word. Measured on 272,466 Spanish Wikipedia paragraphs at threshold=0.1, detection caught the four commonest paragraph-initial forms (El, En, La, Los) and let 52 twins through, including Las (22,040 documents, df 0.081), A (19,701), Se (17,315), Por (12,891) and De (9,800) – the same function word filtered in one casing and indexed as content in the other. So once a spelling is flagged, every other casing of it that the vocabulary holds is flagged with it.

The alternative – pooling document frequencies across casings before comparing to the threshold – was rejected: the pooled value is not observable from these counters (a document containing both de and De is counted twice, and the sum can exceed 1), and it would shift the calibration of a threshold that was measured per spelling.

Casing is folded; diacritics are not. Folding diacritics would merge té/te, más/mas and sí/si, deleting content words – which is what a profile with del_diac=false exists to prevent. The cost of folding case is small and worth naming: an acronym colliding with a function word goes too (ES, 41 documents, follows es), which lowercasing profiles already did.

Example

candidates = stopword_candidates(voc, 0.5)
textconfig = TextConfig(voc.textconfig; pipeline=TokenPipeline(stopwords=Set(candidates)))
source
TextSearch.table — Method
table(model::VectorModel, TableConstructor)

Builds a Tables.jl-compatible table (e.g., a DataFrame) with one row per token, using TableConstructor (e.g. DataFrame) as the row-table constructor. Columns are token, ndocs, occs, and weight.

Example

julia> using DataFrames

julia> corpus = ["hello world", "hello there", "the cat sat"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> model = VectorModel(IdfWeighting(), TfWeighting(), voc);

julia> table(model, DataFrame)
6×4 DataFrame
 Row │ token   ndocs  occs   weight
     │ String  Int32  Int32  Float32
─────┼────────────────────────────────
   1 │ hello       2      2  0.485427
   2 │ world       1      1  1.22239
   3 │ there       1      1  1.22239
   4 │ the         1      1  1.22239
   5 │ cat         1      1  1.22239
   6 │ sat         1      1  1.22239
source
TextSearch.table — Method
table(voc::Vocabulary, TableConstructor)

Builds a Tables.jl-compatible table (e.g., a DataFrame) with one row per token, using TableConstructor (e.g. DataFrame) as the row-table constructor. Columns are token, ndocs, and occs.

Example

julia> using DataFrames

julia> corpus = ["hello world", "hello there", "the cat sat"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> table(voc, DataFrame)
6×3 DataFrame
 Row │ token   ndocs  occs
     │ String  Int32  Int32
─────┼──────────────────────
   1 │ hello       2      2
   2 │ world       1      1
   3 │ there       1      1
   4 │ the         1      1
   5 │ cat         1      1
   6 │ sat         1      1
source
TextSearch.token2id — Method
token2id(voc::Vocabulary, tok::AbstractString)::UInt32

Looks up the id of tok in voc; returns 0 when tok is out of vocabulary.

Example

julia> voc = Vocabulary(TextConfig(), ["hello world"]; verbose=false);

julia> token2id(voc, "hello")
0x00000001

julia> token2id(voc, "unknown")
0x00000000
source
TextSearch.tokenize_and_append! — Method
tokenize_and_append!(voc::Vocabulary, corpus; isnormalized::Bool=false)

Parse each document in the given corpus and appends each token to the vocabulary.

Example

julia> voc = Vocabulary(TextConfig(), 0, 0);

julia> tokenize_and_append!(voc, ["hello world", "hello there"]);

julia> vocsize(voc)
3
source
TextSearch.update_voc! — Method
update_voc!(voc::Vocabulary, another::Vocabulary)
update_voc!(pred::Function, voc::Vocabulary, another::Vocabulary)

Update voc vocabulary using another vocabulary. Optionally a predicate can be given to filter vocabularies.

Note 1: corpuslen remains unchanged (the structure is immutable and a new Vocabulary should be created to update this field). Note 2: Both voc and another vocabularies should had been created with a compatible TextConfig to be able to work on them.

Example

julia> cfg = TextConfig();

julia> voc = Vocabulary(cfg, 0, 0);

julia> update_voc!(voc, Vocabulary(cfg, ["hello world"]; verbose=false));

julia> vocsize(voc)
2
source
TextSearch.vectorize! — Method
vectorize!(buff::VectorizeBuffer, model::VectorModel, text; normalize=true, minweight=1e-6, isnormalized::Bool=false)

Tokenizes text and weights it using model's local/global weighting scheme, returning the result as a SparseVector{Float32,Int32}; entries with a weight below minweight are dropped, and the result is L2-normalized unless normalize=false. buff is used as scratch space (see VectorizeBuffer; tokenization scratch space is borrowed separately via tokenizerbuffer). See vectorize for a version that manages the scratch buffer for you.

Example

julia> corpus = ["hello world", "hello there", "the cat sat"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> model = VectorModel(IdfWeighting(), TfWeighting(), voc);

julia> buff = TextSearch.VectorizeBuffer();

julia> TextSearch.vectorize!(buff, model, "hello world")
6-element SparseArrays.SparseVector{Float32, Int32} with 2 stored entries:
  [1]  =  0.369076
  [2]  =  0.929399
source
TextSearch.vectorize — Method
vectorize(model::VectorModel, text; normalize=true, minweight=1e-6, isnormalized::Bool=false)

Computes the weighted sparse vector (a SparseVector{Float32,Int32}) representation of text under model. text can be a string or a list of strings (a multi-field document).

Example

julia> corpus = ["hello world", "hello there", "the cat sat"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> model = VectorModel(IdfWeighting(), TfWeighting(), voc);

julia> vectorize(model, "hello world")
6-element SparseArrays.SparseVector{Float32, Int32} with 2 stored entries:
  [1]  =  0.369076
  [2]  =  0.929399
source
TextSearch.vectorize_corpus — Method
vectorize_corpus(model::VectorModel, corpus; normalize=true, minweight=1e-6, isnormalized::Bool=false, verbose=true)

Computes the vectorize representation of every document in corpus, processed in parallel across threads (the batch size is picked automatically from the size of corpus, as in SimilaritySearch.getminbatch).

Example

julia> corpus = ["hello world", "hello there", "the cat sat"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> model = VectorModel(IdfWeighting(), TfWeighting(), voc);

julia> vectorize_corpus(model, corpus; verbose=false)[1]
6-element SparseArrays.SparseVector{Float32, Int32} with 2 stored entries:
  [1]  =  0.369076
  [2]  =  0.929399
source
TextSearch.vocabulary_from_thesaurus — Method
vocabulary_from_thesaurus(textconfig::TextConfig, tokens::AbstractVector)

Creates a Vocabulary directly from a list of tokens (a thesaurus), instead of tokenizing a corpus; every token is registered with occs=1 and ndocs=1.

Example

julia> voc = vocabulary_from_thesaurus(TextConfig(), ["cat", "dog", "bird"]);

julia> vocsize(voc)
3

julia> token2id(voc, "cat")
0x00000001
source
TextSearch.with_applied — Method
with_applied(p::TextProfile; stopwords, lemmas, query_expansion) -> TextProfile

p with different artifacts applied, rematerializing the TextConfig. This is how a consumer turns lemmatization off (textsearch search --no-lemmas) or how a refit decides to apply a base's carried map: change the marker, not the pipeline by hand.

source
TextSearch.zip_profile — Function
zip_profile(dir, zippath=dir * ".zip"; compress=true, compression_level=-1) -> zippath

Packages a profile directory (as written by save_profile) into a single .zip archive at zippath, ready to distribute as one file. load_profile reads a .zip produced this way directly (no extraction needed).

Compression

Entries are deflated. They used to be stored uncompressed – zip_writefile has no compression option and always stores – which went unnoticed because a profile zip is within a kilobyte of the sum of its members, and that reads like framing overhead rather than like a missing feature.

It was the single largest saving available to this format until the expansion network moved to ids, and it still costs nothing in compatibility. Measured on the published Spanish profile rebuilt under the current layout (vocsize 730,320, 5,590,091 edges), 48,392,862 bytes of directory against 27,638,434 deflated, a 43% reduction. Per member, the text ones are where it comes from – vocabulary.json 12,566,619 → 5,449,046, weights.json 7,117,381 → 1,283,405 – while the binary ones give up least, which is what one would want: neighbors.bin 22,360,364 → 15,949,444 and the u8-quantized distances.bin 5,590,091 → 4,676,701. Bytes that were already dense stay dense.

The two changes compose rather than compete: that same profile shipped at 163,945,326 bytes, deflating alone would have made it 73,005,239, and the id layout takes it the rest of the way.

compression_level is passed through to ZipArchives (1 fastest, 9 smallest, -1 its default compromise). compress=false restores the old stored behaviour, which is worth keeping reachable for a caller that is about to compress the archive again anyway.

source
TextSearch.AppliedArtifacts — Type
AppliedArtifacts(; stopwords=false, lemmas=false, query_expansion=false)

Which of a profile's artifacts are in play, as opposed to merely carried.

The distinction is the point of a base profile: a generic model computes a lemma map and a query_expansion network, but whether to apply them belongs to the model being tuned from it. A tuned profile that declines lemmatization simply does not apply the map, and one that never needed it does not carry it either.

stopwords and lemmas are tokenization-time and enter the config a profile derives, which serves fitting, indexing and searching alike – they must, since the vocabulary was counted under it. query_expansion is query-time and works on tokens rather than text, so it does not enter the config; it is applied by expand_query!.

There is no entry for orthographic variants, and that is deliberate: a variant map is a pure function of the vocabulary, so a profile does not carry one to apply or decline. A consumer derives it with derive_variants when it wants to correct a query, and whether to correct at all is a QueryPolicy – a property of the query, not of the profile.

source
TextSearch.BOW — Type
BOW = Dict{UInt32,Int32}

A bag of words: a sparse token id => occurrence count mapping for a single document, as produced by bagofwords/bagofwords!.

Example

julia> BOW(0x00000001 => 2, 0x00000002 => 1)
Dict{UInt32, Int32}(0x00000002 => 1, 0x00000001 => 2)
source
TextSearch.EditIndex — Type
EditIndex(index, context, ids, minlength)

A vocabulary indexed for Damerau-Levenshtein lookup, as derive_edits builds it.

  • index: a SimilaritySearch.BKT over the indexed tokens as Vector{Char}.
  • context: the search context the BK-tree is queried with.
  • ids: database position -> vocabulary id, since the indexed set may be a subset.
  • minlength: typed tokens shorter than this are not looked up at all.

Like the variant map, this is derived, never stored: it is a pure function of the vocabulary, so a profile carrying one would be carrying a second copy of vocabulary.json. Building it costs 2.13s over 60,636 Spanish tokens on 8 threads, which is a per-index cost, not a per-query one.

Tokens are indexed as Vector{Char} rather than String on purpose: Dist.Seqs.* index their arguments positionally (a[i]), which on a String means byte offsets and throws StringIndexError on any multi-byte character – and a real Spanish vocabulary is full of them. This is the same reason _morphology_metric collects its tokens in lemmas.jl.

Thread safety

The distance is the plain Dist.Seqs.DamerauLevenshtein(), whose scratch is empty and therefore allocated per call rather than shared, so concurrent lookups compute correct distances. The one piece of shared mutable state is context's distance-evaluation counter, which SimilaritySearch.add_distance_evaluations! increments non-atomically; concurrent queries can therefore lose counts from that statistic. No result depends on it. Give each concurrent caller its own EditIndex if the count matters.

source
TextSearch.EntropyWeighting — Type
EntropyWeighting()

A GlobalWeighting that scores each token by the empirical entropy of its occurrences across document classes/labels, instead of the plain document frequency used by IdfWeighting — tokens whose occurrences concentrate on few classes get a higher weight than tokens spread uniformly across all classes. Since it needs document labels, it is not built via the generic VectorModel(gw, lw, voc) constructor; use VectorModel(::EntropyWeighting, lw, voc, corpus, labels) instead.

Example

julia> corpus = ["me gusta", "me encanta", "no me gusta", "odio esto"];

julia> labels = ["pos", "pos", "neg", "neg"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> model = VectorModel(EntropyWeighting(), TfWeighting(), voc, corpus, labels; verbose=false);

julia> vectorize(model, "me gusta")
6-element SparseArrays.SparseVector{Float32, Int32} with 1 stored entry:
  [1]  =  1.0
source
TextSearch.LineageStep — Type
LineageStep(stage::Symbol, params::Dict{String,Any})

One step in how a profile came to be: :fit from a corpus, :merge of several profiles, or :refit against a dataset sample. params carries the stage's own details (a fit's encoder and corpus size, a merge's source count, a refit's kappa), as JSON-serializable scalars.

This replaces the encoder field, which had drifted into recording lineage anyway – a merge wrote kind=:merged and a refit kind=:refit into a field named for the encoder.

source
TextSearch.NormalizedEntropy — Type
NormalizedEntropy()

A CombineWeighting that weights a token as 1 - entropy / maxent: tokens that discriminate well between classes (low entropy) get a weight close to 1, tokens spread uniformly across classes (entropy close to maxent) get a weight close to 0.

Example

julia> corpus = ["me gusta", "me encanta", "no me gusta", "odio esto"];

julia> labels = ["pos", "pos", "neg", "neg"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> model = VectorModel(EntropyWeighting(), TfWeighting(), voc, corpus, labels; mindocs=1, comb=NormalizedEntropy(), verbose=false);

julia> vectorize(model, "me gusta")
6-element SparseArrays.SparseVector{Float32, Int32} with 2 stored entries:
  [1]  =  0.9996127
  [2]  =  0.0278291
source
TextSearch.QuantizedProjection — Type
QuantizedProjection <: AbstractMatrix{Float32}

An LSI projection held in its quantized form, dequantizing one coordinate at a time on access: codes[i, j] * scales[j] + mins[j].

This exists so the saving is not undone at load. Dequantizing into a Matrix{Float32} would cost 4x the memory the quantization was chosen to avoid – 187 MB against 47 MB for a 64 x 730,320 projection – and nothing in the LSI path needs a dense matrix: _project_sparse! reads P[row, t] element by element, and wordvectors copies into a fresh matrix of its own. LatentSemanticIndexing accepts it because its parameter asks only for an AbstractMatrix{Float32}.

The same codes can also be handed to SimilaritySearch's SQu8Database, which is an AbstractDatabase with distances defined on the quantized form, for a search pipeline that never dequantizes at all.

source
TextSearch.QueryPipeline — Type
QueryPipeline(; policy=QueryPolicy(), variants=nothing, edits=nothing,
                expansion=nothing, distances=nothing)

Everything the query side of a model needs, as plain data: the QueryPolicy to answer a query under, the orthographic variant map and edit index correction bridges with, and the expansion network with its optional distances.

There is one query pipeline and it lives here, which is the point of this type. Before it, the work existed twice with each copy able to do something the other could not:

  • The vector-level path (expand_query!, called from both inverted files) weighted the terms it added, by rank or by distance. But it iterated a query vector's nonzeros – already token ids – so a typed spelling absent from the vocabulary never reached it and correction was impossible by construction.
  • The string-level path (the textsearch CLI's own) corrected first and expanded only from the commonest spelling of each corrected group, but added its terms unweighted, because it fed a Set for grep-like matching.

Unifying them is therefore not a choice between the two: query_tokens runs on strings, where correction is possible, and emits weights, so nothing is lost. What each consumer does with those weights is its own business – see querybow, queryvector and querytokenset.

variants and edits may each be nothing, which disables that half of correction as surely as policy.correction = :off. Both are pure functions of the vocabulary, so derive them once per model rather than once per query: derive_variants costs 0.6s over 479,245 tokens and derive_edits 2.13s over 60,636. They are separate fields rather than one because they answer different questions – variants bridges spellings of the same word, edits guesses which other word was meant – and a consumer may well want the first without the second.

source
TextSearch.QueryPolicy — Type
QueryPolicy(; correction=:auto, expansion=true, expansion_k=0, negligible_ratio=50)

How a query should be treated, as plain data travelling beside the query text.

The shape follows what commercial search does: a query is answered with the most probable reading of it, and the person is always offered the same query answered literally. Neither correcting nor expanding is a property of the profile – the same profile serves both – so it is not baked into the artifact; it is a mark on the query, and every consumer (the CLI, an application, a service endpoint) passes one of these instead of inventing its own flags.

correction – orthographic bridging, see resolve_query_tokens

  • :auto (default) – bridge only where the evidence says the typed spelling is wrong: it is not in the vocabulary, or it is negligible beside a commoner spelling of the same word. Where it bridges it replaces, because that is what correcting means.
  • :off – search exactly what was typed. This is the "search instead for …" escape, and a consumer that corrects by default owes the person a way to reach it.
  • :always – bridge every token whether or not anything suggests it is wrong. Trades precision for reach, and it is the only way to reach an accented alternative of a spelling that is itself common (typed practico is not negligible beside práctico, so :auto leaves it alone).

negligible_ratio

What "negligible" means: a spelling holding less than 1 / negligible_ratio of the documents of its group's commonest spelling. 1 makes every spelling but the commonest negligible, and Inf makes none of them so. Measured on 272,466 Spanish Wikipedia paragraphs, the ratio between a typed spelling and its commonest alternative decays smoothly – of the 517 map keys that are themselves vocabulary tokens, 266 sit in [1,2) and the counts fall through 87, 50, 39, 23 and 15 to [35,50), then 6, 5, 11 and 15 above – so there is no gap to snap to, but the region around the default is sparse and the choice is not delicate. At 50, musica (9 documents against música's 4,404) is corrected while granada (43 against Granada's 1,438, ratio 33) is not.

expansion

Whether to widen the query with the profile's expansion network, and expansion_k how many neighbours per token (0 = all the profile stored). On by default and turned off on request, the same way as correction: both are guesses about intent, so both are answerable literally.

source
TextSearch.QueryResolution — Type
QueryResolution(tokens, resolved)

The outcome of resolve_query_tokens: tokens is what to search with, and resolved records how each typed token got there.

Reportability is the point of the second field. What this does is spelling correction – a simple, deterministic kind, covering case and diacritics but not transpositions or wrong letters – and a search that silently substitutes what the user asked for owes them a way to see it. "Showing results for X" needs this structure; so does deciding not to correct at all.

source
TextSearch.QueryTerm — Type
QueryTerm(token, source, factor, reason)

One term to search for, where it came from, and how much of that source's weight it carries.

reason is :typed for a spelling the person wrote, :derived, :variant or :edit for a correction (see ResolvedToken), and :expansion for a neighbour the network contributed. For all but the last source == token and factor == 1: a correction is the word the person meant, and nothing about having been misspelled makes it a weaker match. That holds for :edit too – the guess is either right, in which case it deserves full weight, or wrong, in which case attenuating it would only make a wrong answer quieter rather than absent.

An expansion term names the query token whose list it came from, and factor is exp(-d) or 1/rank. It is a factor rather than an absolute weight because that is what a weighted representation needs: queryvector gives the neighbour factor times the source term's own weight in the query, which is what expand_query! did and what its tests pin. A neighbour of a rare, high-idf query word should enter heavier than a neighbour of a common one.

source
TextSearch.ResolvedToken — Type
ResolvedToken(typed, ndocs, kept, added, dominant, dominantdocs)

What happened to one token of a query: the form as typed, how many documents hold that exact spelling (ndocs, zero when it is not a vocabulary token), whether that spelling was kept in the search set, every form that was added for it as form => reason, dominant – the commonest spelling of its group, which is the one allowed to contribute query expansion (see expansion_sources) – and dominantdocs, how many documents hold that. dominant is empty only when no spelling of the group is in the vocabulary at all.

Both counts are carried because a correction is only explicable as a comparison. "appears in only 1,020 documents" is not a reason at corpus scale, where 1,020 documents is a perfectly ordinary word; "1,020 against música's 219,000" is.

kept is false exactly when the token was corrected: something was bridged for it and the evidence said the typed spelling was wrong – absent from the vocabulary, or negligible beside a commoner spelling of the same word. Together with ndocs it tells a consumer which of three things to report: a spelling that was not there at all, one that was there but too rare to be what was meant, or an enrichment that left the typed form standing.

Reasons currently produced:

  • :derived – a spelling computed from the folded form, per _derivable_forms: madrid reaching Madrid, usa reaching USA. Nothing is stored for these.
  • :variant – a spelling that had to be stored, because it cannot be computed: practico reaching practicó, leon reaching León.
  • :edit – a spelling reached by guessing: the single vocabulary token within Damerau-Levenshtein distance 1 of what was typed, guerar -> guerra, a transposition no fold can reach. Offered only for a token absent from the vocabulary and only when that neighbour is unique; see edit_candidates. Nothing is stored for these either – the index is derived from the vocabulary by derive_edits.

Keeping the reason a symbol rather than a Bool is what lets these coexist: a deterministic fold and a distance guess are different kinds of claim, and a consumer telling the user what was searched should be able to distinguish them. That is also the order they are tried in – see _candidate_group – so the cheapest and most certain source is exhausted first.

source
TextSearch.SigmoidPenalizeFewSamples — Type
SigmoidPenalizeFewSamples()

Like NormalizedEntropy, but additionally down-weights tokens seen in very few documents (low ndocs) via a sigmoid-like penalty on log2(ndocs), so that rare tokens don't get an unduly high weight just because they appear in a single class.

Example

julia> corpus = ["me gusta", "me encanta", "no me gusta", "odio esto"];

julia> labels = ["pos", "pos", "neg", "neg"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> model = VectorModel(EntropyWeighting(), TfWeighting(), voc, corpus, labels; mindocs=1, comb=SigmoidPenalizeFewSamples(), verbose=false);

julia> vectorize(model, "me gusta")
6-element SparseArrays.SparseVector{Float32, Int32} with 2 stored entries:
  [1]  =  0.99974245
  [2]  =  0.02269659
source
TextSearch.TextProfile — Type
TextProfile(model; stopwords, lemmas, query_expansion, query_expansion_distances, applied, lineage)

A finished, portable text model: the vocabulary and weights in model, plus the artifacts a corpus produced, plus the lineage that says how it got here.

Each artifact is stored once, and model.voc.textconfig is rebuilt by the constructor as the materialization of the profile's policy plus whichever artifacts applied selects. That is a structural guarantee rather than a convention: there is no way to hold a profile whose tokenizer applies a different lemma map than the one it saves.

Whether a profile is a base or a tuned model is read off the lineage rather than declared – see isbase/istuned – so it cannot contradict the facts, and it answers "where did this come from?" at the same time.

Save and load with save_profile/load_profile, combine batches of one corpus with merge_profiles, and adapt one to a dataset with refit_profile.

source
TextSearch.VectorModel — Type
VectorModel{_G<:GlobalWeighting, _L<:LocalWeighting}

Combines a Vocabulary with a local/global term-weighting scheme (e.g. TfWeighting+IdfWeighting for classical TF-IDF) to turn bags of words into weighted sparse vectors (SparseVector{Float32,Int32}) via vectorize/vectorize!. Build one with VectorModel(gw, lw, voc).

Fields

  • global_weighting: the GlobalWeighting scheme (e.g. IDF), applied per-token, corpus-wide.
  • local_weighting: the LocalWeighting scheme (e.g. TF), applied per-token, per-document.
  • voc: the underlying Vocabulary.
  • maxoccs: maximum per-token occurrence count in voc, used by some local weightings.
  • weight: precomputed per-token global weight (weight[tokenID]).

Example

julia> corpus = ["hello world", "hello there", "the cat sat"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> model = VectorModel(IdfWeighting(), TfWeighting(), voc);

julia> vectorize(model, "hello world")
6-element SparseArrays.SparseVector{Float32, Int32} with 2 stored entries:
  [1]  =  0.369076
  [2]  =  0.929399
source
TextSearch.VectorModel — Method
VectorModel(ent::EntropyWeighting, lw::LocalWeighting, voc::Vocabulary, corpus::AbstractVector, labels::AbstractVector;
    mindocs=3,
    smooth=3,
    weights=:balance,
    comb::CombineWeighting=NormalizedEntropy(),
    verbose=true
)

Creates a VectorModel with EntropyWeighting as its global weighting scheme. Unlike the generic VectorModel(gw, lw, voc) constructor, this one needs the actual corpus and a matching labels vector (one label per document) to compute, for each token, its occurrence distribution across classes and the resulting entropy-based weight.

  • mindocs: tokens occurring in fewer than mindocs documents get weight 0.
  • smooth: additive (Laplace-like) smoothing applied to the per-class occurrence counts before computing entropy, to avoid zero counts.
  • weights: how to reweight classes before computing entropy — :balance compensates for class-size imbalance, :none (or nothing) leaves classes unweighted, or pass an AbstractVector of per-class weights directly.
  • comb: the CombineWeighting strategy combining entropy and evidence into the final per-token weight.

Example

julia> corpus = ["me gusta", "me encanta", "no me gusta", "odio esto"];

julia> labels = ["pos", "pos", "neg", "neg"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> model = VectorModel(EntropyWeighting(), TfWeighting(), voc, corpus, labels; mindocs=1, verbose=false);

julia> vectorize(model, "me gusta")
6-element SparseArrays.SparseVector{Float32, Int32} with 2 stored entries:
  [1]  =  0.9996127
  [2]  =  0.0278291
source
TextSearch.VectorModel — Method
VectorModel(gw::GlobalWeighting, lw::LocalWeighting, voc::Vocabulary; weight=nothing)

Creates a VectorModel for the given vocabulary voc using the local weighting lw (e.g. TfWeighting) and global weighting gw (e.g. IdfWeighting). The per-token global weight vector is computed from voc unless weight is given explicitly (e.g. when reusing weights computed elsewhere, such as EntropyWeighting).

Example

julia> corpus = ["hello world", "hello there", "the cat sat"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> model = VectorModel(IdfWeighting(), TfWeighting(), voc);

julia> length(model.weight)
6
source
TextSearch.VectorModel — Method
VectorModel(voc::Vocabulary) -> VectorModel

TF-IDF over voc: shorthand for VectorModel(IdfWeighting(), TfWeighting(), voc).

It has a name because that combination is what nearly every use wants – 31 of the 43 in this repository – and spelling out two weighting schemes to say "the usual one" reads as though a choice were being made. The three-argument form is how the other twelve say what they mean.

source
TextSearch.VectorizeBuffer — Type
VectorizeBuffer(n=128)

Pooled per-thread scratch space for vectorize!: ids accumulates every in-vocabulary token id seen in a document (with repeats), which is then sorted and run-length-encoded to recover per-token occurrence counts — the same merge-based strategy sum(::AbstractVector{<:SparseVector}) uses, avoiding the Dict allocation/hashing a BOW would need for this per-call, performance-sensitive path.

source
TextSearch.Vocabulary — Type
Vocabulary

Holds the token ⇄ id mapping produced while parsing a corpus, along with per-token occurrence and document-frequency counters. A Vocabulary is the entry point of the processing pipeline: it is built from a TextConfig and a corpus, and is later consumed by VectorModel, BM25Scorer, and bagofwords.

Fields

  • textconfig: the TextConfig used to tokenize the corpus that produced this vocabulary.
  • token: id -> token string table.
  • occs: id -> total number of occurrences of the token across the corpus.
  • ndocs: id -> number of documents containing the token.
  • token2id: token -> id reverse mapping (0 means "unknown token").
  • trainsize: number of documents used to build the vocabulary.
  • numtokens: total number of (non-unique) tokens seen while building the vocabulary.

Example

julia> voc = Vocabulary(TextConfig(), ["hello world", "hello there"]; verbose=false);

julia> vocsize(voc)
3

julia> token2id(voc, "hello")
0x00000001
source
TextSearch.Vocabulary — Method
Vocabulary(textconfig::TextConfig, corpus; buffsize=2^16, isnormalized::Bool=false, verbose=true)

Tokenizes corpus under textconfig and builds the resulting Vocabulary. corpus can be any vector of documents (each document a string or a list of strings) or an iterable/generator of documents (useful for corpora too large to fit in memory); in the generator case, documents are consumed and tokenized in batches of buffsize.

Example

julia> voc = Vocabulary(TextConfig(), ["hello world", "hello there"]; verbose=false);

julia> vocsize(voc), gettrainsize(voc)
(3, 2)
source
TextSearch.Vocabulary — Method
Vocabulary(textconfig::TextConfig, trainsize::Int, numtokens::Int)

Creates an empty Vocabulary (no tokens registered yet) preallocated with capacity hints based on trainsize (following Heaps' law). trainsize and numtokens may be 0 when unknown ahead of time; use push_token! or tokenize_and_append! to fill it, or use Vocabulary(textconfig, corpus) to build it directly from a corpus.

Example

julia> voc = Vocabulary(TextConfig(), 0, 0);

julia> TextSearch.push_token!(voc, "cat"; occs=1, ndocs=1)
0x00000001

julia> vocsize(voc)
1
source
TextSearch._VocabularyBatch — Type
_VocabularyBatch

The counts one batch of documents produces, before they reach the shared Vocabulary: tokens in order of first appearance within the batch, with occs/ndocs parallel to it and index mapping a token to its position.

source
TextSearch.PROFILES_RELEASE_TAG — Constant
PROFILES_RELEASE_TAG

The GitHub release that download_profile and list_remote_profiles fetch from by default: "profiles-<format version>", e.g. "profiles-1.1".

It follows the profile format and not the package version, because the format is what decides whether a file can be loaded. A package release that leaves the format alone keeps pointing at the same assets, and one that bumps it points at a release that does not exist until those profiles are refitted and published, instead of silently fetching files it will refuse.

source
TextSearch._ARRAY_DTYPES — Constant
_ARRAY_DTYPES

The element types a binary profile member may hold, by their manifest tag. Deliberately short: an unknown tag is an error rather than a guess, and widening this is a format change.

source
SimilaritySearch.Intersections.bk! — Function
bk!(output, L, P, findpos::Function=doublingsearch) -> int. size

Computes the intersection of a list of posting lists using the Barbay and Kenyon algorithm.

See umerge! if you need to change how matches are captured (onmatch! function).

source
SimilaritySearch.Intersections.bkt! — Function
bkt!(output, L, [P,], [findpos::Function=doublingsearch]; t::Int=length(L)) -> num. matches

Barybay & Kenyon t-thresholds

See umerge! if you need to change how matches are captured (onmatch! function).

source
SimilaritySearch.Intersections.imerge2! — Method
imerge2!(output, A, B) -> num. matches

Computes the intersection of A and B (sorted arrays), using a merge-like algorithm, and stores it in output or calls onmatch(A, i, B, j) on every match. See bk! for how change the behaviour of matches.

source
SimilaritySearch.Intersections.svs_! — Function
svs(postinglists, intersect2=baezayates!) -> output

Computes the intersection of the ordered lists in postinglists using a small vs small strategy. Accepts an intersection algorithm of two sets.

This method does not give explicit support for onmatch2!

source
SimilaritySearch.Intersections.umerge! — Function
 umerge!(output, L, P=ones(Int32, length(L_)); t::Int=1) -> num. matches

Merges posting lists in L and saves the union in output. The merge result is stored into output array. You can customize how to do this specializing the onmatch!(output, L, P, t::Int) function.

Arguments:

  • L: The array of posting lists, the array can be destroyed in the process.

  • P: The array of current positions in posting lists, i.e., initial state as an array of ones of size $|L|$.

  • t: Computes t-thresholds, i.e., t from 1 (union) to |L| (intersection) of posting lists in L using findpos storing the result set in output.

About the callback function

output, L and P are the arguments same than the input, while t is the actual number of lists having the match. Note 1: you should access L[i][P[i]] to get the entry of the ith list, i.e., $1 \leq t \leq |P|$. Note 2: L and P as container lists will be also modified, the contained lists remain untouched.

source
SimilaritySearch.Intersections.xmerge! — Function
xmerge!(output, L, P=ones(Int32, length(L_)); t::Int=1) -> num. matches

Solves t-threshold set operation using other algorithms choosing among them by given t

Arguments:

  • output: vector like to store the t-threshold set
  • L: the list of posting lists to be merged. The posting lists are left untouched but the container is modified.
  • P: indices of the current merging-state (idem to L)
  • t: the threshold, i.e., t=1 (union) ... t=|L| performs intersection)

Simple wrapper around other specific operations depending on t value

See umerge! if you need to modify the output behaviour.

source
Base.length — Method
length(idx::AbstractInvertedFile)

Number of indexed elements (i.e., objects with postings already built). This can be less than length(database(idx)) if db was grown (e.g. via push_item!(database(idx), obj) or append_items!(database(idx), items) directly) without a following index! call to catch up – mirrors SearchGraph's length/len contract.

source
SimilaritySearch.InvertedFiles._index_block! — Method
_index_block!(idx::AbstractInvertedFile, ctx::InvertedFileContext, sp::Int, n::Int)

Per-type hook for index!: builds postings and any bookkeeping (e.g. sizes/doclens) for database(idx)[sp:n], resizing bookkeeping vectors as needed. The caller (index!) updates idx.len[] afterward.

source
SimilaritySearch.InvertedFiles.has_exact_fastpath — Method
has_exact_fastpath(dist::PreMetric)::Bool

Whether the score computed while merging posting lists (via set_distance_evaluate) is already the exact dist value. When false, search_invfile instead evaluates dist directly against the objects stored in the index's db for every merge candidate — see FallbackInvFileOutput in invfilesearch.jl; raise t above the default 1 to bound how many such evaluations happen per query.

source
SimilaritySearch.InvertedFiles.identiterator — Method
identiterator(dist::PreMetric, obj)

Distance-aware id-only iterator for obj. Defaults to the distance-agnostic identiterator(obj) dispatch tree above; overload this for a specific (DistType, ObjType) pair when the same native object type must generate different candidate ids depending on which distance the enclosing index is built for (e.g. a shingle-based candidate encoding for a sequence distance).

source
SimilaritySearch.InvertedFiles.identiterator — Method
identiterator(obj)

Iterator over the plain ids/keys in obj, for callers that only need to know which ids/keys are present (e.g. InvertedFile building/re-sorting/searching its posting lists, which never need a weight: the handful of distances with an exact fast path score from intersection size and set sizes alone, and any other distance is evaluated directly against the full objects kept in db – see InvertedFile). Dense Vectors are not accepted directly – convert to a SparseVector first (e.g. via SparseArrays.sparse) so the reduction to non-zero components is explicit in the caller's code.

source
SimilaritySearch.InvertedFiles.search_invfile — Method

search_invfile(idx::InvertedFile, ctx::InvertedFileContext, q, Q, res::AbstractKnnQueue, t)

Find candidates for solving query Q using idx. It calls callback on each candidate (objID, dist)

Arguments

  • idx: inverted index
  • q: the query object, only used for distances without an exact fast path (see InvertedFiles.has_exact_fastpath)
  • Q: the set of involved posting lists, see select_posting_lists
  • t: threshold (t=1 union, t > 1 solves the t-threshold problem); for distances without an exact fast path, t also bounds how many real evaluate calls happen per query — raise it to reduce cost.
source
SimilaritySearch.InvertedFiles.set_distance_evaluate — Method
set_distance_evaluate(dist::PreMetric, intersection::Integer, size1::Integer, size2::Integer)

Computes a score for a candidate found while merging posting lists, given the intersection size of the matching posting lists and the total number of non-zero entries of each of the two compared elements (size1, size2). Only defined for the handful of distances with an exact closed form (see has_exact_fastpath) — the resulting value is exact, computed purely from these three integers, with no need to touch the original objects. For any other dist, search_invfile does not call this function at all — it evaluates dist directly against the stored objects for every merge candidate instead (see FallbackInvFileOutput in invfilesearch.jl); use t > 1 to bound how many such evaluations happen per query.

source
SimilaritySearch.InvertedFiles.sort_postinglist! — Method
sort_postinglist!(adj::AbstractAdjList, N)

Sorts a single posting list N (as returned by neighbors(adj, tokenID)) back into the order the merge/search algorithms rely on: ascending by id, for plain token adjacency (UInt32). Override for a different concrete adjacency element type (e.g. a compressed encoding).

source
SimilaritySearch.append_items! — Function
append_items!(idx, ctx, items)

Appends all items elements into the index idx. It work in parallel using all available threads. Grows database(idx) then delegates the actual indexing work to index!, which is the sole emitter of the :add! log event for this batch – this function itself does not log, per the exactly-once contract documented on OBSERVE.

Arguments:

  • idx: The inverted index
  • items: The database of sparse objects, it can be only indices if each object is a list of integers or a set of integers, SparseVectors, among other combinations (see identiterator for the exact set of natively supported object types; dense vectors are not accepted directly — convert with SparseArrays.sparse first).
  • n: The number of items to insert (defaults to all)
source
SimilaritySearch.index! — Method
index!(idx::AbstractInvertedFile, ctx::InvertedFileContext)

Builds postings for every object already present in database(idx) but not yet indexed, i.e. the block database(idx)[length(idx)+1 : length(database(idx))]. It is a no-op (nothing is logged) if db has not grown past length(idx). Mirrors SearchGraph's index!: grow database(idx) first (e.g. push_item!(database(idx), obj) / append_items!(database(idx), items)), then call index!(idx, ctx) to catch up. push_item!/append_items! on idx itself already call this internally, so it only needs to be called explicitly when db was grown directly. This is the sole emitter of the :add! log event for the batch it indexes – see the exactly-once contract documented on OBSERVE.

source
SimilaritySearch.push_item! — Function
push_item!(idx::AbstractInvertedFile, ctx::InvertedFileContext, obj)

Inserts a single element into the index. This operation is not thread-safe.

Arguments

  • idx: The inverted index
  • ctx: the index's context
  • obj: The object to be indexed
source
SimilaritySearch.search — Method
search(idx::AbstractInvertedFile, ctx::InvertedFileContext, q, res::AbstractKnnQueue; t=1)

Searches q in idx using the cosine dissimilarity, it computes the full operation on idx. res specify the query

source
SimilaritySearch.InvertedFiles.DictInvertedFile — Type
const DictInvertedFile{DistType, KeyType, DbType} = InvertedFile{DistType, AdjDict{KeyType, UInt32}, DbType}

A dictionary-backed inverted file mapping posting list keys of type KeyType (e.g., String, Vector{UInt8}, NTuple, Int) to document identifiers (UInt32). Empty or non-existent posting lists are never stored in memory or disk, enabling use over arbitrary or massive key spaces.

Constructors

  • DictInvertedFile(::Type{KeyType}, dist::PreMetric=Dist.Sets.Jaccard(); db::AbstractDatabase=VectorDatabase(Any[]), hint_size::Integer=0)
  • DictInvertedFile(dist::PreMetric=Dist.Sets.Jaccard(); KeyType::Type=Any, db::AbstractDatabase=VectorDatabase(Any[]), hint_size::Integer=0)
source
SimilaritySearch.InvertedFiles.InvertedFile — Type
InvertedFile(vocsize::Integer, dist::PreMetric=Dist.Sets.Jaccard(); db::AbstractDatabase=VectorDatabase(Any[]))

Creates an empty InvertedFile with plain token/set-membership posting lists (AdjType's element type is UInt32), for the given vocabulary size and distance function dist (typically one of the set metrics in Dist.Sets, e.g. Jaccard, Dice, Intersection, CosineSet, RogersTanimoto; or any other PreMetric — e.g. Dist.NormCosine() for sparse-vector/MIPS-style cosine search — via the generic direct-evaluate fallback).

Arguments

  • vocsize: the vocabulary size of the index
  • dist: the distance function to be used in searches

Keyword arguments

  • db: the database that will receive a copy of every indexed object (must support push_item!/append_items! for incremental construction, e.g. a VectorDatabase); defaults to an empty, untyped VectorDatabase. If db is passed already non-empty, call index! once before searching to build postings for its contents.
source
SimilaritySearch.InvertedFiles.InvertedFile — Type
struct InvertedFile{DistType<:PreMetric, AdjType<:AbstractAdjList, DbType<:AbstractDatabase} <: AbstractInvertedFile

A general-purpose inverted index: a sparse matrix-like representation mapping component dimensions (or set elements/tokens) to identifiers (AdjType's element type is UInt32, plain token/set membership; other concrete adjacency element types, e.g. a compressed encoding, can be added by extending getcontainer, internal_push!, and sort_postinglist!). It always keeps the original indexed object in db.

Fields

  • dist: distance function used at search time (e.g. Dist.Sets.Jaccard(), Dist.NormCosine()).
  • adj: posting lists (non-zero id-elements, in rows).
  • sizes: number of non-zero values in each element (non-zero values in columns); resized/populated only up to len[].
  • db: the original indexed objects, one per identifier; always populated by push_item!/append_items!, but may hold more objects than have actually been indexed – see len.
  • len: number of objects already indexed (postings built); may be less than length(database(idx)) if db was grown directly without a following index! call to catch up.

For a handful of distances (the set metrics in Dist.Sets, see InvertedFiles.has_exact_fastpath) the score computed while merging posting lists is already exact, at O(1) cost. For any other distance (including Dist.NormCosine), every merge candidate is instead scored by evaluating dist directly against the objects stored in db, so results for that path are exact too — the number of such evaluations (hence cost) is controlled by the t-threshold parameter of search; raise t above the default 1 to bound the number of real evaluations per query.

source
SimilaritySearch.InvertedFiles._index_block! — Method
_index_block!(idx::BM25InvertedFile, ctx::InvertedFileContext, sp::Int, n::Int)

Decoupled indexing path: builds postings and doclens for idx.db[sp:n], reading the already-encoded SparseVecViews directly out of db (no re-tokenization). Used by index! when db was grown directly (e.g. push_item!(database(idx), docvec)) rather than through the fused append_items!/push_item! entry points.

source
SimilaritySearch.append_items! — Method
append_items!(idx::BM25InvertedFile, ctx::InvertedFileContext, corpus; kwargs...)

Adds every document in corpus to idx, computing each one's bag of words under idx.voc first. corpus can hold raw text (AbstractString), already-tokenized TokenizedText, or pre-tokenized string vectors; a corpus of already-computed BOWs is accepted directly by the generic SimilaritySearch.append_items! method without going through this conversion. See also push_item!.

Example

julia> voc = Vocabulary(TextConfig(), ["hello world", "hello there"]; verbose=false);

julia> invfile = BM25InvertedFile(voc);

julia> ctx = InvertedFileContext();

julia> append_items!(invfile, ctx, ["hello world", "hello there"]);

julia> length(invfile)
2
source
SimilaritySearch.push_item! — Method
push_item!(idx::BM25InvertedFile, ctx::InvertedFileContext, doc)

Adds a single document doc to idx, computing its bag of words under idx.voc first. doc can be raw text (AbstractString), already-tokenized TokenizedText, or a pre-tokenized string vector; an already-computed BOW is accepted directly by the generic SimilaritySearch.push_item! method without going through this conversion. See also append_items!.

Example

julia> corpus = ["hello world", "hello there"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> invfile = BM25InvertedFile(voc);

julia> ctx = InvertedFileContext();

julia> append_items!(invfile, ctx, corpus);

julia> push_item!(invfile, ctx, "hello again");

julia> length(invfile)
3
source
SimilaritySearch.search — Method
search(idx::BM25InvertedFile, ctx::InvertedFileContext, qtext, res::AbstractKnnQueue; t::Int=1)

Solves a top-k query over idx for qtext (raw text, TokenizedText, or an already-computed bag of words), accumulating matches into res (an AbstractKnnQueue, e.g. KnnSorted or KnnHeap). Documents are ranked by BM25 score (stored internally as a negative value in res, so lower "distance" still means better match, consistent with SimilaritySearch.jl's convention). Returns res.

Example

julia> corpus = ["hello world", "hello there", "the cat sat"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> invfile = BM25InvertedFile(voc);

julia> ctx = InvertedFileContext();

julia> append_items!(invfile, ctx, corpus);

julia> res = knnqueue(KnnSorted, 2);

julia> search(invfile, ctx, "hello", res);

julia> collect(IdView(res))
UInt32[0x00000001, 0x00000002]
source
TextSearch.BM25._bm25_fused_index_and_grow! — Function
_bm25_fused_index_and_grow!(idx, ctx, items, startID, n, tol=1e-6)

Encodes each raw/BOW document in items[1:n] into its SparseVecView term-frequency representation, registers its postings into idx.adj, and stores the vector into idx.db – all in a single pass (this is the "fused" behavior append_items!/push_item! use for raw-text/ BOW input, kept for efficiency: unlike the generic InvertedFile, BM25InvertedFile's db element type is derived from its input, so there is no cheap way to grow db ahead of indexing for this path – see _index_block! for the decoupled path instead, used when db is grown directly with already-encoded SparseVecViews).

source
TextSearch.BM25.bm25_internal_push_object! — Method
bm25_internal_push_object!(idx, docID, obj, tol) -> (doclen, docvec)

Registers obj (a pair (tokenID, freq) iterable) into idx.adj under docID and builds its SparseVecView term-frequency representation (docvec, to be stored at docID in idx.db by the caller). Returns obj's total token count (doclen) and docvec.

source
TextSearch.BM25.bm25_register_postings! — Method
bm25_register_postings!(idx::BM25InvertedFile, docID::Integer, docvec) -> doclen

Registers an already-encoded document vector docvec (anything pairiterator-compatible, typically the SparseVecView already stored at idx.db[docID]) into idx.adj under docID, and returns its token count (doclen). This is the postings-only half of bm25_internal_push_object! – it does not parse/build a SparseVecView, since docvec is assumed to already be one (e.g. read back from idx.db by _index_block!).

source
TextSearch.BM25.bm25score — Method
bm25score(bm25::BM25Scorer, voc::Vocabulary, query::SparseVectorLike, doc::SparseVectorLike)::Float32

Computes the BM25 relevance score of doc for query – each a term-frequency sparse vector (a SparseVecView, e.g. one of BM25InvertedFile's own db entries via database, or a SparseVector) – by merging their nonzero indices (nzind, assumed sorted ascending) in a single linear pass and summing tokenscore at every token id present in both. Higher is more relevant. query's own frequencies are not used (BM25 doesn't weight by query-side term frequency), only which tokens it contains.

Example

julia> corpus = ["hello world", "hello there", "the cat sat"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> bm25 = BM25Scorer(voc);

julia> invfile = BM25InvertedFile(voc);

julia> ctx = InvertedFileContext();

julia> append_items!(invfile, ctx, corpus);

julia> bm25score(bm25, voc, database(invfile)[1], database(invfile)[1])  # "hello world" scored against itself
2.9917173f0

julia> bm25score(bm25, voc, database(invfile)[1], database(invfile)[2])  # "hello world" scored against "hello there"
0.96917987f0
source
TextSearch.BM25.tokenscore — Method
tokenscore(bm25::BM25Scorer, toknumdocs, doclen, tokfreqindoc)

Computes the BM25 contribution of a single token to a document's score, given the number of documents containing the token (toknumdocs, i.e. its document frequency), the document's length in tokens (doclen), and the token's frequency in the document (tokfreqindoc). Used internally by bm25score and BM25InvertedFile search.

Example

julia> corpus = ["hello world", "hello there", "the cat sat"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> tokenscore(BM25Scorer(voc), 2, 3, 1)
0.8908208f0
source
TextSearch.BM25.BM25InvertedFile — Type
BM25InvertedFile{AdjType<:AbstractAdjList,DbType<:AbstractDatabase} <: AbstractInvertedFile

An inverted-file index (built on top of SimilaritySearch.InvertedFiles) that answers approximate/exact top-k queries ranked by BM25 relevance. Build it with BM25InvertedFile(voc), populate it with append_items!/push_item!, and query it with search (from SimilaritySearch.jl).

Follows the same design as SimilaritySearch.InvertedFiles.InvertedFile: adj only ever stores plain document ids (AdjType's element type is UInt32, exactly like the generic InvertedFile); every other per-document detail needed to score a match – term frequencies – is fetched from db instead of being duplicated into the posting lists.

Fields

  • voc: the Vocabulary shared by every indexed document (also used to tokenize/encode query text).
  • bm25: the BM25Scorer used to rank matches.
  • adj: the adjacency list of posting lists (one per token id), mapping each token to the ids of the documents containing it.
  • doclens: number of tokens per indexed document.
  • db: each indexed document's term-frequency vector, one SparseVecView (token id => UInt32 frequency) per document, always populated by push_item!/ append_items!; may hold more documents than have actually been indexed – see len.
  • len: number of documents already indexed (postings built); may be less than length(database(idx)) if db was grown directly (e.g. push_item!(database(idx), docvec)) without a following index! call to catch up. Growing db directly with pre-computed SparseVecViews and then calling index!(invfile, ctx) is supported and builds postings from the already-stored vectors; the raw-text/BOW-taking append_items!/push_item! methods remain fused (encode+store+register in one pass) for efficiency.
  • query_expansion: nothing, or a query-expansion network (e.g. as produced by LSI.query_expansion) used to enrich queries via expand_query!. Never applied to documents, only to queries.

Example

julia> using SimilaritySearch

julia> corpus = ["hello world", "hello there", "the cat sat"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> invfile = BM25InvertedFile(voc);

julia> ctx = InvertedFileContext();

julia> append_items!(invfile, ctx, corpus);

julia> length(invfile)
3

julia> res = knnqueue(KnnSorted, 2);

julia> search(invfile, ctx, "hello", res);

julia> collect(IdView(res))
UInt32[0x00000001, 0x00000002]
source
TextSearch.BM25.BM25InvertedFile — Method
BM25InvertedFile(p::TextProfile; k1=1.2f0, b=0.75f0, δ=1f0, policy=QueryPolicy(), expansion=p.applied.query_expansion)

Creates an empty BM25InvertedFile from a fitted TextProfile – which is the point of having profiles at all, so it is worth being precise about what comes from where.

The scorer's statistics are the profile's: trainsize and avgdoclen are the corpus the profile was fitted on, and the document frequency behind every idf is read from its vocabulary at search time. The document lengths are the index's, filled in by append_items! as documents arrive. That split is the whole idea: a profile fitted on 6,665,754 Portuguese paragraphs lends its idf and its length normalization to an index holding 20,000 of them, instead of each small index inventing statistics from what little it has.

Tokenization is the profile's too, since the vocabulary's ids and counts came from it.

How queries are answered follows the profile, with one deliberate asymmetry. Expansion is gated by the profile: the network is handed to the index only when applied.query_expansion says the profile endorses it, since it is an artifact the profile may carry without meaning it to be used – pass expansion=true to take it anyway, which is what a base profile needs. Correction is gated by the policy, because it depends on nothing but the vocabulary, which every profile has; the variant map is derived once here rather than once per query, and comes out empty at no cost for a profile that folds case and diacritics.

Example

julia> idx = BM25InvertedFile(profile);                    # corrects, does not expand

julia> idx = BM25InvertedFile(profile; expansion=true);    # ...and expands anyway

julia> idx = BM25InvertedFile(profile; policy=QueryPolicy(correction=:off));   # literal queries
source
TextSearch.BM25.BM25InvertedFile — Method
BM25InvertedFile(voc::Vocabulary; k1=1.2f0, b=0.75f0, δ=1f0, query_expansion=nothing, distances=nothing, query=nothing)

Creates an empty BM25InvertedFile, fitting its BM25Scorer from voc (see BM25Scorer(voc) for k1/b/δ). Populate it with append_items!/push_item!.

How queries are answered is a QueryPipeline, stored on the index. query_expansion (e.g. as produced by LSI.query_expansion) is the short way to say "expand with this network"; pass query instead to also correct spellings, which needs a variant map – and note that deriving one per query would be far too slow, which is why it lives on the index.

Example

julia> voc = Vocabulary(TextConfig(), ["hello world"]; verbose=false);

julia> invfile = BM25InvertedFile(voc);

julia> length(invfile)
0
source
TextSearch.BM25.BM25Scorer — Type
BM25Scorer

Precomputed coefficients for the Okapi BM25 (BM25+) scoring function, used to rank documents against a query given per-token document frequencies and document lengths. Build one with BM25Scorer(voc) or BM25Scorer(trainsize, avgdoclen); score individual (token-frequency, document-length) pairs with tokenscore, or whole query/document bags of words with bm25score. BM25InvertedFile uses a BM25Scorer internally to answer top-k queries efficiently.

Fields

k1_plus_1, k1_mult_1_min_b, and k1_mult_b_div_avg_doc_len are combinations of the BM25 k1/b hyperparameters and the corpus' average document length, precomputed for faster scoring; δ is the BM25+ lower-bound correction term; trainsize is the number of documents the corpus statistics were computed from.

Example

julia> corpus = ["hello world", "hello there", "the cat sat"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> bm25 = BM25Scorer(voc);

julia> bm25.trainsize
3
source
TextSearch.BM25.BM25Scorer — Method
BM25Scorer(trainsize::Integer, avgdoclen::AbstractFloat; k1=1.2f0, b=0.75f0, δ=1f0)

Creates a BM25Scorer for a corpus of trainsize documents with average document length avgdoclen. k1 controls term-frequency saturation, b controls document-length normalization, and δ is the BM25+ lower-bound correction.

Example

julia> bm25 = BM25Scorer(3, 3.5);

julia> bm25.trainsize
3
source
TextSearch.BM25.BM25Scorer — Method
BM25Scorer(voc::Vocabulary; k1=1.2f0, b=0.75f0, δ=1f0)

Creates a BM25Scorer using the training size and average document length already computed in voc (see trainsize and avgdoclen).

Example

julia> corpus = ["hello world", "hello there", "the cat sat"];

julia> voc = Vocabulary(TextConfig(), corpus; verbose=false);

julia> BM25Scorer(voc).trainsize
3
source
Base.empty! — Method
empty!(buff::TokenizerBuffer; normtext=true, tokens=true, unigrams=true)

Clears the requested scratch fields of buff in place. Returns buff.

source
TextSearch.Tokenizer._preprocessing — Method
_preprocessing(config::TextConfig, text) -> AbstractString

Case folding and the three group-substitutions, in one replace pass rather than three.

Each replace allocates a full copy of the text, so chaining them costs three allocations and three scans per document; one call with several pairs costs one of each. Measured on 60,000 Spanish Wikipedia paragraphs (33.1M characters), normalization drops from 8.96s to 8.36s, and the output is byte-identical over 70,000 documents. The pairs stay in their original order because a multi-pair replace tries them in order at each position, which reproduces what the chain did: a URL is consumed whole by re_url before re_num can see the digits inside it.

The lowercase is NOT redundant with the casefold=true that normalize_text passes to Unicode.normalize, which is what it looks like. They differ on the Turkish dotted capital I (U+0130): lowercase maps it to i, while Unicode case folding maps it to i plus a combining dot above, so with del_diac=false the token becomes i̇lhan – which no one can type, so a query for ilhan stops matching. Measured on real text, 28 of 10,000 Spanish articles contain it. It also has to run before the regexes, since re_url is case-sensitive.

source
TextSearch.Tokenizer.alltokengenerators — Method
alltokengenerators(cfg::TokenizationConfig)::Vector{AbstractTokenGenerator}

Builds the full, ordered list of AbstractTokenGenerators cfg runs: the built-in ones implied by cfg.nlist, followed by cfg.generators (any extra/custom generators). Called once per tokenize invocation.

Example

julia> alltokengenerators(TokenizationConfig(nlist=[1, 2]))
2-element Vector{AbstractTokenGenerator}:
 UnigramGenerator()
 NWordGenerator(2)
source
TextSearch.Tokenizer.apply_pipeline — Method
apply_pipeline(p::TokenPipeline, tok) -> Union{Nothing,String}

Runs p's stages over one token, in the fixed order (lemma rewrite, then the stopword filter), returning nothing when the token is dropped. On the hot path: an inactive stage is a === nothing check the compiler can hoist, so an identity pipeline costs two comparisons per token.

source
TextSearch.Tokenizer.flush_token! — Method
flush_token!(buff::TokenizerBuffer, pipe::TokenPipeline, gen::AbstractTokenGenerator, mark_token_type::Bool)

Pushes the token accumulated in buff.io to the token list, applying gen's tokentag (when mark_token_type) and the TokenPipeline's stages; discards empty strings and tokens the pipeline drops.

source
TextSearch.Tokenizer.isemoji — Function
isemoji(c::Char, emojis::Set{Char}=DEFAULT_EMOJIS)::Bool

Tests whether c is one of the emoji characters in emojis (by default, DEFAULT_EMOJIS, the set known to TextSearch, loaded from emojis.txt). Used by normalize_text when TextConfig's group_emo option is set.

Example

julia> isemoji('😀')
true

julia> isemoji('a')
false
source
TextSearch.Tokenizer.normalize_text — Method
normalize_text(config::TextConfig, text::AbstractString, output::Vector{Char}; limits::Bool=true, isnormalized::Bool=false)

Normalizes a given text using the specified transformations of config. If isnormalized=true, skips preprocessing and normalization passes, writing text directly to output.

Example

julia> buff = Char[];

julia> normalize_text(TextConfig(), "Café", buff);

julia> String(buff)
" cafe "
source
TextSearch.Tokenizer.qgrams — Method
qgrams(gen::QgramGenerator, buff::TokenizerBuffer, pipe::TokenPipeline, mark_token_type)

Emits every character gen.q-gram of buff.normtext, reading runs of blanks as one blank (see QgramGenerator). normtext is left as it is, since other generators read it too.

source
TextSearch.Tokenizer.tokenize — Method
tokenize(textconfig::TextConfig, text)
tokenize(copy_::Function, textconfig::TextConfig, text)

tokenize(textconfig::TextConfig, text, buff)
tokenize(copy_::Function, textconfig::TextConfig, text, buff)

Tokenizes text using the given configuration. The tokenize makes heavy usage of buffers, and when these buffers are shared it is mandatory to create a copy of the result (buff.tokens).

Change the default copy function to make an additional filtering of the tokens. You can also pass the identity function to avoid copying.

Example

julia> collect(tokenize(TextConfig(), "Hello world!!"))
["hello", "world", "!!"]
source
TextSearch.Tokenizer.tokenize_corpus — Method
tokenize_corpus(textconfig::TextConfig, arr; isnormalized::Bool=false, verbose=true)
tokenize_corpus(copy_::Function, textconfig::TextConfig, arr; isnormalized::Bool=false, verbose=true)

Tokenize a list of texts. The copy_ function is passed to tokenize as first argument.

Example

julia> corpus = ["hello world", "the cat sat"];

julia> toks = tokenize_corpus(TextConfig(), corpus; verbose=false);

julia> collect(toks[1])
["hello", "world"]
source
TextSearch.Tokenizer.tokenize_paragraphs — Method
tokenize_paragraphs(text::AbstractString)::Vector{String}
tokenize_paragraphs(textconfig::TextConfig, text::AbstractString)::Vector{String}
tokenize_paragraphs([textconfig::TextConfig,] arr::AbstractVector)::Vector{String}

Splits text into paragraphs separated by two or more newlines (\n\n+ or \r\n\r\n+). Trims leading and trailing whitespace from each paragraph and filters out empty paragraphs. When textconfig is provided, normalizes the text before paragraph splitting.

source
TextSearch.Tokenizer.tokenize_sentences — Method
tokenize_sentences(text::AbstractString)::Vector{String}
tokenize_sentences(textconfig::TextConfig, text::AbstractString; isnormalized::Bool=false)::Vector{String}
tokenize_sentences([textconfig::TextConfig,] arr::AbstractVector; isnormalized::Bool=false)::Vector{String}

Splits text into sentences using sentence-ending punctuation (., !, ?) followed by whitespace or newlines. Trims leading and trailing whitespace from each sentence and filters out empty sentences. When textconfig is provided, normalizes each sentence after splitting.

source
TextSearch.Tokenizer.tokenizerbuffer — Method
tokenizerbuffer(f)

Borrows a TokenizerBuffer from Tokenizer's own pool, passes it to f, and returns it to the pool afterwards. Unlike the buffer-less tokenize methods (which release their borrowed buffer before returning), the buffer stays borrowed for the whole extent of f, so it is safe to alias its contents (e.g. via borrowtokenizedtext) as long as they are consumed inside f.

Example

julia> TextSearch.Tokenizer.tokenizerbuffer() do buff
           tokenize(borrowtokenizedtext, TextConfig(), "hello world", buff) |> collect
       end
["hello", "world"]
source
TextSearch.Tokenizer.tokentag — Method
tokentag(gen::AbstractTokenGenerator)::Union{Char,Nothing}

The single-character tag appended (as \ttag) to every token gen produces when mark_token_type=true. Defaults to nothing (untagged).

source
TextSearch.Tokenizer.AbstractTokenGenerator — Type
AbstractTokenGenerator

Abstract type for a single token-producing strategy inside a TokenizationConfig's generators list. TokenizationConfig's nlist keyword argument is convenience sugar that builds the built-in generators below; passing generators directly (or mixing in your own AbstractTokenGenerator subtype) is how new kinds of tokens can be added without touching TokenizationConfig or the tokenizer's dispatch logic (e.g. character q-grams, skip-grams, or collocations, none of which are built-in anymore).

Implementing a new generator kind requires:

  • a struct <: AbstractTokenGenerator holding whatever parameters it needs;
  • needs_unigrams (defaults to false) if it needs the shared word-level unigrams basis computed first;
  • TextSearch.Tokenizer.generate! performing the actual token production;
  • optionally tokentag (defaults to nothing, i.e. untagged) for the single-character tag appended to each token when mark_token_type=true.

A new generator kind works with any TokenPipeline without further changes: the pipeline's stages act on whatever tokens a generator produces. This is also the right home for anything that changes how text becomes tokens – splitting getUserName into three words, keeping H2O whole – since a generator sees the word stream and may emit one token or several, while the pipeline is strictly per-token and data-driven.

source
TextSearch.Tokenizer.NormalizationConfig — Type
NormalizationConfig(;
    del_diac::Bool=true,
    del_dup::Bool=false,
    del_punc::Bool=false,
    group_num::Bool=true,
    group_url::Bool=true,
    group_usr::Bool=false,
    group_emo::Bool=false,
    lc::Bool=true,
    re_user::Regex=DEFAULT_RE_USER,
    re_url::Regex=DEFAULT_RE_URL,
    re_num::Regex=DEFAULT_RE_NUM,
    emojis::Set{Char}=DEFAULT_EMOJIS
)

Defines the text normalization stage of a TextConfig (see its normalization field): utf8 normalization, character removal, whitespace normalization, casing, etc. Consumed by normalize_text.

  • del_diac: indicates if diacritic symbols should be removed
  • del_dup: indicates if duplicate contiguous symbols must be replaced for a single symbol
  • del_punc: indicates if punctuaction symbols must be removed
  • group_num: indicates if numbers should be grouped _num
  • group_url: indicates if urls should be grouped as _url
  • group_usr: indicates if users (@usr) should be grouped as _usr
  • group_emo: indicates if emojis should be grouped as _emo
  • lc: indicates if the text should be normalized to lower case
  • re_user, re_url, re_num: the regexes used to detect @user mentions, URLs, and numbers when their corresponding group_* flag is set (see normalize_text). Override them to customize detection (e.g. for a different language or domain).
  • emojis: the set of emoji characters grouped when group_emo is set (see isemoji).

Example

julia> buff = Char[];

julia> normalize_text(TextConfig(normalization=NormalizationConfig()), "Café", buff);

julia> String(buff)
" cafe "
source
TextSearch.Tokenizer.QgramGenerator — Type
QgramGenerator(q)

Produces character q-grams (tagged 'q') over the whole normalized text, blanks included: they are not sub-words, since a window can span a word boundary ("o de" is a 4-gram of "todo de"), and the boundary blanks the normalizer adds make word starts and ends visible (" to", "do "). Runs of blanks count as a single blank, so layout does not produce q-grams of its own. Each field of a multi-field document is its own text, so no q-gram spans two fields. Texts shorter than q produce none.

Aimed at document-vs-document encodings (classification, clustering, dense encoders built on top) rather than short queries. Several lengths are several generators, and they combine with word tokens freely:

julia> tc = TextConfig(tokenization=TokenizationConfig(generators=[QgramGenerator(3)]));

julia> collect(tokenize(tc, "ab c"))
4-element Vector{String}:
 " ab	q"
 "ab 	q"
 "b c	q"
 " c 	q"

The tag is what keeps the 3-gram que (from porque) and the word que apart when both are generated, and what keeps word-level stopwords and lemmas from matching a q-gram; with mark_token_type=false they are the same token.

source
TextSearch.Tokenizer.TextConfig — Type
TextConfig(;
    normalization::NormalizationConfig=NormalizationConfig(),
    tokenization::TokenizationConfig=TokenizationConfig(),
    pipeline::TokenPipeline=TokenPipeline()
)

Defines a preprocessing and tokenization pipeline, composed of 3 independent stages:

  • normalization: a NormalizationConfig (utf8 normalization, character removal, whitespace normalization, casing, etc.).
  • tokenization: a TokenizationConfig (unigrams, word n-grams, and any extra custom AbstractTokenGenerators).
  • pipeline: a TokenPipeline of per-token stages applied to every generated token (lemma normalization or stopword removal).
  • language: which language this configuration is for, as an ISO 639-1 symbol (:es, :pt, :en, ...) or :unknown.

language currently changes nothing about tokenization – it is recorded, not acted on. It lives here rather than in a TextProfile because the language is not something a corpus produces, it is what selects the policy: whether to strip diacritics, whether suffix-anchored morphology fits, whether function words arrive as free tokens at all. Those are decisions the tokenizer will eventually make from this field, and a field the tokenizer must read belongs in the tokenizer's config.

It also earns its keep immediately: merge_profiles compares policy for equality, so declaring the language is what stops a Spanish profile from silently merging with a Portuguese one – their normalization and tokenization are identical, so nothing else distinguishes them. A detected language distribution ("this corpus turned out 87% Spanish") would be an observation about data, and would belong in a profile's lineage instead.

This is the corpus-independent half of a text model – it can be written by hand with no data. The artifacts a corpus produces (stopword sets, lemma maps, queryexpansion networks) live in a TextProfile, which materializes the pipeline from whichever of them it applies. Query-time expansion is likewise a profile-level decision (`applied.queryexpansion`), not a flag here: it is a search-time behaviour whose data does not live in the tokenizer.

Two ways to say the same thing

Any setting of the two sub-configs can be given directly, so the nesting is there when a whole sub-config is being passed around and absent when only one flag is being changed:

TextConfig(lc=false, del_diac=false)                                     # flat
TextConfig(normalization=NormalizationConfig(lc=false, del_diac=false))  # nested, identical

The flat form exists because the nested one is what actually gets written, over and over, for a single flag – and because nlist=[1] was spelled out in fifty places across this repository while already being the default. Naming a setting both ways is an error rather than one silently winning.

Example

julia> collect(tokenize(TextConfig(), "cats"))
["cats"]

julia> collect(tokenize(TextConfig(lc=false), "Cats"))
["Cats"]
source
TextSearch.Tokenizer.TokenPipeline — Type
TokenPipeline(; lemmas=nothing, stopwords=nothing)

The per-token stages of a TextConfig, as plain data in a fixed order rather than an open set of composable hooks.

lemmas      rewrite a token to its lemma      `Dict{String,String}`
stopwords   drop a token entirely             `Set{String}`

nothing means the stage does not run, and the order above is the order they run in.

Both stages apply to documents and queries alike, so one config serves fitting, indexing and searching. Orthographic bridging – letting a query typed leon reach León – deliberately does not live here: deciding it needs to know whether the typed token is in the vocabulary at all, which is not something a pipeline of plain data can answer. See resolve_query_tokens.

Why a fixed pipeline and not composable transformations

This replaced an AbstractTokenTransformation hierarchy (IgnoreStopwords, LemmaTransformation, ChainTransformation, plus a transform hook dispatching on both the transformation and the token generator). That design was right when TextSearch was more open, and by the time it was removed it held exactly two real stages, both of which are data rather than behaviour: a set to filter by and a map to rewrite through. A generic mechanism for two known things bought nothing and cost three specific problems.

Order. The stages are not commutative and the wrong order fails silently. With the stopword filter first, "las" is not in a set containing "la", survives the filter, and is only then rewritten to "la" – so the stopword lands in the vocabulary through the back door. This was documented backwards once and only measurement caught it. Here the order is in the code, once, and there is nowhere else to express it.

Per-type dispatch. merge_profiles compared transformations through a method per type, and the missing method for LemmaTransformation made it reject two profiles carrying identical lemma maps as incompatible. Comparing two TokenPipelines is comparing two fields and cannot have a missing method.

Cost. ChainTransformation's field was typed AbstractVector{<:AbstractTokenTransformation}, which is not concrete, so every step of every token went through a dynamic dispatch. Measured on 120,000 Spanish Wikipedia paragraphs (71.0M characters): lemmas+stopwords chained took 14.65s and 3.62 GB against 10.12s and 2.61 GB for the stopword filter alone – 45% more time and a gigabyte more garbage for one extra dictionary lookup per token.

Where extensibility lives now

Not here. A new kind of token – character q-grams, skip-grams, collocations, chemical formulas, splitting getUserName into three words – is a AbstractTokenGenerator in TokenizationConfig's generators list, which is the documented extension point and the right place for it: generators see the word stream, so they can emit one token or several.

What has no home here is an algorithmic per-token rewrite or filter that cannot be expressed as data – a stemmer, say, which is exactly what the removed Snowball extension was. Adding one means adding a named field with a documented position in the order, which is cheaper than the mechanism this replaced: that needed a new type, a transform_unigram method, a comparison method (the one that was forgotten), and a decision about where it chained.

Example

julia> p = TokenPipeline(lemmas=Dict("casas" => "casa", "rojas" => "roja"),
                         stopwords=Set(["la"]));

julia> cfg = TextConfig(tokenization=TokenizationConfig(nlist=[1]), pipeline=p);

julia> collect(tokenize(cfg, "las casas rojas"))
["casa", "roja"]

"las" becomes "la" and is then dropped, which is the order this type exists to guarantee.

source
TextSearch.Tokenizer.TokenizationConfig — Type
TokenizationConfig(;
    nlist::Vector=Int8[],
    mark_token_type::Bool=true,
    generators::Vector{<:AbstractTokenGenerator}=AbstractTokenGenerator[]
)

Defines the tokenization stage of a TextConfig (see its tokenization field): unigrams and word n-grams, computed from the output of the normalization stage.

  • nlist: a list of words n-grams to use (1 emits plain unigrams via UnigramGenerator, any other value emits word n-grams via NWordGenerator).
  • mark_token_type: each token is marked with its type (nword) when is true.
  • generators: extra AbstractTokenGenerators to run in addition to the ones nlist builds; this is the extension point for adding new kinds of tokens (e.g. character q-grams, skip-grams, or collocations, none of which are built-in anymore) without needing a new TokenizationConfig keyword argument (see alltokengenerators).

Note: If nlist and generators are both empty, then it defaults to nlist=[1]

Example

julia> cfg = TokenizationConfig(nlist=[1, 2]);

julia> collect(tokenize(TextConfig(tokenization=cfg), "cats sat"))
["cats", "sat", "cats sat	n"]
source
TextSearch.Tokenizer.TokenizedText — Type
TokenizedText(tokens::AbstractVector{String})

Wraps a list of string tokens for a single document. TokenizedText is TextSearch's universal contract for pre-tokenized text. It behaves like an AbstractVector{String} (supporting indexing, iteration, push!, append!, etc.) and is recognized by tokenize, Vocabulary, bagofwords, vectorize, and append_items! to bypass TextSearch's internal normalization and tokenization pipeline.

Use TokenizedText when integrating external tokenizers (e.g., WordTokenizers.jl, HuggingFace/subword tokenizers, spaCy, or custom tokenization functions).

Example

julia> using TextSearch

# Standard tokenization output:
julia> collect(tokenize(TextConfig(), "Hello world!!"))
3-element Vector{String}:
 "hello"
 "world"
 "!!"

# Integrating an external tokenizer:
julia> external_tokens = ["custom", "tokenization", "output"];

julia> tok_doc = TokenizedText(external_tokens);

julia> voc = Vocabulary(TextConfig(), [tok_doc]; verbose=false);

julia> vocsize(voc)
3
source
TextSearch.Tokenizer.TokenizerBuffer — Type
TokenizerBuffer(n=128)

Self-contained scratch space reused across tokenization calls to avoid reallocating on every call: normtext holds the normalized text, tokens accumulates the produced tokens, unigrams holds the word-level basis used by NWordGenerator (and any custom generator needing it), and io is scratch space for building individual token strings.

Tokenizer pools these internally for its own buffer-less convenience API (see tokenize); callers that need to hold a buffer across several calls (e.g. to safely alias its contents via borrowtokenizedtext) should borrow one from the same pool via tokenizerbuffer instead of constructing their own.

source
TextSearch.Tokenizer.UnigramGenerator — Type
UnigramGenerator()

Emits the word-level unigrams themselves as output tokens (untagged). Built from TokenizationConfig's nlist keyword argument when it contains 1. Every other generator that needs the word-level basis (see needs_unigrams) triggers the same underlying computation regardless of whether UnigramGenerator is present — this generator only controls whether the plain words also appear in the output.

source
TextSearch.LSI._lanczos_svd — Method
_lanczos_svd(A, k) -> Union{Nothing,Tuple}

Truncated SVD of A keeping the top k singular triplets via ARPACK's implicitly restarted Lanczos iteration (Arpack.svds): exact to working precision (measured ~3e-7 relative error on the singular values) while never forming a Gram matrix, which is what makes it both the accurate and the fast choice at scale.

ARPACK's own iteration is sequential and it is not re-entrant (unsynchronized static state, so it must not be called concurrently from multiple threads – LSI factorizes one batch at a time, so that is not a constraint here). It is not, however, serial in throughput: the heavy work goes to BLAS, so on a multicore host it does use many cores (~17 of 64 measured), just less effectively than a dense eigen, which is BLAS-3 rather than mostly BLAS-1/2.

Returns nothing when ARPACK cannot deliver k converged triplets – either by failing to converge or by throwing – so the caller can fall back to the exact dense path rather than abort a long fit.

source
TextSearch.LSI._query_expansion_localradius — Method
_query_expansion_localradius(voc, wordvecs, idx, ictx, kk, kcap, rank, q, mingroup)

Assembles a queryexpansion network whose per-token neighbor count is decided by the data instead of by a fixed k, via [`SimilaritySearch.bichromaticmetricjoin`](@ref) as a self-join.

A top-k network gives every token exactly k neighbors whether or not it has k real ones, so a token in a sparse region of the embedding gets filler and a token in a dense one gets truncated. The join instead estimates a cutoff radius per token from the reverse view of the same search: every token that ranked t among its own closest rank candidates votes for t with that distance, and t's cutoff is the q-quantile of its voters. Tokens with fewer than mingroup voters fall back to a pooled global cutoff.

Two consequences worth knowing before choosing this over :topk. The surviving pairs are those where the other token found t in its own top-kk, so the network becomes mutual-ish rather than a plain per-token top-k: a token that nobody's neighborhood reaches gets no query_expansion even if it has close ones of its own. And the output size is data-dependent, so kcap is applied only as a ceiling to keep a pathologically dense token from carrying thousands of neighbors.

source
TextSearch.LSI._select_topk — Method
_select_topk(nzind, nzval, topk::Integer) -> (nzind, nzval)

The topk largest-weighted entries of the parallel (nzind, nzval) arrays, restored to ascending index order afterward (so the result is still a valid sparse-vector index list, and iterating it twice gives the same order). Returns the inputs unchanged, not copied, when there are topk or fewer entries already.

source
TextSearch.LSI.query_expansion — Function
query_expansion(lsi::LatentSemanticIndexing, k::Integer=8;
         dist=Dist.Cosine(), normalize::Bool=true, verbose::Bool=true, approx=:auto,
         construction_recall::Real=0.97, search_recall::Real=0.9) -> Dict{String,Vector{Pair{String,Float32}}}

Builds a queryexpansion network from lsi's vocabulary embeddings (wordvectors); see the (voc, wordvecs, k) method above for the underlying algorithm and for what approx/ `constructionrecall/search_recallcontrol.normalizeis forwarded to [wordvectors`](@ref) before searching.

Example

net = query_expansion(lsi, 5)
net["dog"]   # ["dogs" => 0.02, "puppy" => 0.11, ...]
source
TextSearch.LSI.query_expansion — Function
query_expansion(voc::Vocabulary, wordvecs::AbstractDatabase, k::Integer=8;
         dist=Dist.Cosine(), verbose::Bool=true, approx=:auto,
         construction_recall::Real=0.97, search_recall::Real=0.9)
    -> (; query_expansion::Dict{String,Vector{String}}, distances::Dict{String,Vector{Float32}})

Builds a query_expansion network from voc's token embeddings in wordvecs (column t = embedding of gettoken(voc, t), e.g. from wordvectors or an externally supplied matrix): for every vocabulary token, finds its k nearest neighbors (by dist, cosine by default) among all other tokens' embeddings, via SimilaritySearch.allknn. The token itself is always excluded from its own neighbor list.

The two halves come back separately, as parallel per-token lists sorted by increasing distance (lower means more similar): query_expansion[tok] are the neighbor tokens in rank order, and distances[tok][i] is the distance to query_expansion[tok][i]. They are split because only the ranking participates in the normal query-expansion path – BM25 ignores the query side's weights entirely, and the distances stop being distances in any single space as soon as a network is merged or refitted. Keeping them apart lets a consumer (or a profile on disk) carry the ranking alone, which is where nearly all of a network's size lives.

Two arguments express what the network is for rather than filtering it for quality, which is the distinction that makes them work where eight quality filters did not (see the long note above this function):

  • head_df: a token whose document frequency exceeds this gets no list at all. No default is possible, because a document frequency is a ratio relative to whatever a document is: 0.05 means "in one paragraph in twenty" for a paragraph-level profile and something entirely different for an article-level one. Whoever knows the unit sets it; 0 disables. Measured on 272,466 Spanish Wikipedia paragraphs, 0.05 leaves 57 tokens without a list – the function words plus anos ano parte forma ciudad ser ha donde – at a cost of 0.1% of all pairs.
  • max_target_ratio: drop a neighbour whose document frequency exceeds the source's by more than this factor. This one is scale-invariant, being a ratio of two frequencies within the same corpus, so it carries a real default. planeta -> marte moves toward something rarer (0.29) and survives any value; a maximum-idf token pointing at por jumps 20,000x and does not. Known cost: it also cuts the legitimate rare -> common direction, such as a misspelling pointing at the correct word. 0 disables.

approx selects how the all-pairs search is done, and matters enormously on real vocabularies – an exhaustive search is O(vocabulary²):

  • :auto (default): approximate when length(wordvecs) > QUERY_EXPANSION_APPROX_THRESHOLD, exhaustive below it (where exhaustive is already fast and exact, so there is nothing to gain from approximating).
  • true: always approximate – build a SearchGraph, autotuning construction to MinRecall(construction_recall) and then the search parameters to MinRecall(search_recall).
  • false: always exhaustive, via ParallelExhaustiveSearch. Exact, and unusably slow past a few tens of thousands of tokens.

Example

net = query_expansion(voc, wordvectors(lsi), 5)
net.query_expansion["dog"]   # ["dogs", "puppy", ...]
net.distances["dog"]  # [0.02, 0.11, ...]
source
TextSearch.LSI.wordvectors — Method
wordvectors(lsi::LatentSemanticIndexing; normalize::Bool=true) -> MatrixDatabase{Matrix{Float32}}

Returns the LSI embedding of every vocabulary token, as a (outdim(lsi), vocsize(lsi)) matrix database – column t is the embedding of gettoken(lsi.model, t). This is exactly lsi.P (optionally column-normalized): a document's LSI vector (via vectorize/ vectorize_corpus) is a weighted sum of its tokens' columns of lsi.P, so these per-token vectors live in the same projected space and are directly comparable to each other and to document vectors (e.g. via Dist.Cosine()/Dist.NormCosine()). Set normalize=false to keep the raw (scaling-adjusted) lsi.P columns instead of unit-normalizing them.

Example

X = wordvectors(lsi)   # (outdim(lsi), vocsize(lsi)) MatrixDatabase
X[5]                   # the embedding of gettoken(lsi.model, 5)
source
TextSearch.vectorize! — Method
vectorize!(out::AbstractVector{Float32}, lsi::LatentSemanticIndexing, vec::SparseVectorLike; normalize::Bool=true, minweight::Real=1e-6, isnormalized::Bool=false, topk::Union{Nothing,Integer}=nothing)
vectorize!(out::AbstractVector{Float32}, lsi::LatentSemanticIndexing, text; normalize::Bool=true, minweight::Real=1e-6, isnormalized::Bool=false, topk::Union{Nothing,Integer}=nothing)

Projects a document (sparse vector or raw text) into the lower-dimensional dense LSI space in-place into out.

topk: restrict the projection to the topk heaviest tf-idf entries

topk=nothing (default) projects the full weighted vector, as before. Set topk to an Integer to project only its topk largest-weight entries (ties broken by index, so the result is deterministic) – a cheap ablation with a real, measured effect rather than a theoretical one:

Measured on a 6,175-article Spanish Wikipedia pilot (self near-duplicate retrieval between two disjoint paragraph ranges of the same article, in a separate exploratory sweep outside this repository): topk=4 scores recall@1 0.219 at outdim=64 against 0.182 for the full vector (no topk at all) – fewer, heavier tokens identify a specific document better than the whole weighted bag does. The effect reverses for recall@k at larger k (full vector 0.463 vs topk=4's 0.438): a wider candidate set is better served by more information, a single best guess by less.

Use a larger topk (or none) when indexing documents than when encoding a query at search time. A short query's own generic/template words (e.g. "capital", "government") can crowd out the one entity token that actually disambiguates it out of a small topk, while a document has more legitimate content to choose an anchor set from – measured on the same pilot's real-question evaluation, restricting a query to topk=4 gained almost nothing over the full vector (both near chance), unlike the clear win topk=4 gave on document-vs-document retrieval. There is no universal number this can default to (it trades off against outdim, vocabulary size, and document length), so nothing is enforced – topk is opt-in and symmetric by default (nothing on both sides), and choosing different values for indexing vs. querying is the caller's call to make deliberately.

source
TextSearch.vectorize — Method
vectorize(lsi::LatentSemanticIndexing, text_or_sparsevec; normalize::Bool=true, minweight::Real=1e-6, isnormalized::Bool=false, topk::Union{Nothing,Integer}=nothing)

Projects a raw text or sparse vector into the dense LSI space, returning a Vector{Float32} of length outdim(lsi). See vectorize! for what topk does.

source
TextSearch.vectorize_corpus — Method
vectorize_corpus(lsi::LatentSemanticIndexing, corpus;
                 normalize::Bool=true,
                 minweight::Real=1e-6,
                 isnormalized::Bool=false,
                 verbose::Bool=true,
                 topk::Union{Nothing,Integer}=nothing) -> MatrixDatabase{Matrix{Float32}}

Vectorizes every document in corpus into the dense LSI space in parallel across threads via @BATCHES, returning a MatrixDatabase of size (outdim(lsi), length(corpus)) ready for dense similarity search. See vectorize! for what topk does – typically a larger topk (or nothing) here, at indexing time, than at query time.

source
TextSearch.LSI.LatentSemanticIndexing — Type
LatentSemanticIndexing{M<:AbstractMatrix{Float32}, VM<:VectorModel} <: TextModel

Latent Semantic Indexing (LSI) model that projects sparse vector representations produced by a VectorModel into a lower-dimensional dense semantic space via Truncated Singular Value Decomposition (SVD).

Fields

  • model: The underlying VectorModel used to tokenize and weight text.
  • P: Dense projection matrix of size (k, m) where k = outdim and m = indim = vocsize(model).
  • s: Vector of singular values of length k.
  • k: Output dimension (k <= maxoutdim).
  • maxoutdim: Requested maximum output dimension (default: 128).
  • scaling: Scaling applied to singular vectors (:none, :inv_singular_values, :singular_values).
source
TextSearch.LSI.LatentSemanticIndexing — Method
LatentSemanticIndexing(corpus;
                       config::TextConfig=TextConfig(),
                       gw::GlobalWeighting=IdfWeighting(),
                       lw::LocalWeighting=TfWeighting(),
                       maxoutdim::Integer=128,
                       normalize::Bool=true,
                       minweight::Real=1e-6,
                       isnormalized::Bool=false,
                       verbose::Bool=true,
                       scaling::Symbol=:none)

Convenience constructor that builds an LSI model directly from a text corpus using default or provided TextConfig.

source
TextSearch.LSI.LatentSemanticIndexing — Method
LatentSemanticIndexing(config::TextConfig, corpus;
                       gw::GlobalWeighting=IdfWeighting(),
                       lw::LocalWeighting=TfWeighting(),
                       maxoutdim::Integer=128,
                       normalize::Bool=true,
                       minweight::Real=1e-6,
                       isnormalized::Bool=false,
                       verbose::Bool=true,
                       scaling::Symbol=:none)

Convenience constructor that builds a Vocabulary and VectorModel from config and corpus, then fits and returns a LatentSemanticIndexing model.

source
TextSearch.LSI.LatentSemanticIndexing — Method
LatentSemanticIndexing(model::VectorModel, corpus;
                       maxoutdim::Integer=128,
                       normalize::Bool=true,
                       minweight::Real=1e-6,
                       isnormalized::Bool=false,
                       verbose::Bool=true,
                       scaling::Symbol=:none,
                       factorization::Symbol=:auto)

Computes a Latent Semantic Indexing (LSI) projection matrix from corpus weighted by model. corpus can be a collection of raw texts or pre-vectorized sparse vectors (AbstractVector{<:SparseVectorLike} or AbstractDatabase).

Keyword Arguments

  • maxoutdim: Target embedding dimension (default: 128).
  • normalize: Whether to L2-normalize vectors during intermediate vectorization (default: true).
  • minweight: Threshold below which sparse vector weights are dropped (default: 1e-6).
  • isnormalized: Set to true if input texts are already normalized (default: false).
  • verbose: Whether to display progress bar during corpus vectorization (default: true).
  • scaling: Scaling factor applied to projection coordinates:
    • :none (default): standard orthogonal concept projection P = U_k^T.
    • :inv_singular_values: classical LSI document coordinate scaling P = Σk^{-1} Uk^T.
    • :singular_values: singular value weighted projection P = Σk Uk^T.
  • factorization: how the truncated SVD is computed, which decides whether a large corpus is tractable at all:
    • :auto (default): :full while min(vocsize, length(corpus)) is at most LSI_FULL_FACTORIZATION_MAX, :lanczos above it.
    • :lanczos: _lanczos_svd – ARPACK's restarted Lanczos iteration. Exact to working precision and the fastest option at scale; falls back to :full if ARPACK fails to converge.
    • :full: exact, via a dense Gram matrix and a complete eigen. Costs O(min(m,n)^3) time and min(m,n)^2 memory regardless of maxoutdim (it computes every eigenpair and keeps maxoutdim of them), so it is only appropriate for small corpora.

Both options are exact; the choice is purely about cost, so there is no accuracy knob to tune here.

source
TextSearch.LSI.LSI_FULL_FACTORIZATION_MAX — Constant
LSI_FULL_FACTORIZATION_MAX

Largest Gram-matrix side (min(vocsize, ndocs)) for which factorization=:auto still uses the exact dense :full path. Above it, :auto switches to :lanczos: measured on Spanish Wikipedia slices, :full wins below a couple of thousand documents (n=2000: 4.6s vs 14.8s) and loses badly above (n=8000: 48.8s vs 11.5s), since its cost grows with the cube of this side while ARPACK's is driven by the number of nonzeros.

source
SimilaritySearch.Projections.bitsketch — Method
bitsketch(ri::RandomIndexing, doc; minweight::Real=1e-6, isnormalized::Bool=false) -> Vector{UInt64}
vectorize(::Union{Type{BitSketch}, BitSketch}, ri::RandomIndexing, doc; minweight::Real=1e-6, isnormalized::Bool=false) -> Vector{UInt64}

Computes a SimHash-style binary bit sketch (packed into UInt64 words) from a document projected by RandomIndexing.

source
TextSearch.vectorize! — Method
vectorize!(out::AbstractVector{Float32}, ri::RandomIndexing, vec::SparseVectorLike; normalize::Bool=true, minweight::Real=1e-6, isnormalized::Bool=false)
vectorize!(out::AbstractVector{Float32}, ri::RandomIndexing, text; normalize::Bool=true, minweight::Real=1e-6, isnormalized::Bool=false)

Projects a document (sparse vector or raw text) into the lower-dimensional dense Random Indexing space in-place into out.

source
TextSearch.vectorize — Method
vectorize(m::Module, ri::RandomIndexing, text_or_sparsevec; kwargs...)

Projects a document into the Random Indexing space and quantizes it using SQu8 or SQgu8.

source
TextSearch.vectorize — Method
vectorize(ri::RandomIndexing, text_or_sparsevec; normalize::Bool=true, minweight::Real=1e-6, isnormalized::Bool=false) -> Vector{Float32}

Projects a raw text or sparse vector into the dense Random Indexing space, returning a Vector{Float32} of length outdim(ri).

source
TextSearch.vectorize_corpus — Method
vectorize_corpus(m::Module, ri::RandomIndexing, corpus; kwargs...)

Projects an entire corpus with Random Indexing and quantizes it with SQu8 or SQgu8.

source
TextSearch.vectorize_corpus — Method
vectorize_corpus(ri::RandomIndexing, corpus;
                 normalize::Bool=true,
                 minweight::Real=1e-6,
                 isnormalized::Bool=false,
                 verbose::Bool=true) -> MatrixDatabase{Matrix{Float32}}

Vectorizes every document in corpus into the dense Random Indexing space in parallel across threads via @BATCHES, returning a MatrixDatabase of size (outdim(ri), length(corpus)) ready for dense similarity search.

source
TextSearch.RI.RandomIndexing — Type
RandomIndexing(model::VectorModel, corpus=nothing;
               maxoutdim::Integer=1024,
               method::Symbol=:gaussian,
               rng::AbstractRNG=Random.default_rng())

Constructs a RandomIndexing model from a VectorModel.

Arguments

  • model: The vocabulary and weighting model.
  • corpus: Optional corpus parameter (ignored during construction, provided for API symmetry with LSI).

Keyword Arguments

  • maxoutdim: Target projection dimension (default: 1024).
  • method: Random projection algorithm:
    • :gaussian (default): Gaussian random projection matrix with unit-norm columns.
    • :qr: Orthonormal random projection matrix via QR factorization.
    • :sparse_random (or :ternary): Sparse ternary random projection (±1 with sparse support).
  • rng: Random number generator (default: Random.default_rng()).
source
TextSearch.RI.RandomIndexing — Type
RandomIndexing{M<:AbstractMatrix{Float32}, VM<:VectorModel} <: TextModel

Random Indexing (RI) model that projects sparse vector representations produced by a VectorModel into a lower-dimensional dense semantic space via random projections (SimilaritySearch.Projections), with default output dimension maxoutdim=1024.

Fields

  • model: The underlying VectorModel used to tokenize and weight text.
  • P: Projection matrix of size (k, m) where k = outdim and m = indim = vocsize(model).
  • k: Output dimension (k = maxoutdim).
  • maxoutdim: Target embedding dimension (default: 1024).
  • method: Random projection method used (:gaussian, :qr, :sparse_random).
source
TextSearch.RI.RandomIndexing — Method
RandomIndexing(corpus;
               config::TextConfig=TextConfig(),
               maxoutdim::Integer=1024,
               method::Symbol=:gaussian,
               gw::GlobalWeighting=IdfWeighting(),
               lw::LocalWeighting=TfWeighting(),
               verbose::Bool=true,
               rng::AbstractRNG=Random.default_rng())

Convenience constructor: creates a RandomIndexing model directly from raw corpus using default text configuration.

source
TextSearch.RI.RandomIndexing — Method
RandomIndexing(config::TextConfig, corpus;
               maxoutdim::Integer=1024,
               method::Symbol=:gaussian,
               gw::GlobalWeighting=IdfWeighting(),
               lw::LocalWeighting=TfWeighting(),
               minfreq::Integer=1,
               maxfreq::Integer=0,
               verbose::Bool=true,
               rng::AbstractRNG=Random.default_rng())

Convenience constructor: builds a Vocabulary and VectorModel from corpus using config, and then creates a RandomIndexing model.

source