which documents in a store are retrieval noise, scored from the vectors already in the file
No labels, no model, no second pass over the corpus. The features are document-level aggregates of quantities the chunk store already holds. The blend of them scores 0.988 AUC against hand-marked junk.
The vector legs answer “what is this like”. These answer “is this worth retrieving at all”. A cookie banner is a near-duplicate of every other page’s, sits far from the corpus mean, and is made of words that are everywhere.
def chunk_matrix( store:str='store', # the chunk table limit:int=None, # sample this many chunks instead of using all of them seed:int=0, dtype:type=float16, # the width the store holds)->tuple:
(ids, doc_ids, texts, V): every chunk’s stored vector, L2-normalised, as one array.
Nearest neighbours, centroids and the rest of the arithmetic
Inverse document frequency over chunks: boilerplate is made of words that are everywhere.
The features
Nine of them, all oriented larger-is-noisier. promiscuity is the only one needing anything outside the store. It counts how many separate questions found a document, which a caller passes in as seen if it keeps a log. Without one the feature is zero and the others carry the score.
def noise_scores( db, # an open Database store:str='store', # the chunk table weights:dict=None, # per-feature blend; None -> `NOISE_W` ranker:NoneType=None, # a fitted `Ranker` instead of a fixed blend prefix:str=None, # the tree prefix, for the docs table; None -> from `store`**kw)->L: # forwarded to `noise_features`
Every document, most suspicious first, with the features that put it there.
def noise_features( db, # an open Database store:str='store', # the chunk table seen:dict=None, # doc_id -> the set of queries it was retrieved for, if logged k:int=10, # neighbours per chunk topics:int=64, # topic centroids for the spread features tau:float=0.9, # similarity floor for "this is a duplicate" temp:float=8.0, # softmax temperature for chunk-to-topic spread limit:int=None, # sample this many chunks instead of using all of them exact_max:int=30000, seed:int=0, dtype:type=float16, # the width the store holds)->AttrDict:
Per-document noise features, computed from the vectors already in the file.
The ranker
A logistic model over standardised features, fitted by IRLS with an L2 penalty. Two callers use it. One scores noise over the features above. The other re-ranks retrieval features pairwise. Both are the same twelve lines of arithmetic.
Pairwise linear learning-to-rank over the feedback log.
Tests
from fastcore.test import test_eq, test_closefrom litesearch import Index, hash_embedclass HashEmbed:"`hash_embed` behind an `.encode`, so these tests need no model and no network."def__init__(self, dims=256): self.dims = dimsdef encode(self, xs, **kw): return hash_embed(list(xs), ndim=self.dims)