Ingestion
streaming
- feeds
- Encodertext
The same story kept arriving five times over. Now it arrives once.
Duplicate pair F1
0.94
Embedding ANN with calibrated rerank
Was
0.65
Exact hash + title match
No jargon in this section. The technical write-up is further down.
Syndicated articles came in as near-copies of each other — reworded, retitled, but the same story. Simple duplicate checks missed them entirely, and comparing every article against every other one was far too slow at the volume arriving.
We built a check that compares what two articles mean rather than how they are worded, fast enough to run on every item as it arrives instead of afterwards.
Complaints about a repetitive feed fell to almost nothing, storage and processing costs came down by about a fifth, and the editorial team stopped removing duplicates by hand.
Scroll through the stages. Anything marked as added is a component that did not exist before this project.
streaming
sbert
hnsw
cross-encoder
calibrated
streaming
sbert
We added thishnsw
We added thiscross-encoder
We added thiscalibrated
Dataset, approach, measured results and the stack. Written for whoever has to review it.
Hybrid retrieval and semantic scoring to remove duplicates before downstream pipelines.
Near-duplicate articles from syndication flooded the feed. Exact hashing missed rewrites, and naive pairwise comparison at ingestion volume was computationally impossible.
Next
Answer six questions and we will tell you whether this shape fits your problem — including when it does not.