Search Architecture
This describes how SearchService (backend/src/v1/search/search.service.ts)
actually resolves a query, and — more importantly — the constraints and past
experiments behind why it works this way. Some of this exists only in past
conversation/decision history, not in code comments or commit messages, so
it's written down here for both human developers and AI agents working on
search: getting these decisions wrong silently regresses ranking quality
rather than throwing an error.
The three sources a query can hit
- Postgres full-text search (
SearchEngine.POSTGRES) —ts_rankover ato_tsvector('english', ...)index onsearchableText. The original, still-default engine; a search with noinEngines/ninEnginesgiven resolves to Postgres only, preserving pre-multi-engine behavior exactly. - Elasticsearch full-text search (
SearchEngine.ELASTICSEARCH) — plainmulti_matchacrosscontent.title^5,content.summary^2,searchableText,content.body, boosted by field specificity. Thetitle^5boost specifically exists because crawled pages share a large nav/sidebar boilerplate block in their body text, which otherwise makes a generic query match on some other page's nav mention rather than the actual target page (found in the search-tuning loop's first iteration,lab/alan/search-tuning/eval-history.jsonl). - Vector search — one call per requested
embeddingConfigId, each against its ownembedding_chunks_{id}Elasticsearch index. One index per config, not one shared index, becausedense_vectorfields have a fixed dimension count per field and different embedding configs can use different models/dimensions. This is deliberately not aSearchEngineenum value —embedding_configrows are dynamic and user-created with no fixed enumeration, unlike Postgres/Elasticsearch — so it's an orthogonalembeddingConfigIdsarray on the request instead, combinable with any engine selection.
A query can hit any combination of these simultaneously. The fast path (exactly one text engine, no vector configs) skips merging entirely and returns that engine's results directly, unchanged from pre-hybrid-search behavior.
Merging: reciprocal rank fusion, not raw scores
Once more than one ranked list is in play, results are merged by
reciprocal rank fusion (RRF), not by comparing scores directly. This
is necessary, not a stylistic choice: ts_rank, Elasticsearch's BM25
_score, and cosine similarity from a vector search are not on a
comparable numeric scale, so sorting by "highest score wins" across sources
would be meaningless. RRF instead scores purely by position within each
list — weight / (RRF_K + rank), summed across every list a document
appears in — which works uniformly regardless of how many lists are in play.
RRF_K = 60 is the standard default for this technique, chosen as a
pragmatic starting point, not derived from this project's own data.
TEXT_ENGINE_RRF_WEIGHT = 1 and VECTOR_RRF_WEIGHT = 3 are derived
from this project's own data — the search-tuning loop
(lab/alan/search-tuning/, full results in eval-history.jsonl):
| Vector weight | hit@1 | hit@3 |
|---|---|---|
| 1 (unweighted hybrid) | 64.7% | 88.2% |
| 2 | 76.5% | 94.1% |
| 3 (current) | 82.4% | 94.1% |
| 4 | 76.5% | 94.1% (regressed — an exact-term FAQ query the text engine used to win outright dropped to hit@2) |
Weight 3 is a genuine local maximum, not noise: it's the point where vector similarity wins ties on semantic-mismatch queries (e.g. "backend coding standards" vs. a page titled "Backend Getting Started") while still leaving room for the text engine's own exact-term wins. If the eval set grows or the embedding model/chunking changes, re-run the eval loop and revisit — this weight is only as good as the query set it was tuned against.
Rejected approach: a vector similarity floor
A hard cutoff (VECTOR_SIMILARITY_FLOOR = 0.8, filtering out low-confidence
vector hits before merging) was tried and reverted
(289d2543 → 9c3b5932, recorded in 24cd8da5). It did exactly what it
was calibrated to do — but was a net regression (hit@1 77.4%/hit@3 86.8%,
down from 79.2%/92.5% baseline). Filtering vector's contribution on
weak-confidence queries didn't just remove false vector confidence, it
removed vector's otherwise-adequate covering signal — which exposed that
Elasticsearch's own multi_match ranking has real problems on generic or
paraphrased queries that VECTOR_RRF_WEIGHT=3 had been masking, not just
compensating for. One example: "can I use this commercially without
paying" dropped from hit@1 to rank 5, with an unrelated Wikipedia article
outranking the correct FAQ page on pure text relevance.
Lesson, not just history: a hard confidence cutoff on one signal can silently regress overall ranking by removing that signal's coverage value, even when the cutoff is doing precisely what it was tuned to do in isolation. The RRF weight approach (letting a low-confidence vector hit still contribute something, proportional to its rank) has held up better than a binary include/exclude cutoff. The underlying text-engine ranking weakness this experiment exposed has not been separately fixed — it's currently compensated for by the vector weight, not resolved.
Constraint: content.metadata is never a ranking signal
Do not filter, score, or boost search results based on any
content.metadata field (e.g. the category field the
experiment-wikipedia-ingestion integration writes), in SearchService or
any future ranking work.
Why: metadata is a free-form, per-integration bag with no shared
schema across integrations — the crawler integration doesn't emit
category at all, and nothing enforces that any future integration would.
Building ranking logic on a field only some integrations populate would
silently do nothing for every other source, so it cannot be a
general-purpose scoring signal. metadata exists for other
retrieval/display purposes (showing context to the user in the UI), not as
ranking input.
Ranking/scoring changes should only use signals present on every result
regardless of source: text match strength (ts_rank / Elasticsearch
_score), vector similarity/rank, and RRF weight. A similarity floor or
threshold is the kind of change that fits this constraint (see above for
why that specific one was still rejected on the evidence); anything keyed
off metadata.* does not fit it at all, independent of whether it would
otherwise help.
Debug mode
SearchReadManyInput.debug: true adds a debug field to the response
containing the exact, already-built query sent to each engine/vector index
(not a reconstruction) and a summary of what ran and how it was merged.
Off by default — a vector search's raw debug request includes the full
query embedding array, which isn't free to build or return, and there's no
reason to pay that cost on the normal search path. There is no plan to
expose debug in the frontend; it's an API-level tool, used by calling the
API directly.
Where to re-tune this
lab/alan/search-tuning/ holds the eval harness (run-eval.js, with a
--config flag for comparing embedding models/weights) and the full
historical record of every tuning iteration (eval-history.jsonl). Any
change to RRF_K, the RRF weights, or the embedding model default should
go through that loop and get recorded there — the weights above are only
justified because they're backed by that data, not asserted from
intuition.