Skip to main content

Search Architecture

This describes how SearchService (backend/src/v1/search/search.service.ts) actually resolves a query, and — more importantly — the constraints and past experiments behind why it works this way. Some of this exists only in past conversation/decision history, not in code comments or commit messages, so it's written down here for both human developers and AI agents working on search: getting these decisions wrong silently regresses ranking quality rather than throwing an error.

The three sources a query can hit​

  1. Postgres full-text search (SearchEngine.POSTGRES) — ts_rank over a to_tsvector('english', ...) index on searchableText. The original, still-default engine; a search with no inEngines/ninEngines given resolves to Postgres only, preserving pre-multi-engine behavior exactly.
  2. Elasticsearch full-text search (SearchEngine.ELASTICSEARCH) — plain multi_match across content.title^5, content.summary^2, searchableText, content.body, boosted by field specificity. The title^5 boost specifically exists because crawled pages share a large nav/sidebar boilerplate block in their body text, which otherwise makes a generic query match on some other page's nav mention rather than the actual target page (found in the search-tuning loop's first iteration, lab/alan/search-tuning/eval-history.jsonl).
  3. Vector search — one call per requested embeddingConfigId, each against its own embedding_chunks_{id} Elasticsearch index. One index per config, not one shared index, because dense_vector fields have a fixed dimension count per field and different embedding configs can use different models/dimensions. This is deliberately not a SearchEngine enum value — embedding_config rows are dynamic and user-created with no fixed enumeration, unlike Postgres/Elasticsearch — so it's an orthogonal embeddingConfigIds array on the request instead, combinable with any engine selection.

A query can hit any combination of these simultaneously. The fast path (exactly one text engine, no vector configs) skips merging entirely and returns that engine's results directly, unchanged from pre-hybrid-search behavior.

Merging: reciprocal rank fusion, not raw scores​

Once more than one ranked list is in play, results are merged by reciprocal rank fusion (RRF), not by comparing scores directly. This is necessary, not a stylistic choice: ts_rank, Elasticsearch's BM25 _score, and cosine similarity from a vector search are not on a comparable numeric scale, so sorting by "highest score wins" across sources would be meaningless. RRF instead scores purely by position within each list — weight / (RRF_K + rank), summed across every list a document appears in — which works uniformly regardless of how many lists are in play.

RRF_K = 60 is the standard default for this technique, chosen as a pragmatic starting point, not derived from this project's own data.

TEXT_ENGINE_RRF_WEIGHT = 1 and VECTOR_RRF_WEIGHT = 3 are derived from this project's own data — the search-tuning loop (lab/alan/search-tuning/, full results in eval-history.jsonl):

Vector weighthit@1hit@3
1 (unweighted hybrid)64.7%88.2%
276.5%94.1%
3 (current)82.4%94.1%
476.5%94.1% (regressed — an exact-term FAQ query the text engine used to win outright dropped to hit@2)

Weight 3 is a genuine local maximum, not noise: it's the point where vector similarity wins ties on semantic-mismatch queries (e.g. "backend coding standards" vs. a page titled "Backend Getting Started") while still leaving room for the text engine's own exact-term wins. If the eval set grows or the embedding model/chunking changes, re-run the eval loop and revisit — this weight is only as good as the query set it was tuned against.

Rejected approach: a vector similarity floor​

A hard cutoff (VECTOR_SIMILARITY_FLOOR = 0.8, filtering out low-confidence vector hits before merging) was tried and reverted (289d2543 → 9c3b5932, recorded in 24cd8da5). It did exactly what it was calibrated to do — but was a net regression (hit@1 77.4%/hit@3 86.8%, down from 79.2%/92.5% baseline). Filtering vector's contribution on weak-confidence queries didn't just remove false vector confidence, it removed vector's otherwise-adequate covering signal — which exposed that Elasticsearch's own multi_match ranking has real problems on generic or paraphrased queries that VECTOR_RRF_WEIGHT=3 had been masking, not just compensating for. One example: "can I use this commercially without paying" dropped from hit@1 to rank 5, with an unrelated Wikipedia article outranking the correct FAQ page on pure text relevance.

Lesson, not just history: a hard confidence cutoff on one signal can silently regress overall ranking by removing that signal's coverage value, even when the cutoff is doing precisely what it was tuned to do in isolation. The RRF weight approach (letting a low-confidence vector hit still contribute something, proportional to its rank) has held up better than a binary include/exclude cutoff. The underlying text-engine ranking weakness this experiment exposed has not been separately fixed — it's currently compensated for by the vector weight, not resolved.

Constraint: content.metadata is never a ranking signal​

Do not filter, score, or boost search results based on any content.metadata field (e.g. the category field the experiment-wikipedia-ingestion integration writes), in SearchService or any future ranking work.

Why: metadata is a free-form, per-integration bag with no shared schema across integrations — the crawler integration doesn't emit category at all, and nothing enforces that any future integration would. Building ranking logic on a field only some integrations populate would silently do nothing for every other source, so it cannot be a general-purpose scoring signal. metadata exists for other retrieval/display purposes (showing context to the user in the UI), not as ranking input.

Ranking/scoring changes should only use signals present on every result regardless of source: text match strength (ts_rank / Elasticsearch _score), vector similarity/rank, and RRF weight. A similarity floor or threshold is the kind of change that fits this constraint (see above for why that specific one was still rejected on the evidence); anything keyed off metadata.* does not fit it at all, independent of whether it would otherwise help.

Debug mode​

SearchReadManyInput.debug: true adds a debug field to the response containing the exact, already-built query sent to each engine/vector index (not a reconstruction) and a summary of what ran and how it was merged. Off by default — a vector search's raw debug request includes the full query embedding array, which isn't free to build or return, and there's no reason to pay that cost on the normal search path. There is no plan to expose debug in the frontend; it's an API-level tool, used by calling the API directly.

Where to re-tune this​

lab/alan/search-tuning/ holds the eval harness (run-eval.js, with a --config flag for comparing embedding models/weights) and the full historical record of every tuning iteration (eval-history.jsonl). Any change to RRF_K, the RRF weights, or the embedding model default should go through that loop and get recorded there — the weights above are only justified because they're backed by that data, not asserted from intuition.