Skip to main content

Detailed Roadmap

This page expands on the simplified Roadmap with the reasoning, specifics, and open questions behind each item - grouped the same way the simplified roadmap is grouped (V1, V2, V3). Nothing here is a separate plan; it's the "why" and "how" behind what the simplified roadmap says will happen.


V1 details - search and homelab deployment​

Why Elasticsearch moved from V2 into V1​

The original roadmap scoped V1 around PostgreSQL full-text search only, with Elasticsearch deferred to V2's "Search Improvements." In practice, a self-hosted homelab user judges the product by whether search actually works well - a worse search engine with a cleaner roadmap sequence isn't a better product. Elasticsearch, hybrid text+vector search, and the ranking-quality work below all shipped as part of V1 instead.

Search architecture, as it exists today​

  • Hybrid search: Elasticsearch full-text relevance and vector similarity are merged via reciprocal rank fusion (RRF), not compared on raw score (BM25 and cosine similarity aren't on a comparable scale). The vector side is weighted more heavily than the text side (empirically tuned, not a default) - text alone under-uses semantic matches, vector alone loses exact-term precision.
  • Explicit Elasticsearch mapping: the content index used to be dynamically mapped, which meant it silently inherited Elasticsearch's default standard analyzer - no stopword filtering at all. At small scale this didn't matter; once the corpus grew large and topically diverse, stray common words (a, the, to) in a boosted title field could occasionally outrank a real match, because they were rare enough within that specific field to get inflated relevance weight. The fix was an explicit mapping using the english analyzer (stopwords + stemming), plus keyword-only typing for ID fields (never full-text searched, so no reason to analyze them) and leaving free-form per-integration metadata unindexed entirely.
  • Metadata is never a ranking input. Content metadata (e.g. a category field one integration happens to populate) is free-form and has no shared schema across integrations - using it for filtering or scoring would only work for whichever integration happens to populate that specific field, silently doing nothing for every other source. Metadata is for display only.
  • A hand-rolled Elasticsearch migration system, mirroring the existing TypeORM/Postgres migration system as closely as the two systems' capabilities allow: timestamped migration files with up()/down(), an applied-state ledger index, the same AUTO_MIGRATION gate at boot, and the same manual production CI job pattern. The one real asymmetry from Postgres: an Elasticsearch mapping change on a live index isn't reversible in place, so a mapping migration creates a new physical index and swaps an alias rather than mutating anything - down() means "point the alias back," not "undo the mapping."
  • Embedding model choice is not fixed. Multiple self-hosted (Ollama) and comparable models were benchmarked head-to-head against a real eval query set: nomic-embed-text (the original default), qwen3-embedding:8b, qwen3-embedding:0.6b, mxbai-embed-large, and bge-m3. qwen3-embedding:8b was chosen as the new default because it was the only model of the five that beat the nomic-embed-text baseline on both hit@1 and hit@3 at once (84.9% / 96.2%, later re-confirmed at 86.8% / 96.2%) - not because it scored highest on every individual metric. qwen3-embedding:0.6b actually had the single best hit@3 of anything tested (a perfect 100%), at roughly a seventh the size, but tied the baseline rather than beating it on hit@1 (83.0%) - worth keeping in mind as a much cheaper alternative if resource cost matters more than top-1 precision. This benchmarking approach - real eval set, side-by-side comparison, not a single theoretical pick - is the intended process for any future embedding or ranking change, not a one-time exercise.

Vector search is optional and configurable​

This is a direct consequence of the homelab deployment goal below: running an embedding model is real resource cost, and a Raspberry-Pi-class self-hoster shouldn't be forced into it just to get working search. Elasticsearch text search alone (with the mapping/analyzer fix above) is the baseline; vector search and hybrid ranking are an opt-in improvement on top, for anyone who can and wants to run an embedding provider (self-hosted via Ollama, or a hosted API).

Homelab deployment​

A single, self-contained docker-compose.yaml - not a Terraform-style infra/ module - with everything needed to run Yew Search on someone's own hardware: Postgres, Redis, Elasticsearch, the backend, and the frontend. This is deliberately distinct from any one person's actual homelab configuration (which mixes in unrelated services); the goal is a file someone with no prior Yew-specific infrastructure can bring up as-is.

Two things this surfaces that don't exist yet and need solving as part of building it, not after:

  1. Elasticsearch portability. A real running homelab's Elasticsearch setup has host-specific baggage (auth configuration, index namespacing to avoid collisions with unrelated services on the same cluster) that a shippable compose file can't assume. This needs to be genuinely self-contained, not a copy of one person's live setup.
  2. Whether vector search ships enabled by default. Given the point above, the honest default for a fresh homelab deployment is probably Elasticsearch text search only, with vector search documented as an opt-in step (add an Ollama service, or point at a hosted embedding provider) rather than assumed.

V2 details - permissions and the unified auth model​

Why permissions come before the MCP server, not alongside it​

The MCP server (V3) is meant to be usable by API keys, not just logged-in human sessions - both by Yew's own conversational assistant and by external AI tools a company connects. Building that auth story before real permissions exist would mean inventing a throwaway authorization model now and migrating everything later. Sequencing it after V2 avoids that.

API keys are first-class principals, not scoped delegations​

The initial framing considered here was an API key as a narrower delegation of whatever user created it (the GitHub personal-access-token pattern - a key can do less than its owner, never more). What's actually planned is different and simpler: an API key is assigned groups and permissions directly, the same mechanism used for a user, independent of who created it. There's no "is this narrower than its creator" relationship to validate, because there's no inheritance relationship at all - closer to how AWS IAM treats users and service principals as peers under the same policy-attachment mechanism, rather than one being a constrained view of the other.

The unified auth-principal shape​

Two authentication paths, one authorization shape:

  1. Cookie present → look up the User → resolve their permission object.
  2. API key present → look up the ApiKey → resolve its permission object (same shape as a user's).

Auth middleware normalizes both into one standard object before any downstream code runs. Every authorization check in the codebase gets written once, against that normalized shape, with no branching on whether the caller authenticated via session or key. This is what makes "the MCP server's auth is just another consumer of the permission system" true in practice, not just true in theory with if (isApiKey) checks scattered through every service.

The exact shape of the permission object itself is not yet designed - role-based access control (owner/admin/member/viewer, per the simplified roadmap's V2 section) is the starting assumption, but the concrete shape is deliberately left open until V2 work actually begins.


V3 details - MCP server and the conversational assistant​

The MCP server is scoped to knowledge retrieval, not controlling Yew​

This is a narrower, more defensible scope than "an MCP server for Yew" might suggest, and it's a deliberate choice: every tool it exposes is read-only and about getting knowledge out (search, fetch a document, list what sources exist), never about managing integrations, users, or configuration. That distinction matters against this repo's own current operating instructions, which say plainly "do not introduce MCP servers yet" - this roadmap is the reason that changes: a read-only retrieval surface is a fundamentally smaller blast radius than a general control-plane, and building real permissions first (V2) is what makes doing this safely, rather than casually, possible. The operating instructions themselves should get revisited once V2 lands, not worked around silently before then.

Sketched tool surface (not finalized): search, getDocument, listSources. The expectation is that this ends up being a thin protocol adapter over retrieval capability that already exists (the same search/content-read logic the REST API uses today), not a new retrieval engine - an agent doing "more complicated searches across knowledge" doesn't need fundamentally different primitives than a human typing into the search box, it needs to call the same search primitive multiple times in a reasoning loop, refining its query as it goes.

Two consumers, one surface​

  1. Yew's own conversational assistant - built as an MCP client against Yew's own server, not a separate bespoke retrieval implementation. This is a deliberate dogfooding choice: one well-tested tool surface improves both consumers at once, instead of maintaining two retrieval code paths that can drift apart.
  2. External AI assistants from other products - for companies that want to plug their existing AI tooling (their own Claude/ChatGPT deployment, or other agents) directly into their Yew knowledge base, without needing Yew's own UI at all. This is what makes the MCP server valuable independent of whether anyone uses Yew's search UI or conversational assistant - it's a standalone value delivery mechanism, not just plumbing for the assistant.

The conversational assistant needs grounding, not just retrieval​

A chat interface over search results is not enough on its own for something positioned as a trustworthy knowledge engine - if the assistant can produce an answer that isn't traceable back to a real source document, wrong or hallucinated answers undermine the whole premise. Citations back to source documents are a requirement of the feature, not a nice-to-have polish item added later.

Summarization is the smallest of the three new capabilities​

Search result summarization reuses the exact provider-routing pattern already built for embedding configs (LangChain adapters over OpenAI, Ollama, Voyage, Google GenAI, and Mistral AI - see build-embeddings.ts; note Anthropic has no native embeddings API, so it isn't one of the routed providers there). The equivalent for a completion/chat model is architecturally the same shape, just routing to a different LangChain capability - and a chat-model router would be the natural place for Anthropic to actually show up, unlike the embeddings one. This is why it's realistic to ship ahead of the MCP server and assistant, even though all three are grouped under the same "knowledge engine capabilities" heading - it doesn't have the same permissions dependency the other two have, since it doesn't need a new auth surface at all.


V5 details - turnkey identity and infrastructure-as-code​

Why this is V5, not folded into V4​

V4's SSO/SAML and SCIM items were originally scoped as V4 Enterprise features, but they're identity/provisioning concerns, not infrastructure/compliance ones - V4 is genuinely about dedicated infrastructure and compliance certification. Moved here because V5 has an actual dependency reasoning V4 didn't: a Terraform provider needs the V2 permission API to be stable through real usage first (V2 through V4), not still actively changing shape - building a provider on a churning API means constant breakage. SSO and SCIM don't strictly share that dependency, but they're the same customer profile and usually adopted together, so grouping them here tells a coherent "turnkey" story instead of three disconnected bullet points.

SSO and SCIM are two different things, usually bought together​

SSO (SAML 2.0/OIDC) only answers "how does someone log in" - it doesn't create accounts or manage group membership. SCIM is the separate protocol that lets an IdP push user creation, deactivation, and group membership into Yew automatically, mapping IdP groups to Yew teams. Enterprises adopt both together almost universally, but they're distinct integrations, not one feature - worth keeping them named separately so neither gets silently assumed to cover the other.

SSO doesn't conflict with the existing "cookie sessions only, never JWT" decision (docs/docs/backend/authorization.md, chosen for instant remote logout) - it only changes how the initial login handshake happens. The session that results afterward is the same cookie-based session as any other login.

Terraform provider​

A thin plugin over Yew's own REST API (companies, teams, folders, integration assignments, permission grants), the same pattern already proven by products like Grafana - not a novel piece of infrastructure to invent. The one design implication worth carrying into V2 API work now, even though the provider itself is a V5 deliverable: resources need a stable, human-meaningful identifier (a slug, not just a UUID) for terraform import and idempotent re-apply to work the way Terraform users expect.

ACL sync - a real fork, not a default​

Whether Yew should mirror a connected source's own ACLs (the way Glean does - inherit, don't own) rather than only using Yew's own permission model, for integrations where the underlying source genuinely has its own access control (e.g. a domain-wide Google Workspace connection). This is explicitly not a committed V5 feature - it's the point at which it becomes worth reconsidering, because the customer sophisticated enough to want Terraform-managed provisioning and SCIM sync is the same customer who'd actually have a source-system ACL worth syncing from. See docs/docs/backend/permissions.md for the full tradeoff.


Open questions, not yet decided​

These are named explicitly so they don't get silently assumed one way or the other later:

  • Whether summarization, the MCP server, and the conversational assistant are self-hosted-available or business-tier-only. The simplified roadmap currently lists them under the Business tier with an explicit "TBD" - MCP in particular has an argument for being available to self-hosters too, since a big part of its value (letting a company's own AI tooling plug into their knowledge) doesn't require Yew's SaaS infrastructure at all.
  • Whether "V1 done" requires search + summarization only, with the MCP server and assistant as a clearly-separate fast-following phase, or whether all of the knowledge-engine capabilities need to land together before calling that phase complete.
  • The final MCP tool surface (sketched above, not committed).
  • The concrete shape of the permission object referenced in V2 - now drafted in docs/docs/backend/permissions.md, still marked proposed.
  • Whether ACL sync (V5) gets built at all, versus staying a permanently rejected idea in favor of Yew always owning its own permission model.
  • The exact scope of the V5 Terraform provider - which resources it covers on day one versus later.