Search Architecture¶
This document describes the architectural principles, domain model, and technical design of Badger Commerce's search and discovery platform. It is intended for engineers building on the platform, operators who want to understand what the system can express, and anyone evaluating future direction.
For the customer-facing feature surface, see Search and Discovery. For tenant isolation fundamentals that underpin search, see Multi-Tenancy Architecture.
Why a Search Platform¶
Most ecommerce platforms treat search as a bolt-on: a free-text input box that runs alongside Mongo-driven category pages, manual collections, and faceted filters wired into bespoke admin screens. The result is fragmentation — multiple authoring concepts, multiple data paths, and inevitable drift between how a category page works and how the search results page works.
Badger Commerce takes the opposite stance. Search is the platform's discovery substrate. Category pages, search results, virtual categories, megamenu queries, and ad-hoc landing pages all resolve through the same search layer over the same index. There is one model to learn, one place where ranking and faceting are configured, and one performance profile for everything customer-facing.
Core Concepts¶
Taxonomy Collection
(what the customer is (manually curated
selling — classification) product groupings)
│ │
│ ↑ existing primitives │
▼ ▼
┌────────────────────────────────────────────────┐
│ MerchandisingTree │
│ (how the customer sells — projection) │
│ │
│ MerchandisingNode MerchandisingNode │
│ │ │ │
│ NodeBacking NodeBacking │
│ (one of five) (one of five) │
└────────────────────────────────────────────────┘
│
▼
┌─────────────────────┐
│ Typesense index │
│ (per-tenant) │
└─────────────────────┘
Taxonomy¶
The taxonomy describes what the customer is selling: classes of product, the attributes those classes carry, and how attributes are validated. It is structural metadata that lives once per tenant and is the source of truth for product classification. See Taxonomy, TaxonomyLevel, and TaxonomyAttribute in commerce-core.
Collection¶
A Collection is a manually curated grouping of products, identified by a seoName and optionally arranged into a parent/child hierarchy via parentCollectionIDs. Products declare collection membership via Product.collectionIDs (a list of collection seoNames — the field name is a historical misnomer). Collections existed before the merchandising tree and remain in active use.
MerchandisingTree¶
A named, hierarchical projection of the catalogue that customers actually browse. A site has one or more trees; one is marked primary and drives canonical URLs and SEO. Other trees can power gift guides, seasonal navigation, or megamenus.
A tree is just an ordered set of root MerchandisingNodes. Each node has children. The tree decides nothing about products on its own — that's the backing's job.
MerchandisingNode¶
A node is the unit of merchandising. It carries display data (name, description, image, SEO metadata, facet configuration) and a NodeBacking that specifies how products resolve into it.
NodeBacking¶
The five backing types are the heart of the model. Every node is exactly one:
| Backing | Resolution |
|---|---|
ManualCollectionBacking |
Products in the named Collection (via Product.collectionIDs). |
TaxonomyProjectionBacking |
Products assigned to a taxonomy level (or any descendant), optionally narrowed by attribute filters. |
VirtualQueryBacking |
Products matching a VirtualQuery predicate — e.g. "orange clothes under £100". |
ManualListBacking |
A specific list of products in a specific order. Used for landing pages, gift guides, hand-curated promotions. |
GroupBacking |
No products of its own — a structural container that aggregates children for navigation and SEO. |
Mixing backings inside one tree is encouraged. A typical fashion retailer's primary tree might be a taxonomy projection at the top, with virtual queries pinned to season nodes, manual collections at the leaves, and a manual list for "Featured Today" on the homepage.
VirtualQuery¶
VirtualQuery is a small JSON-serialisable predicate DSL — not a raw query string. It supports clauses like AttributeEquals, AttributeIn, AttributeRange, PriceRange, InStock, InCollection, InTaxonomyLevel, TextMatch, and Nested(VirtualQuery), composed with AND / OR. The DSL is internal, validated, renderable in admin UI, and translated into Typesense filter_by expressions by VirtualQueryTranslator.
The Five Principles¶
These are the load-bearing rules. If you find yourself violating one, stop and check with the maintainers — there's almost certainly a better path.
1. Search is the source of truth for discovery; Mongo is the source of truth for state.¶
Anything customer-facing that browses, filters, faceted-navigates, or full-text-searches goes through Typesense. That includes classic category/collection pages — they are not a special case. Anything that mutates (write a product, update inventory, change a price) goes through Mongo first; the index follows.
This means the rendering layer has one query model and one performance profile. It also means we never debug "why does the category page disagree with the search results page" — they're the same thing.
2. One model, many backings.¶
Manual collections, taxonomy projections, virtual queries, manual lists, and groups are all the same node-shaped thing. They differ only in how products resolve into them. Admin UX, SEO, facet configuration, breadcrumbs, and indexing logic all operate on the abstract node — not on five parallel concepts.
This is what makes the model both expressive enough for large fashion retailers and approachable enough for a small charity shop. Start with one manual collection. Add a virtual query when the use case appears. The platform doesn't make you commit to a paradigm up front.
3. Multi-tenant by isolation, not by filter.¶
Each tenant gets its own Typesense collection, named products__{catalogueId}, fronted by a stable alias. Cross-tenant queries are impossible by construction; tenant deletion is one API call; per-tenant schema evolution is supported (different tenants can have wildly different facetable attribute sets without compromise).
We rejected the "single collection with siteId filter" approach despite its simplicity: it makes noisy-neighbour problems harder to isolate, forces the schema to a lowest-common-denominator, and weakens compliance posture. One Typesense cluster, many tenant collections, is the right balance for both hosted and SaaS deployments.
4. Events drive freshness; reconciliation guarantees correctness.¶
Mongo lifecycle events flow through ApplicationEventPublisher → RabbitMQ → worker → Typesense. Updates land in seconds. But events can be lost (network blips, broker restarts, deployments mid-batch), so a nightly reconciliation job re-asserts the world: it streams everything modified in the last 25 hours and upserts, then diffs Typesense IDs against Mongo IDs and deletes orphans.
Neither mechanism is sufficient alone. Events without reconciliation drift silently. Reconciliation without events makes the catalogue feel stale. Together they give us best-effort freshness with eventual correctness — the right trade-off for ecommerce.
5. Variants roll up to products in the index.¶
Each product is one document. Variant attributes (colours, sizes, SKU codes) are collapsed into multi-valued facet fields on the parent. A search for "red shirts" returns the shirt once, not once per red variant. SKU-level lookups are still possible via dedicated fields — but listings, facets, and dedup logic all flow naturally because the index reflects how shoppers think about products.
The exception we're not making: B2B catalogues sometimes need variant-level documents. We'll cross that bridge when a customer needs it; the schema is not painted into a corner.
Indexing Pipeline¶
Mongo write ──► ApplicationEventPublisher (in-process)
│
▼
SearchIndexQueueProducer
│
▼
RabbitMQ
(search.index.events)
│
▼
SearchIndexEventListener (worker)
│
▼
SearchIndexer ◄── MerchandisingResolver
│
▼
Typesense
Events¶
| Event | Triggered by | Effect |
|---|---|---|
ProductIndexEvent |
ProductService.save / delete |
Upsert or remove a single product doc. |
CollectionIndexEvent |
CollectionService.save / delete |
Reindex products in that collection (breadcrumb / collectionSeoNames refresh). |
MerchandisingNodeIndexEvent |
merch tree mutation | Recompute merchandisingNodeIds on affected products. |
TaxonomyIndexEvent |
active taxonomy change | Schema migration via alias swap, or full rebuild. |
PageIndexEvent |
page save (Mongo mapping event) / PageService.delete |
Upsert or remove the page in the tenant's content index — only for sites with Content Search. |
ContentReindexRequestedEvent |
stereotype save | Rebuild the tenant's content index. |
Schema migration via alias swap¶
When the facetable attribute set changes — typically because the active taxonomy gained or renamed an attribute — the platform doesn't mutate the live index. Instead, SchemaManager builds a fresh collection at the next version (products__{catalogueId}__v3), bulk-indexes everything, then atomically swaps the products__{catalogueId} alias to point at it. Zero downtime, zero half-migrated state, and rollback is one alias flip.
Reconciliation¶
A nightly @Scheduled job per worker node runs ReconciliationJob. For each site:
- Stream all enabled products with
lastModifiedDate > now − 25h. Upsert. - Page through Typesense and Mongo IDs in parallel; delete orphans.
- Emit metrics on drift count and duration.
The 25-hour lookback gives a 1-hour overlap with the previous run for safety. Drift count is a leading indicator of pipeline health: trending up means events are being lost.
Schema & Facets¶
Base fields (always present)¶
id, siteId, name, description, miniDescription, seoName, canonicalPath, imageUrls, price, comparePrice, onSale, inStock, manufacturer, rating, featured, publishedDate, displayOrder, collectionSeoNames, taxonomyLevelIds, merchandisingNodeIds, searchTerms.
taxonomyLevelIds is denormalized: a product assigned to "Womens > Outerwear > Coats" has all three level IDs in the array, so a query against any ancestor matches without recursive lookups.
merchandisingNodeIds is the magic that makes node-driven listings cheap. At index time, MerchandisingResolver walks every merch tree and asks "does this product resolve into this node?" The answer set is stored on the doc. Browsing a node is then a single equality filter — no per-request resolution.
Dynamic fields (per-site)¶
| Field | Type |
|---|---|
attributes (nested) |
One sub-field per facetable taxonomy attribute. Schema generated from the active taxonomy at collection-creation time. |
variant.colours, variant.sizes |
Rolled up from Variant.variantGroupValuesAndCode. |
variant.skuIds |
For SKU lookup. |
variant.priceRange |
When variants override the parent price. |
Each tenant gets exactly the facets it needs and nothing more. No lowest-common-denominator schema.
Facet configuration cascade¶
A node renders facets from a three-level fallback: node.facetConfig → tree.defaultFacetConfig → taxonomy-derived default (every facetable attribute, ordered by taxonomy display order). Most sites will set the tree default and rarely touch individual nodes.
REST API¶
The search layer is exposed to headless/storefront-API consumers through two read
endpoints (tag Search in the API reference). Both delegate to the same
MerchandisingQueryService.search(...) as the storefront pages, so their products, facet
counts, and pagination match the web /search and /c/** pages exactly.
GET /v1/search/products— catalogue-wide faceted search. Mirrors the storefront/searchpage. Parameters:q(free text),sort,filter,pageNumber,pageSize.GET /v1/merchandising/products?path=/c/...— products within a merchandising-tree node, resolved by canonical path. Mirrors a/c/**category page.404if no node exists at that path or it's hidden. Samesort/filter/paging parameters.
Facet selection uses a repeated filter parameter carrying facetId:value pairs — e.g.
?filter=colour:Black&filter=colour:Blue&filter=manufacturer:Nike. Values for the same
facet are ORed, different facets are ANDed. The facetId is the id of a FacetGroup
returned in the response, so a client round-trips a facet straight back as a filter. Which
facets are aggregated is resolved server-side via the facet configuration
cascade — there's no client control over the facet set.
Both responses are a ProductSearchResults: a page of products (data +
paginationParams), the facets aggregated over the matched set, and a query echo.
pageNumber is 1-based.
Authoring Model¶
Operators don't think about indexes, schemas, or events. They think about:
- "What's my catalogue shape?" → Define a taxonomy, classify products.
- "How do I want customers to browse?" → Build a primary merchandising tree by mixing collections, taxonomy projections, virtual queries, and curated lists.
- "What other entry points do I need?" → Add additional named trees for gift guides, megamenus, seasonal landing pages.
The platform handles index lifecycle, schema migrations, breadcrumb computation, alias swaps, drift reconciliation, and node resolution. None of these are exposed as authoring surface.
What This Is Not (Yet)¶
Round 1 of the search platform delivers the foundation: domain model, persistence, admin scaffolding, and the indexing pipeline. The following are deliberately deferred:
- Rendering off search —
CollectionControllerstill hits Mongo today. Round 2 rewires it. - URL & SEO strategy — canonical paths, redirects from old
/collection/{seoName}to new/c/{path}, sitemap generation, hreflang. - Personalisation and ranking signals — popularity, recency, user-history boosts. Foundation supports it; we haven't built it.
- A/B testing of merchandising trees — the multi-tree model exists partly to enable this later.
- SaaS Sett authoring tooling — collaborative editing, drafts, scheduled publishing.
- Migration tooling — automatic conversion of existing Collection hierarchies into a default merchandising tree. Existing collections continue to work via
ManualCollectionBacking; explicit migration is a future, opt-in tool.
Glossary¶
- Tree —
MerchandisingTree. A named hierarchical projection of the catalogue. A site has one or more. - Primary tree — the tree marked
primaryfor a site. Drives canonical URLs and SEO. - Node —
MerchandisingNode. A position in a tree with display metadata and a backing. - Backing — the strategy by which a node resolves products. One of five types (manual collection, taxonomy projection, virtual query, manual list, group).
- Virtual category — a node backed by a
VirtualQuery. Authored as a saved filter, behaves to the customer as just another category. - Projection — a node backed by a taxonomy level + optional attribute filters. The merchandising tree's view of a slice of the taxonomy.
- Reconciliation — the nightly job that re-asserts Typesense state against Mongo. Catches drift from missed events.
- Alias swap — zero-downtime reindex pattern: build a new versioned Typesense collection, then atomically repoint the stable alias.
- Roll-up — variant attributes collapsed onto the parent product doc, so listings dedupe and facets count by product.
Content index¶
Sites that opt into Content Search get a second per-tenant collection,
pages__{catalogueId}, holding one document per CMS page / blog post (title, headings, body text,
tags). It shares the client, queue and reconciliation machinery above but none of the product schema,
so product search is unaffected whether it's on or off.
Related Documentation¶
-
Content Search — optional search over pages and blog posts.
-
Search and Discovery — customer-facing search features.
- Multi-Tenancy Architecture — tenant isolation principles.
- Product Catalog Management — products, variants, collections.
- Extension System — how rendering extensions consume the search layer.