Content Search¶
Content search makes a site's CMS pages and blog posts searchable from the same storefront search
box as its products. Shoppers typing "returns" or "care guide" see matching pages and articles in the
as-you-type dropdown, and the /search results page gains a Pages & articles tab with highlighted
snippets and tag filtering.
It is an optional, per-site service: off by default, switched on per site, and built on the same Typesense cluster as product search (see Search Architecture) but with its own index, so turning it on never changes product search or ranking.
Turning it on¶
- Make sure the site has Typesense configured (it already does if product search works).
- Set Enable Content Search (
contentSearchEnabled, search-config group) totruein the site configuration admin (/admin/shopSettings/guided-config), or with the MCPmanageSiteConfigtool. - Run Reindex search (Merchandising > Trees) once to index the site's existing pages. From then on pages are indexed as they're saved, and the nightly reconciliation keeps the index honest.
Turning it off hides content results immediately (the storefront checks the flag on every query). The site's content index is left in place, so switching it back on needs only a reindex.
What shoppers see¶
- Search box dropdown — a Pages & articles group below products and categories: the page title, a page/article icon, and the matching passage with the searched words in bold. Configured per search-bar instance (Show page & article suggestions, Max page & article suggestions, default on / 3); it only ever appears on sites with content search enabled. See Search Bar & Autocomplete.
/searchtabs — Products (n) and Pages & articles (n). The content tab (/search?q=…&tab=content) lists results with type (Page / Article), publish date and author for articles, the best-matching snippet, tags, and the article's masthead image as a thumbnail.- Tags — chips above the results show the tags across the match set with counts; clicking one
filters to it (
contentTag=…), clicking it again clears it. Each result's own tags link the same way, so/search?tab=content&contentTag=Guidesbrowses a topic with no query at all (newest first). - Nothing in the catalogue? If a search finds no products but does find content,
/searchopens on the Pages & articles tab rather than showing an empty products page.
Every customer-facing theme is covered: Nova and the themes layered on it (Brock, Depot, Pop, Atelier,
Spec) share nova/fragment/content-results.html; the Bootstrap family uses
bootstrap/fragment/content-results.html.
What gets indexed¶
One document per page, in a per-tenant Typesense collection pages__{catalogueId}.
| Field | From |
|---|---|
title |
The page name (falling back to its first heading) — weighted highest |
headings |
Markdown # headings, HTML <h1>–<h6>, JSON-component headings, accordion titles and tab labels, the hero's title slot, FAQ title and questions |
tags |
The page's tags (also a facet) |
description |
The page description, else its metaDescription attribute |
body |
Everything else the page shows as text (capped at 50,000 characters) |
type |
article for blog posts (blogPost stereotype), else page |
url |
The page's canonical path, else /p/{seoName} |
imageUrl |
The blog masthead image |
author, publishedDate |
attributes.author; the page's creation date |
Text comes from the extensions the page actually shows — its own enabled placements plus those
inherited from its stereotype, merged exactly as rendering merges them
(ExtensionService.findEffectivePlacements). It is read from each extension's stored configuration,
not from a render: indexing runs in the worker, which has no templates or request, and a render would
also pull in navigation, personalisation and analytics.
| Extension | Indexed |
|---|---|
markdownFragment |
Markdown source, syntax stripped; # lines as headings |
htmlFragment |
HTML with scripts/styles/comments removed; <h*> as headings |
jsonComponent, hero |
Literal text, quote/author, image alt props; headings as above. Data-bound props ({"$ref": …}) and button/link labels are skipped |
faqExtension |
Title and questions as headings, answers as body |
blogMastheadExtension |
Its image becomes the result thumbnail |
Product grids, forms and other functional blocks contribute nothing — they aren't content. To make a
new content extension searchable, add a @Component implementing ContentTextExtractor
(commerce-core, search/content).
Left out: disabled pages; pages ticked Exclude from Site Search (page editor, next to Exclude
from Sitemap); global system pages (catalogue ALL); platform pages (search, basket, checkout,
account, login); and pages with no text and no description.
How it stays fresh¶
| Trigger | Effect |
|---|---|
Any page save (admin, MCP, import, extension edits — anything through MongoTemplate) |
PageSearchIndexListener → PageIndexEvent → the page is re-extracted and upserted, or removed if it no longer qualifies |
PageService.delete |
The page's document is removed |
| Stereotype save | StereotypeSearchIndexListener → ContentReindexRequestedEvent → the site's content index is rebuilt, since every page using the stereotype may have changed |
| Reindex search (Merchandising > Trees) | Rebuilds products and content |
| Nightly reconciliation | Rebuilds content for every enabled site and prunes documents whose page is gone (including deletes that bypassed the service) |
All of these travel the existing search.index.* RabbitMQ route to the worker's ContentIndexer,
and every one is a no-op for a site without content search — the decision lives in one place,
ContentSearchSettings, read by both the worker and the storefront.
Ranking¶
Queries search title, headings, tags, description, body with weights 8, 5, 4, 2, 1, so a page
titled "Returns" beats a page that mentions returns in passing. Ties, and tag browsing with no query,
sort newest first. Snippets come from Typesense highlighting of the body (else the description) and are
returned as plain-text runs. The storefront escapes them and wraps only the matched words in
<mark>/<strong>, so nothing from the index is ever rendered as HTML.
Not yet¶
- A REST endpoint (
/v1/search/content) for headless storefronts. It will be added spec-first toapi.yamlalongside/v1/search/products. - Indexing products' own long-form content, or collections/merchandising-node descriptions.
- Synonyms and per-site ranking tweaks.