Site Data Transfer¶
BadgerCommerce provides built-in export and import functionality for transferring site configuration and catalogue data between environments. This feature is essential for copying data from production to development, creating site templates, or migrating between installations.
Key Features¶
Data Export¶
- Background Export: Exports run asynchronously in the worker, not on the web request. The archive is streamed straight off a Mongo cursor to a temp file and uploaded to S3, so even very large sites export with a flat, bounded memory footprint (no out-of-memory failures).
- Queue → poll → download: Requesting an export creates an
ExportJob; the admin UI polls its status and downloads via a short-lived pre-signed S3 URL once it reportsCOMPLETED. - Selective Export: Only exports catalogue and configuration data - excludes PII and transactional data
- Password Protected: Requires admin password confirmation for security
- Portable Format: Standard ZIP file with JSON data, easily inspectable and version-controllable
Data Import¶
- Validation Preview: See what will be imported before committing
- Conflict Resolution: Automatically replaces existing data in the target site
- ID Remapping: All MongoDB IDs and cross-references are automatically remapped
- Tenant Isolation: Imported data is fully owned by the target site
Privacy by Design¶
The export explicitly excludes all personally identifiable information (PII) and transactional data:
| Exported (Safe) | Excluded (PII/Transactional) |
|---|---|
| Products & Variants | Orders |
| Collections | Users/Customers |
| Pages (CMS) | Donations |
| Taxonomies | Gift Aid Declarations |
| Menus | Gift Cards |
| Media metadata | Subscriptions/Invoices |
| Delivery Options | Fundraiser Requests |
| Promotions (automatic and code-entered) | Analytics/Page Views |
| Stereotypes | AI Conversations |
| Variant Groups (Colour, Size…) | Sensitive config values |
| Site Text overrides | Sales counts |
| Stock on hand per SKU | |
| Site Configuration* |
*Site configuration excludes sensitive values like API keys, tokens, and secrets. It also leaves
behind keys that describe the source site itself: its Badger Sett organisation id (which routes
subscription webhooks and federation to every site carrying it) and the legal-entity sign-off
(legal-entity-confirmed-at / -by), so the target site's owner confirms their own details.
Storefront theme: the manifest records the source site's theme (sourceThemeName) and the import
switches the target site to it, provided that theme is installed on the target database. If it isn't
(say, a bespoke theme), the target keeps its current theme and the validation preview warns about it.
Not carried: the rest of the site (tenant) document — logo, social accounts, product image path, domains — and merchandising trees. The target site keeps its own.
Export File Structure¶
site-export-{siteId}-{timestamp}.zip
├── manifest.json # Export metadata (version, timestamp, source site and theme)
├── catalogue/
│ ├── products.json # All products with variants
│ ├── collections.json # Product collections/categories
│ ├── media.json # Media metadata (URLs, not binaries)
│ ├── variant-groups.json # Colour / Size / … definitions the variants refer to
│ └── stock.json # Stock on hand per SKU and warehouse
├── content/
│ ├── pages.json # CMS pages
│ └── menus.json # Navigation menus
├── taxonomy/
│ └── taxonomies.json # Product taxonomies
└── configuration/
├── site-configurations.json # Site settings (non-sensitive)
├── stereotypes.json # Product/page stereotypes
├── delivery-options.json # Shipping methods
├── site-text.json # Site Text overrides (Admin > Shop Settings > Site Text)
└── promotions.json # Promotions and discounts (Jackson default typing)
How ID Remapping Works¶
When importing data from Site A into Site B, all identifiers are automatically remapped:
| Field | On Export | On Import |
|---|---|---|
catalogueId |
Source site's ID | Replaced with target site's ID |
siteId |
Source site's ID | Replaced with target site's ID |
MongoDB _id |
Original IDs | New ObjectIds generated |
| Cross-references | Original IDs | Remapped to new IDs |
Some references are remapped specially: a product's images (media ids), any extension's mediaId
setting (the blog masthead image, which blog cards and og:image also use) on products, collections,
pages and stereotypes, and the site's active taxonomy (taxonomy-active-id, which names a
taxonomy by id). Collections, pages and other items that arrive
without an id each get their own new one.
This ensures: - No ID collisions between sites - Complete data ownership by the target site - No lingering references to the source site
Usage¶
Accessing the Feature¶
- Navigate to System > Tenants
- Click on the tenant you want to export from or import into
- Find the Data Transfer panel on the right side
Exporting Site Data¶
- Click Export Site Data
- Enter your admin password when prompted
- The export is queued to the worker and the dialog shows progress while it generates in the background. Large catalogues may take a little while.
- When it completes, the download starts automatically (via a pre-signed S3 link). The file is
named
site-export-{siteId}-{timestamp}.zip
The export job is processed by the worker, so you can navigate away — the generated archive is stored in S3 and the download link remains valid for 24 hours.
Importing Site Data¶
- Click Import Site Data and select a ZIP file
- Enter your admin password when prompted
- Review the validation summary showing:
- Source site information
- Entity counts to be imported
- Warnings about data replacement
- Click Import Now to proceed
- Wait for the import to complete
Once the import succeeds, a full search reindex is queued for the site (products, and pages and blog posts for content search), since the imported documents were never indexed one by one.
Use Cases¶
Development Environment Setup¶
Export production catalogue data to set up a realistic development environment without exposing customer data.
Site Templates¶
Create a "template site" with products, pages, and configuration, then export it for use as a starting point for new sites.
Staging/QA Testing¶
Copy catalogue data to a staging environment to test changes with real product data.
Disaster Recovery¶
Maintain offline backups of site configuration and catalogue data (note: for full backups including transactional data, use the S3 archive feature).
Multi-Environment Deployment¶
Promote catalogue changes from development through staging to production by exporting and importing.
Technical Details¶
Service Layer¶
The feature is implemented in the SiteTransferService interface with these operations:
public interface SiteTransferService {
// Streams the ZIP straight into the provided stream off a Mongo cursor — driven by the worker,
// which writes to a temp file and uploads it to S3. Memory-bounded regardless of catalogue size.
Map<String, Long> writeSiteExport(SiteContext siteContext, OutputStream out) throws IOException;
ImportValidationResult validateImport(InputStream zipFile);
SiteImportResult importSiteData(SiteContext targetSite, InputStream zipFile);
}
Both directions are streaming and memory-bounded:
- Export — the worker (
ExportProcessingJob) picks up a pendingSITE_TRANSFER_ZIPjob, callswriteSiteExportto stream each entity collection (via a Mongo cursor + JacksonSequenceWriter) into a temp ZIP, then uploads it to S3 withRequestBody.fromFile. - Import — the uploaded ZIP is spooled to a temp file and each entry is deserialised with a
Jackson
MappingIterator, persisting in fixed-size batches. ID-remap tables hold only string pairs, so peak heap stays flat even for large catalogues.
Sensitive Data Filtering¶
Configuration keys are automatically filtered if they:
- Are marked with ofSecretString() in ConfigKeyDefaults
- Contain patterns like: apikey, apisecret, token, secret, password, credential, stripe, recaptcha
Import Order¶
Data is imported in dependency order to maintain referential integrity:
- Media (referenced by products)
- Taxonomies (referenced by products)
- Collections (referenced by products)
- Products
- Pages
- Menus
- Stereotypes
- Delivery Options
- Promotions
- Site Configuration
Limitations¶
- Media files: Only metadata is exported, not the actual binary files. Media URLs are preserved but files must exist at those URLs.
- Product images:
@DBRefimage references are cleared during import; re-associate images in admin if needed. - Global promotions: Only site-specific promotions are exported; global promotions are excluded.
- User data: No user accounts are transferred; users must be created separately on the target site.
Security Considerations¶
- Export/import operations require admin authentication
- Password confirmation is required for both operations
- Sensitive configuration values are automatically stripped
- All PII and transactional data is excluded by design
- Audit logging captures export/import operations
Related Features¶
- Multi-Tenancy Architecture - Understanding tenant isolation
- Site Customization - Configuring site settings
- Content Management - Managing pages and menus