Dashboard

Products
4679
Export-ready
3386
Excluded
0
Merged away
0
Merge queue
0
Sites
24
How the pipeline works

Four idempotent stages move data through one S3 bucket (layers are key roots: raw/ structured/ export/ export-images/ logs/) and Postgres:

  1. fetch โ€” crawls each site, saves raw responses to raw/. Runs only on disposable scraper boxes, never from here, so our IPs aren't blacklisted. Their runs still show up under Runs.
  2. parse โ€” turns a site's latest raw run into structured/ parquet (pure, no network).
  3. merge โ€” normalizes INCI, dedupes, and upserts canonical products into Postgres. A scraped change that would overwrite a manual edit goes to the Merge queue instead.
  4. export โ€” writes qualifying products (name + full INCI + โ‰ฅ1 image) to export/ parquet and their images to export-images/. Use Reuse existing images to skip re-downloading.

Editing a product's INCI fields auto-normalizes them and rebuilds its ingredient rows. Control matching (add/correct/delete overrides, then re-normalize) on the INCI page.

Recent jobs

All runs โ†’
#StageStatusRunStarted
13 merge succeeded 20260721T054627Z 2026-07-21 05:46:27
12 parse succeeded 20260721T054556Z 2026-07-21 05:45:56
11 export succeeded 20260720T165724Z 2026-07-20 16:57:24
10 export succeeded 20260720T153849Z 2026-07-20 15:38:49
9 merge succeeded 20260720T153709Z 2026-07-20 15:37:09
8 parse succeeded 20260720T153639Z 2026-07-20 15:36:39
7 merge succeeded 20260720T120144Z 2026-07-20 12:01:44
5 merge succeeded 20260720T115003Z 2026-07-20 11:50:03