Dashboard
Products
4679
Export-ready
3386
Excluded
0
Merged away
0
Merge queue
0
Sites
24
How the pipeline works
Four idempotent stages move data through one S3 bucket (layers are key roots:
raw/ structured/ export/ export-images/ logs/) and Postgres:
- fetch โ crawls each site, saves raw responses to
raw/. Runs only on disposable scraper boxes, never from here, so our IPs aren't blacklisted. Their runs still show up under Runs. - parse โ turns a site's latest raw run into
structured/parquet (pure, no network). - merge โ normalizes INCI, dedupes, and upserts canonical products into Postgres. A scraped change that would overwrite a manual edit goes to the Merge queue instead.
- export โ writes qualifying products (name + full INCI + โฅ1 image) to
export/parquet and their images toexport-images/. Use Reuse existing images to skip re-downloading.
Editing a product's INCI fields auto-normalizes them and rebuilds its ingredient rows. Control matching (add/correct/delete overrides, then re-normalize) on the INCI page.
Recent jobs
All runs โ| # | Stage | Status | Run | Started |
|---|---|---|---|---|
| 13 | merge | succeeded | 20260721T054627Z | 2026-07-21 05:46:27 |
| 12 | parse | succeeded | 20260721T054556Z | 2026-07-21 05:45:56 |
| 11 | export | succeeded | 20260720T165724Z | 2026-07-20 16:57:24 |
| 10 | export | succeeded | 20260720T153849Z | 2026-07-20 15:38:49 |
| 9 | merge | succeeded | 20260720T153709Z | 2026-07-20 15:37:09 |
| 8 | parse | succeeded | 20260720T153639Z | 2026-07-20 15:36:39 |
| 7 | merge | succeeded | 20260720T120144Z | 2026-07-20 12:01:44 |
| 5 | merge | succeeded | 20260720T115003Z | 2026-07-20 11:50:03 |