Manual competitor price checks cap out at a few hundred products and a lag measured in weeks — and the harder problem is matching, since the same product carries a different name on every site. We built a pipeline that parses, visually matches and classifies competitor catalogues: more than fifty competitors and over one hundred thousand products, refreshed daily.
A retailer competing on price needs to know what its competitors charge today, not what they charged when someone last checked. Done manually, that means staff opening competitor sites, finding the equivalent product, recording a price, and repeating — which caps coverage at a few hundred items and introduces a lag measured in weeks. The harder problem is not collection but matching: the same product appears under different names, different photographs and different specification formats on every site, so a naive name-based comparison produces a table nobody trusts. The retailer needed price and assortment intelligence across a wide competitive set, refreshed frequently enough to act on, with matching reliable enough that category managers would base decisions on it.
Per-competitor parsers extract catalogue structure, product attributes and prices, isolated so that a layout change on one site degrades one source rather than the whole pipeline. Runs are queued and retried independently.
A dedicated worker normalises product imagery — fetching, converting and storing to S3 — decoupling the slow, bandwidth-heavy image work from parsing so neither stage blocks the other.
Visual matching resolves the core problem that names cannot: identifying that two differently-titled listings are the same physical product. This is what makes the resulting price comparison defensible to the category managers who act on it.
Matched products are tagged and classified into the retailer's own category structure, so comparisons happen along the dimensions the business actually manages rather than along competitors' taxonomies.
The front end where category managers work: price positioning by category, assortment gaps, and movement over time. The pipeline's output is only useful at the point where someone can query it without asking an analyst.
The retailer tracks more than fifty competitors across more than one hundred thousand products with a daily refresh — coverage and latency that manual monitoring cannot reach at any realistic headcount. Because matching is visual rather than name-based, category managers work from comparisons they trust, which is the difference between a report that informs pricing and one that gets ignored. The system is in production and under active development, with the most recent work extending the vision and tagging services.
Tell us what the process looks like today and we will tell you what can be automated — and what should not be.
LET'S TALK