The Problem
Finding reliable nutritional information for supermarket products is harder than it should be. Most data is scattered across individual product pages, presented inconsistently, and locked behind aggressive bot protection. There’s no open, queryable database of Tesco’s full food catalogue.
This project solves that.
How It Works
The scraper uses Playwright with Firefox to navigate Tesco’s product pages. Tesco uses Akamai bot protection which blocks most automated tools, but with the right browser configuration and session management, every product page yields structured data.
Each product is captured with:
- Identity: name, GTIN barcode, SKU, brand
- Price: in GBP with availability status
- Nutrition per 100g: energy, fat, saturates, carbohydrates, sugars, fibre, protein, salt
- Per-serving nutrition: with Reference Intake percentages
- Metadata: category hierarchy, star rating, review count, product image
The data is stored as JSON and rendered through this Astro-powered frontend. The site is deployed on Cloudflare Workers for global distribution with zero cold starts.
The Path Forward
- Scale the dataset — from 20 to 50, then 100, then 1,000+ products
- SQL database — migrate from flat JSON to a proper relational schema
- Price tracking — historical price data to detect trends, sales, and inflation
- Search and filtering — by category, brand, dietary requirement, nutritional profile
- API access — structured endpoints for third-party integration
This is an MVP. The foundation is solid — the rest is execution.