
Sign up to save your podcasts
Or


In this episode of Lead Prompt, I dive into the engineering realities of scaling Greppr.org past 41 million indexed documents and 0.5TB of data. Moving past management theory, I break down the engineering optimizations required to survive the chaos of crawling the web at scale. Discover how I bypassed geo-location redirects using proxies, defeated crawler tar pits with custom fail-fast timeouts, implemented rapid early-stage filters for NSFW content and non-text rich media, and eliminated index bloat through automated canonical URL deduplication.
Show notes are here: https://leadprompt.sh/a/737-Lessons-from-crawling-the-web-at-scale-2026w17
Keywords:
greppr, web crawling, systems engineering, solo founder, apache solr, web scraping, backend development, data indexing, web crawler at scale, canonical urls, crawler tar pits, proxy routing, geo ip redirection, nsfw filtering, rich media filtering, solr architecture, software engineering podcast, lead prompt, tech infrastructure
By John CollinsIn this episode of Lead Prompt, I dive into the engineering realities of scaling Greppr.org past 41 million indexed documents and 0.5TB of data. Moving past management theory, I break down the engineering optimizations required to survive the chaos of crawling the web at scale. Discover how I bypassed geo-location redirects using proxies, defeated crawler tar pits with custom fail-fast timeouts, implemented rapid early-stage filters for NSFW content and non-text rich media, and eliminated index bloat through automated canonical URL deduplication.
Show notes are here: https://leadprompt.sh/a/737-Lessons-from-crawling-the-web-at-scale-2026w17
Keywords:
greppr, web crawling, systems engineering, solo founder, apache solr, web scraping, backend development, data indexing, web crawler at scale, canonical urls, crawler tar pits, proxy routing, geo ip redirection, nsfw filtering, rich media filtering, solr architecture, software engineering podcast, lead prompt, tech infrastructure