Notes from running pipelines in production.
Not marketing posts. What we have actually learned keeping web data correct over hundreds of consecutive delivery days, including the parts that argue against hiring us.
The gate every batch passes before it leaves us.
- Type conformanceEvery field matches its declared type
- Range and sanityValues fall inside agreed bounds
- Null-rate thresholdsCoverage flagged when it drops
- Duplicate resolutionResolved against your chosen key
- Schema driftSource structure changes caught early
- Volume varianceUnexpected record counts halt delivery
How to specify a schema so the first delivery is actually usable
Most first deliveries disappoint for the same reason, and it is never the extraction. Nobody wrote down what the fields meant. Four things each one needs.
Build vs buy: what an in-house scraping team actually costs
The prototype costs an afternoon, which is why build-vs-buy decisions get made on the wrong number. Here is the rest of the invoice, and the cases where building really is correct.
What actually breaks in a scraper running 600 consecutive days
Writing a scraper is easy. Keeping one correct for two years is a different job entirely, and the failures that cost most are the ones that never throw an error.
More as we have something worth saying. If you want a specific question answered here, ask us.
Rather see it than read about it?
Send us three sources. You'll get a feasibility read, a proposed schema, and a real sample extracted from your own targets, before any commitment.
No retainer required to find out whether your sources are feasible.