How these projects actually run.
Each write-up follows the same structure we'd use on your project: what the team was stuck on, how we scoped it, what we built, and what landed in the end.
The pipelines the write-ups below are drawn from.
- 0 errorsEvent ticketingDaily · Parquet → SFTP
- 0 errorsRestaurant deliveryWeekly · Parquet → SFTP
- 0 errorsElectronic componentsMonthly · TSV in, ZIP out
- 0 errorsQuick-service restaurantsScheduled · Parquet → SFTP
Examples of work we’ve delivered. Figures come from the delivery manifests behind each pipeline. Clients are anonymised by sector — we name one only with their agreement.
628 million pricing records across 205 unbroken delivery days
The problem
Event and listing prices move constantly, and a headline price is not what a buyer pays: fees and tax land at checkout. The team needed the full price picture, every day, at a volume that made manual collection impossible.
What we built
- Daily event snapshots plus ad-hoc listing pulls against a fixed pre-market window
- Full fee breakdown captured per listing: display price, buyer fee, other fees, sales tax, and all-in checkout price
- 20-field schema keyed on listing ID, with venue, section, row, and seat detail retained
- Delivered as Parquet over SFTP, with byte-for-byte verification on arrival
The outcome
205 consecutive delivery days between January and August 2026, 475 files, and zero failed transfers. 1.87 million listing rows compress to 1.6 MB, so a day of data moves in seconds.
Engagement shape
Store-level menu pricing and review data across a national footprint
The problem
Menu prices vary store by store on delivery platforms, so a single national price list tells you almost nothing. Coverage had to be resolved per location, and paired with ratings and review data for the same stores.
What we built
- Per-store collection rather than a single national vantage point, using geolocation-aware requests
- Menu pricing and ratings/reviews split into separate delivery streams against a shared store key
- Normalisation of item naming and modifier pricing so stores stay comparable
- Weekly delivery as Parquet over SFTP
The outcome
79 verified files delivered with no transfer errors, giving per-store pricing that can be compared across a national footprint rather than averaged into uselessness.
Engagement shape
Manufacturer part lists resolved to live availability and pricing
The problem
The team held long lists of manufacturer part numbers but no reliable way to resolve them to current availability and price. Lookups were manual, inconsistent, and out of date by the time they were compiled.
What we built
- Manufacturer part lists accepted as TSV input, one file per monthly cycle
- Each part resolved against live search, with availability and pricing captured
- Unresolvable parts reported explicitly rather than dropped silently
- Results returned as compressed archives on a monthly cadence
The outcome
A repeatable monthly cycle that turns a static part list into current market data, with every input file accounted for in the delivery manifest.
What we'll publish, and what we won't.
Named only with permission
A client's name appears here when they've agreed to it in writing. Nothing gets published because it would look good on our site.
Numbers we can evidence
Any figure in a case study comes from a delivery report we can produce. We won't round a result upward because it reads better.
Including what went wrong
Real projects have sources that had to be dropped and estimates that moved. Write-ups that mention none of that aren't worth reading.
What working with us is actually like.
We'd rather show you how we operate than fill this space with quotes. Every point below is something you can hold us to from the first conversation.
You talk to the people building it
No account layer between you and the engineers. The person who scoped your schema is the person who fixes it when a source changes.
We tell you when something won't work
If a source is unreliable, a field isn't consistently available, or a request isn't something we'll take on, you hear it during scoping, not after an invoice.
Small on purpose
A four-person team means every project gets senior attention. We take on work we can do properly rather than filling a pipeline.
The maintenance is the commitment
Anyone can deliver a dataset once. We stay responsible for it: monitoring sources, catching drift, and repairing extraction before it reaches you.
Your project would get the same treatment.
Tell us the sources and the fields you need. We'll come back with a feasibility read, a proposed schema, and a real sample.
No retainer required to find out whether your sources are feasible.