Data
We get you the data, clean, on a schedule.
Scrapers, feeds and pipelines that keep working after launch. We keep a client's high-volume collection system reachable in production against sources built to block it. That discipline transfers to pipelines of any size.
- Projects
- 28
- Figures published
- 9
- Free first step
- Data Map
- Fixed price from
- $1.5K
7 written up on the work page
listed on the work page
built from the market you need data from, plus a sample of what you have
Scraper and Data-Feed Rescue
Is this you
The week this fixes.
Sentences we hear from the people who end up sending this practice a bottleneck, and what each one is costing while it stays true.
“The scraper broke again and nobody noticed for a week.”
Decisions made on stale data, and a silent failure found three days later in the numbers.
“Their price changed Tuesday and nobody noticed until next week.”
Spend already out the door against the wrong price.
“The analyst starts in March. The data is a mess now.”
The first months of that seat spent on plumbing rather than on analysis.
What a project covers
Scoped to the smallest useful system, judged against a number.
- Scraper and data-feed rescue: take over the failing pipeline, stabilize it, keep clean data moving, with an optional monthly run
- Collection pipelines with storage, scheduling, and monitoring designed in from the start
- Sync into the systems that use the data: warehouses, dashboards, ERPs, storefronts
What that has looked like
- Competitor pricing collected continuously, on a high-volume system that stays reachable against sources built to block it
- Product data collected, normalized, and synced to your store on a schedule
- Market data collected continuously instead of exported by hand
The measures we agree before building
One of these is picked with you at the scoping step, and reported against.
- Feed uptime and how stale the data gets before anyone notices
- Coverage: how much of the target set you actually capture
- Manual interventions per week to keep the pipeline running
No charge, no call
4 free ways to start.
Each is a document or a sample, written by the person who would do the work. Send one input, get something back you can act on or ignore.
Free
Data Map
Where the data would come from, how often, what breaks first, and what running it costs a month.
You send: The market you need data from, plus a sample of whatever you have now.
You get back:
- The sources worth collecting, and the ones that are not worth the trouble
- A collection frequency, and what goes stale between runs
- The failure that will happen first, and what it costs to survive it
Free
Diagnosis of one broken source
Pick the feed that keeps breaking. We tell you why it breaks and what it would take to stop.
You send: The source, plus whatever error output or logs you have.
You get back:
- What is actually failing: blocks, layout changes, rate limits, or the schedule
- Whether it is worth stabilising or worth replacing
- What a rescue would involve, with the price band it falls in
Free
One-week competitor sample
Name three competitors. We collect them for a week and send you what changed.
You send: Three competitor names or URLs, and what you care about: prices, availability, catalog, or copy.
You get back:
- A week of collected records for the three, in a file you keep
- What moved during the week, and what did not
- What a continuous version would cost to set up and run
Free
Thousand-row cleanup sample
Send a thousand rows of the messy file. Get them back deduped, corrected, and filled in.
You send: A thousand rows of the list, catalog, or contact file that is not usable as it stands.
You get back:
- The same thousand rows, cleaned, with a change log beside them
- The rules we applied, so you can argue with any of them
- What the rest of the file would cost at the same standard
Every free first step across the five practices is listed together, if the problem sits between two of them.
Entry points
The ways in, each priced before work starts.
Start with one. None of them requires the next one, and each is small enough to judge on its own.
Scraper and Data-Feed Rescue
Take over the pipeline that keeps breaking and keep clean data moving: blocks, captchas, layout changes, retries, alerts.
$1.5K–$3K · Fixed price · optional monthly run · 1 to 3 weeks
Competitor and Market Monitor
Know what competitors changed before your team notices: prices, availability, catalog, or copy, collected on a schedule.
$1.5K–$3K · Setup · optional monthly run · 1 to 3 weeks
Reporting Layer
One dashboard your team actually opens, built off the data you already have rather than off a new system.
$2K–$5K · Fixed price · 2 to 4 weeks
Data Cleanup and Enrichment
The list, catalog or contact file deduped, corrected and filled in. Once, properly, with the rules written down.
$1K–$3K · One-off · 1 to 2 weeks
Every price, fixed or by the hour, is agreed before work starts and built the same way: the hours, the tools, and a share of company cost. See how a price is built
The free first step
What the first move looks like.
Examples of what we sent back after a bottleneck came in.
One place to read a restaurant's point of sale, social, and web analyticsWhat happened: Discovery only. No build.
A restaurant group wanted its point-of-sale data, social channels, and web analytics readable in one place, so marketing could be decided from the numbers rather than from instinct.
What it covered
- A diagram of the data flow from the point of sale through social and web analytics into one store
- What an AI layer over that store could answer, and what it could not
- The discovery scope, and the smallest channel test that would prove the value first
- The first move
- A channel test before the platform.
Treasury forecast automation for a logistics groupWhat happened: Discovery only. No build.
A family-run logistics group built its weekly cash-flow forecast by hand from two reports out of its freight system, with stakeholders in two countries.
What it covered
- An intent summary agreed before the workshops
- Discovery workshops on the forecast as it was assembled, report by report
- A bilingual non-disclosure agreement, and a discovery output the group could take to a build
- The first move
- Automate the two report pulls before touching the forecast model.
The work
What this looks like shipped.
A high-volume hotel pricing-data system
169 properties, same day
Scheduled collection, retries and monitoring that keep a travel-market data system running at more than eight million live requests a month, without adding headcount to maintain the feed.
A retail catalog pipeline that keeps a Shopify store current
Scrapes Costco product data on a schedule, normalizes and stores it, and syncs price and stock to Shopify.
Read the case study
A self-hosted email verification engine built for explainable results
34 reason codes · 25 output columns
An internal product: a bulk verification CLI with DNS and multi-stage SMTP checks that explains each classification instead of emitting a black-box score, deployed on its own verification node and run on real lead lists.
Read the case study
A batch product-data API for marketplace catalogs
Takes a batch of marketplace product URLs, up to 100 per request, and returns each product's barcode, images, price, and description as a CSV attachment, with explicit success, partial, and failure contracts per item.
Public-records search and retrieval, automated
Drives the Hamilton County clerk-of-court foreclosure search with Selenium and retrieves each matching PDF, retrying failed downloads: a job otherwise done by hand, one case at a time.
County-scale property and cadastral data extraction
Extracts Montana cadastral and property records from the cadastral API and its HTML pages into CSV, with a notebook analysing extraction performance.
A cold-outreach engine we ran on our own pipeline first
247 qualified leads
Contact sourcing from Apollo, verification, warmed secondary sending domains, sequenced copy, reply triage, and weekly deliverability reporting. We ran it at scale on our own program through the second half of 2025, then for a partner under their brand, as a pilot for a client, and on hourly contracts.
Read the case study
6 more data projects
What else we have built in this area, and what changed as a result.
Engagements in this catalogue, described without numbers and, in most cases, without the client: they were built under confidentiality or under another firm's brand. For numbers with sources behind them, see the work page.
A supplier feed that changed shape every quarter
Suppliers sent the same data in their own formats and quietly changed them. Every change broke the import, and it was usually a person downstream who noticed first.
What we built
- A per-supplier mapping layer, so a format change is a config edit rather than a code change
- Schema validation at the door, rejecting a bad file instead of half-importing it
- Alerting that names the supplier and the field that moved
- A quarantine area, so a rejected file can be inspected rather than lost
What changed
A supplier changing their export stopped being an outage, and the team found out from an alert rather than from a customer.
One number, instead of three departments disagreeing
Finance, sales, and operations each reported a different figure for the same month, and every meeting started by arguing about which was right.
What we built
- Extraction from each source system on a schedule
- A modelled layer where the definitions are written down and applied once
- Reconciliation checks that fail loudly when two sources disagree
- Dashboards built on the modelled layer, so nobody reports off a raw export
What changed
The definition argument moved out of the meeting and into a documented model everyone could read.
Inheriting a scraper that had stopped working
A collection pipeline built by a departed contractor had been failing intermittently for months. Nobody knew whether a quiet day meant no data or no run.
What we built
- Diagnosis of why it was failing, before touching the collection logic
- Retry, backoff, and rotation appropriate to how the source actually defends itself
- Monitoring that distinguishes 'ran and found nothing' from 'did not run'
- A written runbook so the next failure is not an investigation
What changed
A silent failure became a notification, and the team stopped discovering gaps weeks after the fact.
A decade of back-archive PDFs turned into a queryable dataset
The answers a team needed were already in thousands of documents sitting in a shared drive, findable only by someone who remembered which file it was in.
What we built
- Bulk text and table extraction across mixed document quality, including scans
- A schema for the fields that mattered, with extraction confidence recorded per field
- Search across the extracted set, linked back to the page it came from
- A re-run path, so improving the extraction does not mean starting over
What changed
Questions that used to require someone's memory of the archive became a search.
A scraping prototype in a day, then the developer who kept it running
A fitness-product company needed sports data scraped quickly, and had struggled to find remote developers it could work with.
What we built
- A working scraping prototype for the target data set, delivered within a day of the ask
- A developer screened and placed on the client's work
- A search audit and keyword plan later in the engagement
What changed
It had its data the next day, praised the developer in writing, and later introduced us to a third firm unprompted.
List building and contact extraction for a US industrial supplier
A US supplier needed email-marketing lists built in Apollo, and web addresses and contacts extracted for its target accounts.
What we built
- Apollo list building against the supplier's segments
- Automated extraction of web addresses and contacts for the target set
What changed
The supplier's outreach had lists behind it rather than a spreadsheet someone was still typing.
Elsewhere on the site
Read next
- PrototypesLabsWorking prototypes on sample data, including the kind of surface this practice builds. Nothing there is a client system.
- The workThe workCase pages, the systems that could not be named, and every figure with its source.
- The workThe evidence registerEvery record behind the work page, what kind of evidence each is, and what is withheld.
Where to next
The free first step: a Data Map.
Send the market you need data from, plus a sample of what you have. The reply says what to do first, and whether it is worth doing at all.