Skip to content

Custom data collection

Web Data Collection Services

Custom data collection and structured datasets, assembled from the sources that actually answer your question.

Collected records aggregated into 500 m bins for coverage reviewSparseDense

Overview

The dataset you need rarely exists in one place

A useful dataset is usually a join. Store locations from one source, catchment demographics from a statistical agency, road geometry from an open mapping project, opening hours from map listings. Each piece is easy; making them agree on identifiers, geography and time is where projects stall.

We start from the decision you are trying to make and work backwards to the smallest dataset that supports it. That keeps scope honest and avoids paying for fields nobody will use.

The output is a documented dataset: a schema, a source note per field, a collection date and a validation report. If you later want it refreshed, extended to another country or joined to your own systems, the structure is already there.

Sample output

The same data, mapped

A static preview rather than an embedded map SDK, so the page stays fast. Interactive maps are built on request as part of a visualization project.

Collected records aggregated into 500 m bins for coverage reviewSparseDense

Scope

What a delivery includes

Every custom dataset ships with the same set of artefacts.

  • Schema definition

    Field names, types, units, coordinate system and allowed values, agreed before collection starts.

  • Source register

    Where each field came from, when it was collected and any access conditions that apply.

  • Primary dataset

    The records themselves, in the formats and destinations you specified.

  • Geocoded coordinates

    Latitude and longitude in WGS 84 for any record with an address, with match quality recorded.

  • Validation report

    Completeness by field, duplicate handling, outliers detected and records excluded.

  • Coverage summary

    What the dataset does and does not cover, by geography, category and time.

  • Change log

    For recurring projects, the added, removed and modified records since the previous run.

  • Reproducible pipeline

    The collection can be re-run on demand rather than rebuilt from scratch.

Output formats

Delivered the way your stack expects

CSV
Flat delivery for spreadsheets, BI tools and quick statistical work.
Excel
Multi-sheet workbook including the data dictionary and coverage notes.
JSON
Nested structures for records with repeating groups or variable fields.
GeoJSON
Geographic delivery for anything with coordinates or geometry.
Database or warehouse
Direct load into PostgreSQL/PostGIS, BigQuery, Snowflake or object storage.

Sample dataset

What you actually receive

Sample project data. Geocode quality is always returned so you can filter on match confidence rather than trusting every coordinate equally.

Sample project · Enriched location dataset · schema excerpt
Site IDAddressLatitudeLongitudeGeocode qualityPopulation 15 minCompetitors 1 km
SITE-001412 W Broadway, Vancouver BC49.26346-123.10389Rooftop184,20012
SITE-002980 Granville St, Vancouver BC49.27906-123.12481Rooftop231,80027
SITE-0033105 Main St, Vancouver BC49.25774-123.10088Interpolated148,6009
SITE-0042288 Kingsway, Vancouver BC49.24463-123.06821Rooftop126,4006
SITE-0051490 Lonsdale Ave, North Vancouver BC49.32819-123.07214Rooftop88,3004

Delivered work

This service on a real project

Sample projects built on this service, with the numbers they produced and what each one settled.

POI DataData CollectionMarket Research

POI Data Analysis

Building a comparable category census across five metro areas, and the standardisation work that made the comparison valid.

Raw records
94,200
Unique POIs
71,480
Category labels
14 → 1

What it showed

  • Ranked by absolute count, Metro A led by a wide margin. Ranked per capita, it placed third, and the two markets the team had considered marginal turned out to be the most densely served.
  • Chain share varied from 17.9% to 51.4% across markets that industry commentary had treated as broadly similar. That spread became the central finding of the research rather than a footnote.
Read the full case study
POI DataLocation IntelligenceHeatmap Analysis

Restaurant Location Intelligence

Mapping a city's food and drink offer, measuring competitive density, and producing a ranked shortlist of neighbourhoods for a new venue.

Venues collected
3,180
Coverage recovered
+6%
Areas shortlisted
42 → 6

What it showed

  • Market Square, the area the operator had assumed was strongest because it felt busy, had the highest competitive density in the city and the highest incumbent ratings. Entering there would have meant competing with well-established venues for demand that was already fully served.
  • Riverside North had a third of the competitive density with over half the reachable demand of Market Square. It scored highest overall despite feeling quieter on a weekend visit, because much of its demand is workplace-based and shows up on weekdays.
Read the full case study

How it works

Five steps, every project

  1. Step 01

    Tell us what data you need

    Describe the decision first. The field list follows from it.

  2. Step 02

    We define the data scope

    We propose sources, a schema and a coverage definition, then agree the acceptance criteria.

  3. Step 03

    We collect and process the data

    Sources are collected, standardised, geocoded where relevant and joined on agreed keys.

  4. Step 04

    We validate the dataset

    Completeness, duplicates, ranges and spatial sanity checks run before anything is delivered.

  5. Step 05

    We deliver the final result

    Dataset, documentation and validation report, in your formats and destinations.

Projects are quoted on record volume, number of sources, enrichment depth and refresh frequency. Describe the dataset you need and you get a fixed price against a written scope.

FAQ

Questions we get asked

What sources do you combine?

Publicly available web pages, open government and statistical data, open geospatial datasets such as OpenStreetMap and national mapping agencies, map listings, and any data you provide or are licensed to use. Each field's source is recorded in the delivery.

Can you work from a list we already have?

Yes, and it is often the cheapest starting point. We clean and deduplicate your list, geocode it, then add the attributes you are missing. You keep your identifiers so the result joins straight back into your systems.

How do you measure data quality?

Against criteria agreed before collection: field completeness thresholds, duplicate rates, geocoding match quality and, where a reference total exists, coverage against it. The validation report states the achieved numbers rather than claiming perfection.

Can the dataset be updated later?

Yes. Projects are built as repeatable pipelines, so a refresh is a re-run rather than a rebuild. Refreshes include a change log of added, removed and modified records.

How do you handle personal data?

We avoid collecting personal data unless it is strictly necessary and lawful for your stated purpose. Business contact details published by the business itself are treated as business data. We do not build datasets about private individuals.

What if a field turns out to be unavailable?

We tell you during scoping if we can already see the problem, and during the pilot batch if it only appears at scale. The field is either dropped from scope with a price adjustment or replaced with the closest reliable proxy, with your agreement.

Next step

Have a specific data requirement?

Tell us the geography, the fields and the cadence you need for data collection. You get a scoped plan, a sample and a fixed price before any work starts.