Data and methodology
Every figure in MarketGround traces back to a public, authoritative source, allocated onto a fine hexagonal grid so that any shape you draw, a radius, a drive time, a hand-drawn polygon, gets an answer built from the same auditable foundation. This page explains what the numbers are made of and how we check them.
Sources
| Dataset | What it provides | Cadence |
|---|---|---|
| 2020 Decennial Census blocks | Exact enumerated population at the finest geography the Census publishes | Decennial |
| American Community Survey (5-year) | Demographics: income, housing, education, race and ethnicity, commuting | Annual releases |
| Census LODES | Where people work relative to where they live, at block level | Annual |
| Census Population Estimates (PEP) | County population totals for the current year | Annual, revised |
| LandScan USA (ORNL) | Measured day and night population, 2016 through 2021 | Archived at 2021 |
| Overture Maps and Foursquare OS Places | Businesses and points of interest, with openings and closings over time | Monthly |
| OpenStreetMap and agency GTFS | The road network and transit timetables behind travel-time areas | Continuous, timetables weekly |
How a shape becomes numbers
Census data is published for fixed geographies (block groups) that rarely match the shape you care about. We allocate each block group's values down to a hexagonal grid of roughly 0.1 km² cells, using the 2020 census blocks inside it as the weights. Blocks are enumerated rather than modelled, and they nest exactly inside block groups, so the allocation is anchored to counted people rather than to assumptions. When you draw a shape, we intersect it with the grid and reconstruct every statistic from the cells it captures, weighting partial cells by their overlap.
Sums are summed. Medians are never summed: they are rebuilt from the underlying distribution brackets, the same way the Census computes them. A statistic that can only be approximated carries a visible caveat next to the number, in the interface and in exports.
Day and night population
Through 2021 the day and night surfaces are measured: LandScan USA, the only national product that separates the two. From 2022 onward no measured product exists, so we model daytime from the commuting ledger (residents, minus workers who leave, plus workers who arrive), adjusted for people who work from home, and anchor every county to the Census Bureau's own estimate for the target year. The modelled surface is then calibrated against the years where it overlaps the measured one, so its known biases are corrected by evidence rather than by judgment.
In the product these years appear as one continuous history, because that is how they are meant to be read: the same grid, the same resolution, the same census foundation, anchored to the same county totals. The seam is here, not on the chart. Years through 2021 are measured by LandScan USA (Oak Ridge National Laboratory archived the product after 2021), and years from 2022 onward are modelled as described above. Where the two overlap, they agree within a few percent, which is the calibration's evidence that the handover is sound. Two years of LandScan, 2019 and 2020, are identical releases, so the chart draws the repeat as a hollow point rather than pretend it is a second observation.
One caveat we state rather than bury: modelled daytime counts jobs, not visitors. A shopping centre is credited with its staff, not its customers, so treat modelled daytime in heavy retail areas as a lower bound. This is precisely what the calibration measures and partially corrects.
Travel-time areas
Drive, walk, cycle and transit areas are computed on our own routing infrastructure from OpenStreetMap roads and published agency timetables. A transit area depends on when you leave, so it requires a departure day and time. Where a point has no transit service, or an agency's timetable is not yet in our graph, the result is returned as a walking area and labelled as one. We refuse to present a walking shape as a transit answer.
What we check
- Loaded census blocks must sum exactly to each state's published decennial count. Not approximately: exactly.
- Allocation must conserve people. The population allocated to the grid is checked against the block totals it came from, per block group, within floating-point tolerance.
- Reconstructed medians are compared against the published values: over 99% fall within the Census Bureau's own margin of error.
- Every modelled year must match the Census county estimates for that year within 2%, state by state, or the load fails.
- A polygon drawn exactly on a block group boundary reproduces that block group's published counts, person for person.
These checks run as a gate, not a report. A state that fails any of them is not served.
What we do not publish
The specific calibration factors, model parameters, and allocation engineering are proprietary. The sources, the structure of the method, and the checks above are public because your confidence in the numbers should not require trusting us blindly.
Sources: US Census Bureau (public domain); LandScan USA, Oak Ridge National Laboratory (CC BY 4.0); Overture Maps Foundation (CDLA-Permissive 2.0); Foursquare OS Places (Apache 2.0); OpenStreetMap contributors (ODbL); Protomaps.