Skip to content

Tooling

Data Tooling

The initial part of Phase II of the project focusses on the concept of peer hex identification: given a hex cell, can we identify, nationally, other cells that are very similar in terms of built-environment and sociodemographic features? The cells are small so there are a lot of them, and a lot of features.

We've put a lot of thought over the last few months into how best to build the necessary datasets and implement this. We've homed in on a very flexible and efficient way of sharing this data (subject to licencing and confidentiality restrictions).

We've also built an automated ETL (extract-transform-load) process to build these datasets. This post describes what goes into it, how it works, and how others can get at the data.