
Scope
Catalogued 42 legacy datasets
Datasets ranged from a handful of rows to several hundred thousand, many maintained manually in spreadsheets with no documented source.
Public sector partner + CIDS data readiness
A data cataloguing and lineage project for a public sector team preparing several long-running spreadsheet-based datasets for a move into a shared data platform.

We finally know where a number in a report actually came from.
Data governance lead
Public sector partner

Scope
Datasets ranged from a handful of rows to several hundred thousand, many maintained manually in spreadsheets with no documented source.

Outcome
The catalogue's findings on duplicate and conflicting data sources became the business case for consolidating onto a single platform.
Project overview
The team’s datasets had accumulated over several years, mostly as spreadsheets passed between staff with little record of where the underlying numbers came from or which version was current.
CIDS worked through each dataset with the team that owned it, documenting source, update frequency, known quality issues, and downstream reports that depended on it, building a lightweight catalogue rather than a heavyweight governance tool the team wouldn’t maintain.
The exercise surfaced several datasets that were quietly duplicating each other with slightly different numbers, which had been an open question inside the team for some time without anyone having the evidence to resolve it.

Before and after
2 operational shifts from the project
Before
Dataset provenance
Source and update history known only informally, if at all.
Duplicate data
Suspected but unconfirmed duplication across teams.
After
Dataset provenance
42 datasets catalogued with documented source and lineage.
Duplicate data
Duplicate and conflicting sources identified and flagged.