Individual project

Public sector partner + CIDS data readiness

Data governance uplift

A data cataloguing and lineage project for a public sector team preparing several long-running spreadsheet-based datasets for a move into a shared data platform.

Timeline
January 2025 - June 2025
Datasets catalogued
42
Fields with documented lineage
6% to 89%
Printed data schema diagrams being annotated at a desk.
Image adapted from the CIDS project image set for the project detail page.
We finally know where a number in a report actually came from.
DG

Data governance lead

Public sector partner

Recognition and proof

Project signals

A dashboard display showing decision support charts.

Scope

Catalogued 42 legacy datasets

Datasets ranged from a handful of rows to several hundred thousand, many maintained manually in spreadsheets with no documented source.

A research workspace connected to production-ready work.

Outcome

Built the case for a shared platform

The catalogue's findings on duplicate and conflicting data sources became the business case for consolidating onto a single platform.

Project overview

Making years of undocumented spreadsheets auditable.

The team’s datasets had accumulated over several years, mostly as spreadsheets passed between staff with little record of where the underlying numbers came from or which version was current.

CIDS worked through each dataset with the team that owned it, documenting source, update frequency, known quality issues, and downstream reports that depended on it, building a lightweight catalogue rather than a heavyweight governance tool the team wouldn’t maintain.

The exercise surfaced several datasets that were quietly duplicating each other with slightly different numbers, which had been an open question inside the team for some time without anyone having the evidence to resolve it.

A public sector team reviewing data readiness material.
Each dataset's lineage was mapped by hand with the team that owned it, since the source systems didn't have any machine-readable metadata to draw on.89% of fields now have documented lineage

Before and after

2 operational shifts from the project

Data Governance Uplift

Before

Dataset provenance

Source and update history known only informally, if at all.

Duplicate data

Suspected but unconfirmed duplication across teams.

After

Dataset provenance

42 datasets catalogued with documented source and lineage.

A report's numbers can now be traced back to where they came from.

Duplicate data

Duplicate and conflicting sources identified and flagged.

Gives the team a concrete list to reconcile before platform migration.