CAPGov, Public Data Platform
ETL pipelines turning large public datasets into evidence policy researchers can cite.
The record
- Role
- Undergraduate researcher, ETL
- Year
- 2023 – 2026
- Stack
- Python
- SQL
- Docker
- ETL
- Proof
- analytical dashboards in use
Overview
Three years building the data infrastructure behind evidence-based policy research: pipelines that ingest large public datasets, validate them, and land them somewhere analysts can trust.
What I built
- Built the ETL pipelines in Python and SQL that ingest and process the large public datasets behind the group's policy research.
- Moved a legacy setup onto containerized Docker services, so a researcher gets a working environment from one command and development matches production.
- Replaced manual data entry with validation at ingest, so a bad row is caught before it reaches an analytical dashboard rather than after someone has already cited it.