Skip to content

Latest commit

 

History

History

README.md

ETL

The Data Department uses the ETL pipeline to retrieve and clean all of the third-party data and first-party data stored on shared drives rather than iasWorld needed for modeling and reporting.

We use a new copy of this excel workbook to track data refresh progress each year.

This pipeline is generally only used once a year to refresh as much data as possible prior to modeling. Whoever is responsible for working through the data refresh should work as much of the raw scripts as they can, then switch over to running warehouse scripts to clean raw data that has been updated, or warehouse scripts that don't have raw counterparts. Once that's done Glue Crawlers can be tiggered to add new data to Athena.

Raw Scripts

While working through the raw scripts, keep in mind that many refrence data that is either static (no longer updated) or public open data that could possibly be updated, but likely isn't. A good example of the later is Enterprise Zone in spatial-economy.R. If we search the City of Chicago open data portal for that data asset, we can see the data hasn't been updated since 2021. While we still needed to check to make sure our data is as recent as is available, because it already is, we don't actually have to run that part of the script.

Warehouse Scripts

Once the raw scripts have been run, these scripts should be run for whichever raw bucket has new data in it. There are also some warehouse scripts that operate independent of a raw counterpart (spatial-census.R) and need to be run regardless of progress working through the raw scripts. Ideally, there shouldn't be any updates that need to be made for these scripts, but it's possible something about new raw data will cause an error that needs to be addressed.

Glue Crawlers

Run the glue crawler for the corresponding warehouse bucket once all the necessary warehouse scripts for that bucket have been run successfully.

DBT

After all of the new data in the spatial warehouse bucket has been successfully added to Athena we need to trigger a rebuild of our Athena location database using DBT. This can be done using the build-and-test-dbt workflow by passing it location.*.