Reddit Batch Data Pipeline

A data pipeline to extract Reddit data from r/dataengineering.

Output is a Google Data Studio report, providing insight into the Data Engineering official subreddit.

Motivation

Project was based on an interest in Data Engineering and the types of Q&A found on the official subreddit.

It also provided a good opportunity to develop skills and experience in a range of tools. As such, project is more complex than required, utilising dbt, airflow, docker and cloud based storage.

Architecture

Extract data using Reddit API
Load into AWS S3
Copy into AWS Redshift
Transform using dbt
Create Google Data Studio Dashboard
Orchestrate with Airflow in Docker
Create AWS resources with Terraform

Setup

Follow below steps to setup pipeline. I've tried to explain steps where I can. Feel free to make improvements/changes.

NOTE: This was developed using an M1 Macbook Pro. If you're on Windows or Linux, you may need to amend certain components if issues are encountered.

As AWS offer a free tier, this shouldn't cost you anything unless you amend the pipeline to extract large amounts of data, or keep infrastructure running for 2+ months. However, please check AWS free tier limits, as this may change.

First clone the repository into your home directory and follow the steps.

git clone 
cd Reddit-batch-data-pipeline

Name		Name	Last commit message	Last commit date
Latest commit History 2 Commits
airflow		airflow
images		images
instructions		instructions
terraform		terraform
.DS_Store		.DS_Store
LICENSE		LICENSE
README.md		README.md
requirements.txt		requirements.txt

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Repository files navigation

Reddit Batch Data Pipeline

Motivation

Architecture

Setup

About

Uh oh!

Releases

Packages

Uh oh!

Contributors 2

Uh oh!

Languages

License

jrdegbe/Reddit-Data-Pipeline

Folders and files

Latest commit

History

Repository files navigation

Reddit Batch Data Pipeline

Motivation

Architecture

Setup

About

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Uh oh!

Contributors 2

Uh oh!

Languages

Packages