Jorvik is a collection of utilities for creating and managing ETL pipeline in Pyspark. Build from Data Engineers for Data Engineers.
The Jorvik project welcomes your expertise and enthusiasm!
Writing code isn’t the only way to contribute. You can also:
- review pull requests
- suggest improvements through issues
- let us know your pain-points and repetitive tasks
- help us stay on top of new and old issues
- develop tutorials, videos, presentations, and other educational materials
See How to Contribute for instructions on setting up your local machine and opening your first Pull Request.
Jorvik is available in Pypi and can be installed with pip
pip install jorvik- Storage: Interact with the storage layer
- Pipelines: Build and test etl pipelines with ease
- Data Lineage: Track data lineage
See the full power of jorvik when all the features come together in the examples bellow:
- Transactions: A multi step pipeline that creates customer statistics from customers and transaction data.