Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Databricks Formula 1 Data Engineering Project

An end-to-end Azure Databricks Data Engineering project that builds a modern data pipeline using the Medallion Architecture (Bronze, Silver, and Gold layers). The project ingests Formula 1 datasets, performs data cleansing and transformations using PySpark, and creates curated analytical tables for reporting and insights.

Key Features

  • Ingest CSV and JSON datasets into Azure Databricks.
  • Implement the Medallion Architecture (Bronze, Silver, Gold).
  • Perform data cleansing, validation, and transformations using PySpark.
  • Store processed data using Unity Catalog and Delta Lake.
  • Build optimized Gold layer tables for analytics and reporting.

Technologies Used

  • Azure Databricks
  • PySpark
  • Delta Lake
  • Unity Catalog
  • Azure Data Lake Storage Gen2 (ADLS Gen2)
  • Git & GitHub

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages