This pipeline builds a medallion lakehouse (bronze, silver and gold Delta tables on S3), orchestrated by Apache Airflow 3 on Kubernetes. PySpark jobs, started by Airflow as SparkApplications on the Spark Operator, load raw JSONL and CSV files from the landing zone into bronze; to stress-test the cluster, one job multiplies 3 million flights 330 times into 990 million rows. dbt (dbt-spark through the Spark Thrift Server) turns bronze into typed, deduplicated silver tables and builds the gold table, and every model runs its data tests in the same dbt build. A second DAG simulates a daily delivery: PySpark lands one day of 1 million flights in bronze and MERGEs it into silver, so a re-run updates instead of duplicating. GitLab CI converts the notebooks to Python and copies them and the dbt project to S3 after every merge, and Airflow pulls the code from there at run time.