掌握PySpark核心技能,从零构建可扩展的ETL管道。本课程通过项目实战,涵盖DataFrame操作、数据清洗、性能优化及Databricks云端应用,助你精通大数据分析。
原始标题:Big Data Analytics & Designing ETL Pipelines with Pyspark

本课程是 Christ Raharja 于 2026 年 8 月推出的《Big Data Analytics & Designing ETL Pipelines with Pyspark》实战课。课程时长 4 小时 6 分钟(共 21 讲),采用**完全基于项目驱动(Project-based)**的教学模式。它将大数据分析与 PySpark 紧密结合,带你从零开始在本地及云端(Databricks)环境构建高效、可扩展的数据管道(ETL/ELT)。
Published 8/2026
Created by Christ Raharja
MP4 | Video: h264, 2560×1440 | Audio: AAC, 44.1 KHz, 2 Ch
Level: All Levels | Genre: eLearning | Language: English | Duration: 21 Lectures ( 4h 6m ) | Size: 3 GB
Learn big data analytics, DataFrame Operation, data transformation, ETL pipelines, SparkSQL, Pyspark, Databricks, SQLite
What you’ll learn
⚡ Learn the basic fundamentals of big data analytics and ETL pipelines
⚡ Learn how to create Spark DataFrame, load CSV file, and convert data types using Pyspark
⚡ Learn about DataFrame operations, select columns, filter and sort data using Pyspark
⚡ Learn how to clean data, handle missing values, and remove duplicates using Pyspark
⚡ Learn about data partitioning and performance optimization
⚡ Learn about data transformation and data aggregation
⚡ Learn how to merge datasets using Join and Union operations
⚡ Learn how to build and design ETL pipelines for flight analytics project
⚡ Learn how to extract flight data from multiple sources using Pyspark
⚡ Learn how to clean and transform flight data using Pyspark
⚡ Learn how to load flight data to SQLite database
⚡ Learn how to analyze logistics data using SparkSQL
⚡ Learn how to build machine learning for predicting restaurant revenue using MLlib
⚡ Learn how to set up Databricks workspace and load sport analytics data
⚡ Learn how to analyze football player performance on Databricks
⚡ Learn how to build and design ETL pipelines for energy consumption analytics project
⚡ Learn how to extract and transform energy consumption data using Pyspark
⚡ Learn how to load energy consumption data to SQLite database
Requirements
❗ No previous experience in data engineering is required
❗ Basic knowledge in Python and Pyspark
Description
This course contains the use of artificial intelligence
Disclosure: AI tools were used only to assist in creating the course outline and course thumbnail. All instructional content, explanations, and project walkthroughs were fully created manually by the instructor.
Welcome to Big Data Analytics & Designing ETL Pipelines with Pyspark course. This is a comprehensive project based course where you will learn how to process and analyze large datasets, perform complex data transformation, and build ETL pipelines. This course is a perfect combination between big data and Pyspark, making it an ideal opportunity to practice your programming skills while improving your technical knowledge in data engineering. In the introduction session, you will learn the basic fundamentals of big data analytics and ETL, you will learn big data workflow and the difference between ETL and ELT. Then, in the next section, we will learn basic concepts of Pyspark like creating Pyspark DataFrame, reading CSV files, DataFrame operations, selecting columns, filtering, sorting data, converting data types, handling missing values and removing duplicates. Afterward, in the next section, we will learn how to transform data using aggregation techniques and combine datasets using union and join operations. These skills will enable us to efficiently process large datasets and prepare data for developing ETL pipelines. Following that, we are going to design a complete ETL pipeline for flight and airport data. We will start with data extraction and ingestion, where we will get data from three different sources, flight records CSV, airport information CSV, and aircraft specifications JSON. Then, for the transformation stage, we are going to clean, standardize, and integrate the datasets to create structured flight analytics data. For the last step, we are going to load the transformed data into a SQLite database for downstream analysis and reporting. Then, in the next section, we are going to analyse a logistics dataset using Spark SQL, where we will create SQL queries to evaluate shipment performance, compare shipping cost and identify logistics trends and patterns. After that, we are going to build a machine learning model for predicting restaurant revenue using PySpark and MLlib. This model will be able to predict future restaurant revenue based on factors like customer volume, menu pricing, and marketing spend. In the next section, we are going to set up a Databricks workspace, load a sports analytics dataset into Databricks, and perform analysis and transformation on football data using PySpark, the objective is to analyse player performance and team statistics. Lastly, at the end of the course, we are going to work on real world data engineering projects, where we will design ETL pipelines for energy consumption data. We will extract and transform the data, then load the processed results into a SQLite database, we are going to create three separate pipelines for energy consumption per household, energy consumption per appliance, and energy consumption based on weather conditions.
First of all, before getting into the course, we need to ask this question to ourselves, why should we learn about big data analytics? Well, here is my answer. Nowadays, almost all industries generate massive amounts of data from applications, transactions, sensors, websites, and other digital systems. As the volume and complexity of data continue to grow, businesses need efficient technologies to process and analyse this data at scale. This is where big data analytics comes in, allowing us to work with large datasets and extract valuable insights that can support better decision making.
Below are things that you can expect to learn from this course
✨ Learn the basic fundamentals of big data analytics and ETL pipelines
✨ Learn how to create Spark DataFrame, load CSV file, and convert data types
✨ Learn about DataFrame operations, select columns, filter and sort data using Pyspark
✨ Learn how to clean data, handle missing values, and remove duplicates using Pyspark
✨ Learn about data partitioning and performance optimization
✨ Learn about data transformation and data aggregation
✨ Learn how to merge datasets using Join and Union operations
✨ Learn how to build and design ETL pipelines for flight analytics project
✨ Learn how to extract flight data from multiple sources using Pyspark
✨ Learn how to clean and transform flight data using Pyspark
✨ Learn how to load flight data to SQLite database
✨ Learn how to analyze logistics data using SparkSQL
✨ Learn how to build machine learning for predicting restaurant revenue using MLlib
✨ Learn how to set up Databricks workspace and load sport analytics data
✨ Learn how to analyze football player performance on Databricks
✨ Learn how to build and design ETL pipelines for energy consumption analytics project
✨ Learn how to extract and transform energy consumption data using Pyspark
✨ Learn how to load energy consumption data to SQLite database
Who this course is for
⭐ Data engineers who are interested in building and designing ETL pipelines using Pyspark
⭐ Data scientists who are interested in processing and analyzing large dataset using Pyspark
此处内容需要权限查看
会员免费查看



