This intensive three-day course centers on the construction and optimization of high-efficiency data processing workflows utilizing PySpark, Pandas, and Polars within Kubernetes-based ecosystems.
Attendees will acquire a hands-on grasp of Spark application execution on Kubernetes, exploring how configuration choices at the application level directly impact performance, scalability, resource utilization, and operational costs. Key optimization domains covered include executor sizing, memory distribution, dynamic allocation mechanisms, partitioning methodologies, shuffle dynamics, resolving small-file bottlenecks, and the efficient handling of Parquet files.
The curriculum also tackles typical obstacles encountered when leveraging Pandas, such as memory constraints and out-of-memory errors, while presenting Polars as a high-performance alternative for specific processing tasks. Through practical exercises, participants will learn to diagnose performance and memory issues, evaluate various configuration strategies, and implement optimization techniques in realistic ETL and machine learning contexts.
The course places a strong emphasis on practical decision-making: equipping learners with the ability to pinpoint bottlenecks, select the right tools, configure Spark effectively, and strike a balance between performance and infrastructure resource consumption and cost.
Read more...