Get in Touch
 Duration 21 hours

Course Outline

Introduction:

  • Apache Spark within the Hadoop Ecosystem
  • Brief overview of Python and Scala

Foundational Theory:

  • System Architecture
  • RDD Concepts
  • Transformations vs. Actions
  • Stages, Tasks, and Dependencies

Practical Application in Databricks (Hands-on Workshop):

  • Exercises utilizing the RDD API
  • Core action and transformation functions
  • Working with PairRDDs
  • Join Operations
  • Caching Strategies
  • Exercises utilizing the DataFrame API
  • SparkSQL
  • DataFrame operations: select, filter, group, and sort
  • UDFs (User Defined Functions)
  • Exploring the DataSet API
  • Streaming Data Processing

Deployment Strategies in AWS (Hands-on Workshop):

  • Foundations of AWS Glue
  • Analyzing the differences between AWS EMR and AWS Glue
  • Implementing example jobs in both environments
  • Evaluating the advantages and limitations of each

Additional Topics:

  • Introduction to Apache Airflow for orchestration

Requirements

Programming proficiency (preferably in Python and Scala)

Foundational knowledge of SQL

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories