Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction, Objectives, and Migration Strategy
- Course goals, participant profile alignment, and success metrics
- High-level migration approaches and risk assessments
- Configuration of workspaces, repositories, and lab datasets
Day 1 — Migration Fundamentals and Architecture
- Lakehouse concepts, Delta Lake overview, and Databricks architecture
- Differences and implications between SMP and MPP for migration
- Medallion (Bronze→Silver→Gold) design principles and Unity Catalog overview
Day 1 Lab — Translating a Stored Procedure
- Practical migration of a sample stored procedure to a notebook
- Mapping temp tables and cursors to DataFrame transformations
- Validation and comparison against the original output
Day 2 — Advanced Delta Lake & Incremental Loading
- ACID transactions, commit logs, versioning, and time travel features
- Auto Loader, MERGE INTO patterns, upserts, and schema evolution
- OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage optimization
Day 2 Lab — Incremental Ingestion & Optimization
- Implementation of Auto Loader ingestion and MERGE workflows
- Application of OPTIMIZE, Z-ORDER, and VACUUM; result validation
- Measurement of read/write performance enhancements
Day 3 — SQL in Databricks, Performance & Debugging
- Analytical SQL features: window functions, higher-order functions, JSON/array handling
- Interpreting the Spark UI, DAGs, shuffles, stages, tasks, and identifying bottlenecks
- Query tuning patterns: broadcast joins, hints, caching, and spill reduction
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring a resource-intensive SQL process into optimized Spark SQL
- Using Spark UI traces to identify and resolve skew and shuffle issues
- Benchmarking before/after performance and documenting tuning steps
Day 4 — Tactical PySpark: Replacing Procedural Logic
- Spark execution model: driver, executors, lazy evaluation, and partitioning strategies
- Converting loops and cursors into vectorized DataFrame operations
- Modularization, UDFs/pandas UDFs, widgets, and reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Refactoring a procedural ETL script into modular PySpark notebooks
- Introducing parametrization, unit-style tests, and reusable functions
- Code review and application of best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Databricks Workflows: job design, task dependencies, triggers, and error handling
- Designing incremental Medallion pipelines with quality rules and schema validation
- Integration with Git (GitHub/Azure DevOps), CI, and testing strategies for PySpark logic
Day 5 Lab — Build a Complete End-to-End Pipeline
- Assembling a Bronze→Silver→Gold pipeline orchestrated via Workflows
- Implementing logging, auditing, retries, and automated validations
- Executing the full pipeline, validating outputs, and preparing deployment documentation
Operationalization, Governance, and Production Readiness
- Unity Catalog governance, lineage, and access control best practices
- Cost management, cluster sizing, autoscaling, and job concurrency patterns
- Deployment checklists, rollback strategies, and runbook creation
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations on migration work and lessons learned
- Gap analysis, recommended follow-up activities, and handoff of training materials
- References, further learning paths, and support options
Requirements
- A solid grasp of data engineering concepts
- Proficiency with SQL and stored procedures (Synapse / SQL Server)
- Knowledge of ETL orchestration concepts (ADF or similar tools)
Audience
- Technology managers with a background in data engineering
- Data engineers looking to transition procedural OLAP logic to Lakehouse patterns
- Platform engineers overseeing Databricks adoption
35 Hours