Modern Data Engineering with Databricks, Spark & Snowflake Course
Engineer modern data platforms with Spark, Databricks and Snowflake across batch, streaming, lakehouse, warehouse and governance.
Talk to an advisor on WhatsApp
How you will learn
What you will learn, module by module
Design end-to-end data platforms spanning batch and streaming pipelines, Spark, Databricks, lakehouse architecture, Snowflake, cloud warehousing, orchestration and governance. Progress from Modern Data Platform Architecture to Modern Data Engineering Capstone through guided labs, assessed projects, and portfolio evidence.
01Module 1 · 6 hoursModern Data Platform ArchitectureDesign a target architecture for a multi-source analytics platform with SLAs and data quality expectations.
- OLTP versus analytics
- Warehouse/lake/lakehouse
- Medallion architecture
- Batch versus streaming
- Data contracts
- Governance
- Platform tradeoffs
- Tools and platforms
- Architecture diagrams, cloud object storage concepts
- Portfolio evidence
- Modern data platform blueprint
- Assessment
- Architecture case
02Module 2 · 10 hoursApache Spark Engineering FoundationsBuild distributed transformations on a multi-file dataset and inspect execution behaviour.
- Spark architecture
- DataFrames
- Transformations/actions
- Partitions
- Joins
- Shuffle
- Spark SQL
- Tools and platforms
- Apache Spark, PySpark, Spark SQL
- Portfolio evidence
- PySpark data pipeline
- Assessment
- Spark lab
03Module 3 · 12 hoursDatabricks Lakehouse and Delta EngineeringBuild bronze-silver-gold Delta tables with data quality checks.
- Workspaces
- Notebooks/repos
- Delta tables
- ACID
- Schema enforcement/evolution
- Medallion layers
- Catalog/governance concepts
- Tools and platforms
- Databricks, Delta Lake, Unity Catalog concepts
- Portfolio evidence
- Databricks medallion pipeline
- Assessment
- Lakehouse lab
04Module 4 · 10 hoursBatch ETL/ELT and Pipeline OrchestrationCreate an incremental pipeline with retries, checkpoints and data quality gates.
- Incremental ingestion
- CDC concepts
- Idempotency
- Scheduling
- Dependencies
- Retries
- Parameters
- Tools and platforms
- Databricks workflows/Lakeflow concepts, Airflow optional, dbt optional
- Portfolio evidence
- Production-style batch pipeline
- Assessment
- Pipeline review
05Module 5 · 10 hoursStreaming Data EngineeringBuild a streaming pipeline that aggregates events and writes curated tables.
- Event streaming concepts
- Windows
- Watermarks
- State
- Late data
- Exactly-once concepts
- Stream-batch unification
- Tools and platforms
- Spark Structured Streaming, Kafka concepts
- Portfolio evidence
- Real-time pipeline demo
- Assessment
- Streaming lab
06Module 6 · 10 hoursSnowflake Data Warehousing EngineeringLoad, transform and optimise a warehouse workload in Snowflake with role-based access.
- Snowflake architecture
- Databases/schemas
- Warehouses
- Stages
- File loading
- Streams/tasks concepts
- Snowpark awareness
- Tools and platforms
- Snowflake, SQL, Snowpark concepts
- Portfolio evidence
- Snowflake analytics pipeline
- Assessment
- Warehouse lab
07Module 7 · 8 hoursAnalytics Engineering, Quality and GovernanceCreate tested curated models with documentation, lineage and ownership rules.
- Dimensional models
- Transformation layers
- dbt-style testing/documentation
- Lineage
- Metadata
- Data quality SLAs
- Access controls
- Tools and platforms
- SQL, dbt concepts, data catalog/governance tooling
- Portfolio evidence
- Governed analytics layer
- Assessment
- Quality checkpoint
08Module 8 · 8 hoursPerformance, Cost and Reliability EngineeringTune one Spark and one Snowflake workload, quantify performance/cost changes and document tradeoffs.
- Partitioning/clustering
- File sizing
- Caching
- Query plans
- Autoscaling
- Warehouse sizing
- Spark tuning
- Tools and platforms
- Databricks/Spark, Snowflake monitoring
- Portfolio evidence
- Performance and cost benchmark
- Assessment
- Optimisation report
09Module 9 · 18 hoursModern Data Engineering CapstoneBuild an end-to-end enterprise data platform using Databricks/Spark and Snowflake with curated BI-ready outputs.
- Ingestion
- Batch/stream processing
- Lakehouse
- Warehouse
- Orchestration
- Quality
- Governance
- Tools and platforms
- Databricks, Spark, Snowflake, SQL, Git
- Portfolio evidence
- Enterprise data engineering portfolio project
- Assessment
- Capstone architecture and demo
Projects you will build
2 portfolio projects plus module evidence
Lakehouse-to-Warehouse Data Platform
Ingest raw data, process with Spark/Databricks, curate in Snowflake and serve analytics.
Pipeline code · architecture · quality tests · lineage · performance reportReal-Time Customer Event Platform
Process streaming events into curated real-time and historical tables.
Streaming pipeline · medallion tables · monitoring · BI-ready outputsWhy this course
Modern data engineering requires dependable ingestion, transformation, orchestration, quality, streaming, warehousing, and governance across the full platform lifecycle.
The curriculum progresses from Modern Data Platform Architecture to Modern Data Engineering Capstone, with guided labs, assessments, and two portfolio projects: Lakehouse-to-Warehouse Data Platform and Real-Time Customer Event Platform.
Who this course is for
SQL/Python practitioners moving into enterprise data pipelines, lakehouse and warehouse engineering.
What you will be able to do
- Design a target architecture for a multi-source analytics platform with SLAs and data quality expectations.
- Build distributed transformations on a multi-file dataset and inspect execution behaviour.
- Build bronze-silver-gold Delta tables with data quality checks.
- Create an incremental pipeline with retries, checkpoints and data quality gates.
- Build a streaming pipeline that aggregates events and writes curated tables.
- Tune one Spark and one Snowflake workload, quantify performance/cost changes and document tradeoffs.
- Build an end-to-end enterprise data platform using Databricks/Spark and Snowflake with curated BI-ready outputs.
Technology you will use in this course
Data Engineering & Analytics Engineer
This course supports the development of skills relevant to roles such as Data Engineer, Analytics Engineer, Databricks/Spark Engineer, and Cloud Data Engineer. The strongest learner outcome is a portfolio that shows the problem, implementation, testing or evaluation, documentation and a clear explanation of decisions—not a certificate alone.
Course evidence and instruction
Ranjeet Kumar
Advisor, Brightnest AI Academy · Innovation & Growth LeaderA technologist and data leader with 15+ years of experience applying data, artificial intelligence and machine learning to complex problems, scalable products and business growth.
Learner experience
I started with basic Excel knowledge. The SQL, Power BI and Python projects helped me explain business insights clearly and move into an analyst role.
Industry and technology ecosystem
Clear answers before you enrol
Engineer modern data platforms with Spark, Databricks and Snowflake across batch, streaming, lakehouse, warehouse and governance.
Is the Modern Data Engineering course suitable for beginners?
This course progresses from intermediate to advanced level. Learners should understand SQL and Python fundamentals. Basic data modelling or analytics experience is recommended.
What will I build during the course?
You will complete guided labs in every module and build two portfolio projects: Lakehouse-to-Warehouse Data Platform and Real-Time Customer Event Platform. Deliverables include working files or code, documentation, testing or evaluation evidence, and a final presentation.
Which tools and platforms are covered?
Key tools include Architecture diagrams, cloud object storage concepts, Apache Spark, PySpark, Spark SQL, Databricks, Delta Lake, and Unity Catalog concepts. Additional platforms are introduced in relevant modules through practical tasks, and the toolset may evolve as industry practice changes.
How long does the course take?
The course includes approximately 92 guided learning hours across 9 modules, normally delivered over 12–14 weeks depending on batch intensity and learner practice time.
Which career paths can this course support?
The curriculum supports the development of skills relevant to roles such as Data Engineer, Analytics Engineer, Databricks/Spark Engineer, and Cloud Data Engineer. Career outcomes depend on prior experience, project quality, interview readiness and market conditions; employment is not guaranteed.
Will I receive mentor and career support?
The course includes live instruction, lab support, assignment feedback, project reviews and career preparation covering portfolio development, resume writing, LinkedIn profile improvement, and interview guidance.
Ready to start your Modern Data Engineering with Databricks, Spark & Snowflake journey?
Review the full curriculum, experience a live class and confirm the right starting point before enrolling.
