School of Data Engineering & Analytics · Intermediate–Advanced

Modern Data Engineering with Databricks, Spark & Snowflake Course

Engineer modern data platforms with Spark, Databricks and Snowflake across batch, streaming, lakehouse, warehouse and governance.

Talk to an advisor on WhatsApp
Modern Data Engineering with Databricks, Spark & Snowflake course illustration at Brightnest AI Academy
Modern Data Platform ArchitectureApache Spark Engineering FoundationsDatabricks Lakehouse and Delta EngineeringBatch ETL/ELT and Pipeline Orchestration
Architecture diagramscloud object storage conceptsApache SparkPySparkSpark SQL
Duration12–14 weeks92 hours
Batch startsConfirm with admissionsOpen for registration
Learning formatLive mentor-led instruction, guided labs, assignments, feedback and project reviewsLive online / classroom
Curriculum9 modulesLabs and assessed capstone
Portfolio2 projectsPlus module evidence
LevelIntermediate–AdvancedCourse level
PathwayData Engineering & Analytics EngineerRelated career pathway

How you will learn

Live instructor-led sessions that connect concepts to real workplace decisions.
Guided labs and workshops in every module.
Assignments, checkpoints and practical feedback.
Portfolio documentation, demonstrations and capstone review.
Access to recordings and LMS resources according to the published batch policy.
Career preparation based on completed work and target roles.
Course curriculum

What you will learn, module by module

Design end-to-end data platforms spanning batch and streaming pipelines, Spark, Databricks, lakehouse architecture, Snowflake, cloud warehousing, orchestration and governance. Progress from Modern Data Platform Architecture to Modern Data Engineering Capstone through guided labs, assessed projects, and portfolio evidence.

01Module 1 · 6 hoursModern Data Platform ArchitectureDesign a target architecture for a multi-source analytics platform with SLAs and data quality expectations.
Topics you will cover
  • OLTP versus analytics
  • Warehouse/lake/lakehouse
  • Medallion architecture
  • Batch versus streaming
  • Data contracts
  • Governance
  • Platform tradeoffs
Tools and platforms
Architecture diagrams, cloud object storage concepts
Portfolio evidence
Modern data platform blueprint
Assessment
Architecture case
02Module 2 · 10 hoursApache Spark Engineering FoundationsBuild distributed transformations on a multi-file dataset and inspect execution behaviour.
Topics you will cover
  • Spark architecture
  • DataFrames
  • Transformations/actions
  • Partitions
  • Joins
  • Shuffle
  • Spark SQL
Tools and platforms
Apache Spark, PySpark, Spark SQL
Portfolio evidence
PySpark data pipeline
Assessment
Spark lab
03Module 3 · 12 hoursDatabricks Lakehouse and Delta EngineeringBuild bronze-silver-gold Delta tables with data quality checks.
Topics you will cover
  • Workspaces
  • Notebooks/repos
  • Delta tables
  • ACID
  • Schema enforcement/evolution
  • Medallion layers
  • Catalog/governance concepts
Tools and platforms
Databricks, Delta Lake, Unity Catalog concepts
Portfolio evidence
Databricks medallion pipeline
Assessment
Lakehouse lab
04Module 4 · 10 hoursBatch ETL/ELT and Pipeline OrchestrationCreate an incremental pipeline with retries, checkpoints and data quality gates.
Topics you will cover
  • Incremental ingestion
  • CDC concepts
  • Idempotency
  • Scheduling
  • Dependencies
  • Retries
  • Parameters
Tools and platforms
Databricks workflows/Lakeflow concepts, Airflow optional, dbt optional
Portfolio evidence
Production-style batch pipeline
Assessment
Pipeline review
05Module 5 · 10 hoursStreaming Data EngineeringBuild a streaming pipeline that aggregates events and writes curated tables.
Topics you will cover
  • Event streaming concepts
  • Windows
  • Watermarks
  • State
  • Late data
  • Exactly-once concepts
  • Stream-batch unification
Tools and platforms
Spark Structured Streaming, Kafka concepts
Portfolio evidence
Real-time pipeline demo
Assessment
Streaming lab
06Module 6 · 10 hoursSnowflake Data Warehousing EngineeringLoad, transform and optimise a warehouse workload in Snowflake with role-based access.
Topics you will cover
  • Snowflake architecture
  • Databases/schemas
  • Warehouses
  • Stages
  • File loading
  • Streams/tasks concepts
  • Snowpark awareness
Tools and platforms
Snowflake, SQL, Snowpark concepts
Portfolio evidence
Snowflake analytics pipeline
Assessment
Warehouse lab
07Module 7 · 8 hoursAnalytics Engineering, Quality and GovernanceCreate tested curated models with documentation, lineage and ownership rules.
Topics you will cover
  • Dimensional models
  • Transformation layers
  • dbt-style testing/documentation
  • Lineage
  • Metadata
  • Data quality SLAs
  • Access controls
Tools and platforms
SQL, dbt concepts, data catalog/governance tooling
Portfolio evidence
Governed analytics layer
Assessment
Quality checkpoint
08Module 8 · 8 hoursPerformance, Cost and Reliability EngineeringTune one Spark and one Snowflake workload, quantify performance/cost changes and document tradeoffs.
Topics you will cover
  • Partitioning/clustering
  • File sizing
  • Caching
  • Query plans
  • Autoscaling
  • Warehouse sizing
  • Spark tuning
Tools and platforms
Databricks/Spark, Snowflake monitoring
Portfolio evidence
Performance and cost benchmark
Assessment
Optimisation report
09Module 9 · 18 hoursModern Data Engineering CapstoneBuild an end-to-end enterprise data platform using Databricks/Spark and Snowflake with curated BI-ready outputs.
Topics you will cover
  • Ingestion
  • Batch/stream processing
  • Lakehouse
  • Warehouse
  • Orchestration
  • Quality
  • Governance
Tools and platforms
Databricks, Spark, Snowflake, SQL, Git
Portfolio evidence
Enterprise data engineering portfolio project
Assessment
Capstone architecture and demo
Applied portfolio

Projects you will build

2 portfolio projects plus module evidence

Portfolio project 1

Lakehouse-to-Warehouse Data Platform

Ingest raw data, process with Spark/Databricks, curate in Snowflake and serve analytics.

Pipeline code · architecture · quality tests · lineage · performance report
Portfolio project 2

Real-Time Customer Event Platform

Process streaming events into curated real-time and historical tables.

Streaming pipeline · medallion tables · monitoring · BI-ready outputs
Course value

Why this course

Modern data engineering requires dependable ingestion, transformation, orchestration, quality, streaming, warehousing, and governance across the full platform lifecycle.

The curriculum progresses from Modern Data Platform Architecture to Modern Data Engineering Capstone, with guided labs, assessments, and two portfolio projects: Lakehouse-to-Warehouse Data Platform and Real-Time Customer Event Platform.

Course fit

Who this course is for

SQL/Python practitioners moving into enterprise data pipelines, lakehouse and warehouse engineering.

Intermediate–AdvancedData Engineering & Analytics Engineer
PrerequisitesLearners should understand SQL and Python fundamentals. Basic data modelling or analytics experience is recommended.
Practical capabilities

What you will be able to do

  • Design a target architecture for a multi-source analytics platform with SLAs and data quality expectations.
  • Build distributed transformations on a multi-file dataset and inspect execution behaviour.
  • Build bronze-silver-gold Delta tables with data quality checks.
  • Create an incremental pipeline with retries, checkpoints and data quality gates.
  • Build a streaming pipeline that aggregates events and writes curated tables.
  • Tune one Spark and one Snowflake workload, quantify performance/cost changes and document tradeoffs.
  • Build an end-to-end enterprise data platform using Databricks/Spark and Snowflake with curated BI-ready outputs.
Tools and platforms

Technology you will use in this course

Architecture diagramscloud object storage conceptsApache SparkPySparkSpark SQLDatabricksDelta LakeUnity Catalog conceptsDatabricks workflowsLakeflow conceptsAirflowdbtSpark Structured StreamingKafka conceptsSnowflakeSQL
Career relevance

Data Engineering & Analytics Engineer

This course supports the development of skills relevant to roles such as Data Engineer, Analytics Engineer, Databricks/Spark Engineer, and Cloud Data Engineer. The strongest learner outcome is a portfolio that shows the problem, implementation, testing or evaluation, documentation and a clear explanation of decisions—not a certificate alone.

Course evidence and instruction

Academy advisor

Ranjeet Kumar

Advisor, Brightnest AI Academy · Innovation & Growth Leader

A technologist and data leader with 15+ years of experience applying data, artificial intelligence and machine learning to complex problems, scalable products and business growth.

What our learners say

Learner experience

I started with basic Excel knowledge. The SQL, Power BI and Python projects helped me explain business insights clearly and move into an analyst role.
Anisha SharmaBusiness Analyst · Analytics & Consulting

Industry and technology ecosystem

MicrosoftAmazon Web ServicesDeloitteTech MahindraTata Consultancy ServicesWipro
Course FAQs

Clear answers before you enrol

Engineer modern data platforms with Spark, Databricks and Snowflake across batch, streaming, lakehouse, warehouse and governance.

Is the Modern Data Engineering course suitable for beginners?

This course progresses from intermediate to advanced level. Learners should understand SQL and Python fundamentals. Basic data modelling or analytics experience is recommended.

What will I build during the course?

You will complete guided labs in every module and build two portfolio projects: Lakehouse-to-Warehouse Data Platform and Real-Time Customer Event Platform. Deliverables include working files or code, documentation, testing or evaluation evidence, and a final presentation.

Which tools and platforms are covered?

Key tools include Architecture diagrams, cloud object storage concepts, Apache Spark, PySpark, Spark SQL, Databricks, Delta Lake, and Unity Catalog concepts. Additional platforms are introduced in relevant modules through practical tasks, and the toolset may evolve as industry practice changes.

How long does the course take?

The course includes approximately 92 guided learning hours across 9 modules, normally delivered over 12–14 weeks depending on batch intensity and learner practice time.

Which career paths can this course support?

The curriculum supports the development of skills relevant to roles such as Data Engineer, Analytics Engineer, Databricks/Spark Engineer, and Cloud Data Engineer. Career outcomes depend on prior experience, project quality, interview readiness and market conditions; employment is not guaranteed.

Will I receive mentor and career support?

The course includes live instruction, lab support, assignment feedback, project reviews and career preparation covering portfolio development, resume writing, LinkedIn profile improvement, and interview guidance.

Ready to start?

Ready to start your Modern Data Engineering with Databricks, Spark & Snowflake journey?

Review the full curriculum, experience a live class and confirm the right starting point before enrolling.

A-56, Sector-64, Noida, Uttar Pradesh – 201301