This course is part of Data Engineering Foundations.
This comprehensive course covers the fundamentals of modern data engineering platforms including Hadoop, Spark, and Snowflake. Students learn to optimize and manage data pipelines, execute analytics using Databricks, and implement machine learning workflows with MLflow. The curriculum includes hands-on experience with PySpark for data science, along with best practices in DataOps and DevOps methodologies. Through practical exercises and real-world scenarios, participants develop skills in building scalable, efficient data solutions.
Instructors:
English
English
What you'll learn
Master essential data engineering platforms like Hadoop, Spark, and Snowflake
Execute advanced data analytics using Databricks and PySpark
Implement end-to-end machine learning workflows with MLflow
Optimize and manage scalable data pipelines
Apply DataOps and DevOps methodologies to data engineering projects
Build efficient ETL processes using modern tools
Skills you'll gain
This course includes:
PreRecorded video
Graded assignments, exams
Access on Mobile, Tablet, Desktop
Limited Access access
Shareable certificate
Closed caption
Get a Completion Certificate
Share your certificate with prospective employers and your professional network on LinkedIn.
Created by
Provided by

Top companies offer this course to their employees
Top companies provide this course to enhance their employees' skills, ensuring they excel in handling complex projects and drive organizational success.





There are 4 modules in this course
This comprehensive data engineering course focuses on mastering essential platforms and tools for modern data processing. Beginning with foundational concepts in Hadoop and Spark, students progress through advanced topics including Snowflake data warehousing, Databricks analytics, and MLflow implementation. The curriculum emphasizes practical skills in building data pipelines, executing analytics, and managing machine learning workflows, while incorporating industry best practices in DataOps and project management.
Overview and Introduction to PySpark
Module 1
Snowflake
Module 2
Azure Databricks and MLFlow
Module 3
DataOps and Operations Methodologies
Module 4
Fee Structure
Individual course purchase is not available - to enroll in this course with a certificate, you need to purchase the complete Professional Certificate Course. For enrollment and detailed fee structure, visit the following: Data Engineering Foundations
Instructors
Executive in Residence and Founder of Pragmatic AI Labs at Duke University
Noah Gift is the founder of Pragmatic AI Labs and serves as an Executive in Residence at Duke University, where he lectures in the Master of Interdisciplinary Data Science (MIDS) program. He specializes in designing and teaching graduate-level courses on machine learning, MLOps, artificial intelligence, and data science, while also consulting on machine learning and cloud architecture for students and faculty. A recognized expert in the field, Gift is a Python Software Foundation Fellow and an AWS Machine Learning Hero, holding multiple AWS certifications, including AWS Certified Solutions Architect and AWS Certified Machine Learning Specialist. He has authored several influential books, such as Practical MLOps, Python for DevOps, and Pragmatic AI, and has published over 100 technical articles across various platforms, including Forbes and O'Reilly. His extensive industry experience includes roles as CTO and Chief Data Scientist for notable companies like Disney Feature Animation, Sony Imageworks, and AT&T, contributing to major films like Avatar and Spider-Man 3. Gift's work has generated millions in revenue through product development on a global scale. He actively consults startups on machine learning and cloud architecture while leading initiatives to enhance data science education.
Senior Data Engineer and Educator at Duke University
Kennedy Behrman is a Senior Data Engineer at Duke University, where he also serves as an instructor for several online courses focused on data engineering and visualization. With decades of experience in Python and data management across various fields, including film, computing, and machine learning, he has established himself as a leading figure in the industry. Behrman has developed and taught courses such as "Data Visualization with Python" and "Linux and Bash for Data Engineering," equipping students with essential skills for the evolving data landscape. His expertise extends to big data processing technologies, where he covers platforms like Apache Spark and Snowflake. In addition to his teaching roles, Behrman has authored educational materials that contribute to the understanding of data science principles. His commitment to fostering learning and innovation in data engineering makes him a valuable asset to both Duke University and the broader academic community.
Testimonials
Testimonials and success stories are a testament to the quality of this program and its impact on your career and learning journey. Be the first to help others make an informed decision by sharing your review of the course.
Frequently asked questions
Below are some of the most commonly asked questions about this course. We aim to provide clear and concise answers to help you better understand the course content, structure, and any other relevant information. If you have any additional questions or if your question is not listed here, please don't hesitate to reach out to our support team for further assistance.





