EdX

Serverless Data Processing with Dataflow: Foundations (edX)

Offered by Google Cloud,
Serverless Data Processing with Dataflow: Foundations (edX)

This course is part 1 of a 3-course series on Serverless Data Processing with Dataflow. This course is part 1 of a 3-course series on Serverless Data Processing with Dataflow. In this first course, we start with a refresher of what Apache Beam is and its relationship with Dataflow.

Class Deals by MOOC List - Click here and see EdX's Active Discounts, Deals, and Promo Codes.

Next, we talk about the Apache Beam vision and the benefits of the Beam Portability framework. The Beam Portability framework achieves the vision that a developer can use their favorite programming language with their preferred execution backend. We then show you how Dataflow allows you to separate compute and storage while saving money, and how identity, access, and management tools interact with your Dataflow pipelines. Lastly, we look at how to implement the right security model for your use case on Dataflow.
This course is part of the Google Cloud Data Engineer Learning Path Professional Certificate.

What you'll learn

  • Demonstrate how Apache Beam and Cloud Dataflow work together to fulfill your organization’s data processing needs
  • Summarize the benefits of the Beam Portability Framework and enable it for your Dataflow pipelines
  • Enable Shuffle & Streaming Engine for batch & streaming pipelines respectively for maximum performance
  • Enable Flexible Resource Scheduling for more cost efficient performance
  • Select the right combination of IAM permissions for your Dataflow job
  • Implement best practices for a secure data processing environment

Syllabus

  1. Introduction

This module covers the course outline and does a quick refresh on the Apache Beam programming model and Google’s Dataflow managed service.

  1. Beam Portability

In this module we are going to learn about four sections, Beam Portability, Runner v2, Container Environments, and Cross-Language Transforms.

  1. Separating Compute and Storage with Dataflow

IIn this module we discuss how to separate compute and storage with Dataflow. This module contains four sections Dataflow, Dataflow Shuffle Service, Dataflow Streaming Engine, Flexible Resource Scheduling.

  1. IAM, Quotas, and Permissions

In this module, we talk about the different IAM roles, quotas, and permissions required to run Dataflow.

  1. Security

In this module, we will look at how to implement the right security model for your use case on Dataflow.

  1. Summary

In this course, we started with the refresher of what Apache Beam is, and its relationship with Dataflow.

Go to Class
MOOC List is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Related Courses

Data Wrangling with Python Project (Coursera) Coursera
University of Colorado Boulder

Data Wrangling with Python Project (Coursera)

The "Data Wrangling Project" course provides students with an opportunity to apply the knowledge gained throughout the specialization in a real-life data wrangling project of their interest. Participants will follow the data wrangling pipeline step by step, from identifying data sources to processing and integrating data, to achieve a fine dataset ready for analysis. This course enables students to gain hands-on experience in the data wrangling process and prepares them to handle complex data challenges in real-world scenarios.

Aug 31st 2026
5-12 Weeks
Data Processing and Analysis with Excel (edX) EdX
Rochester Institute of Technology,RITx

Data Processing and Analysis with Excel (edX)

Learn to use Excel to organize and clean data so it can be manipulated and analyzed. In this course, you will learn how to organize your data within the Microsoft Office Excel software tool. Once organized, we will discuss data cleaning. You will learn how to identify outliers and anomalies in the data, and how to identify and change data-types. Together we will develop a data analysis plan, after which we will apply analysis methods and tools, including exploratory analysis, evaluation of results, and comparison with other findings.

Self Paced
Self-Paced
Introduction to Serverless on Kubernetes (edX) EdX
Linux Foundation,LinuxFoundationX

Introduction to Serverless on Kubernetes (edX)

Learn how to build serverless functions that can be run on any cloud, without being restricted by limits on the execution duration, languages available, or the size of your code. With the advent of systems like AWS Lambda, the term serverless gained much popularity. However, many people are still unsure what it is for, and how it can help them build applications faster than traditional approaches. Other potential users are turned off by the arbitrary limits and lock-in of cloud-based serverless products.

Self Paced
Self-Paced
Serverless Data Processing with Dataflow: Foundations (Coursera) Coursera
Google Cloud

Serverless Data Processing with Dataflow: Foundations (Coursera)

This course is part 1 of a 3-course series on Serverless Data Processing with Dataflow. In this first course, we start with a refresher of what Apache Beam is and its relationship with Dataflow. Next, we talk about the Apache Beam vision and the benefits of the Beam Portability framework. The Beam Portability framework achieves the vision that a developer can use their favorite programming language with their preferred execution backend.

Aug 24th 2026
2 Weeks
Google Cloud Big Data and Machine Learning Fundamentals (edX) EdX
Google Cloud

Google Cloud Big Data and Machine Learning Fundamentals (edX)

Data Analysts, Data Engineers, Data Scientists, and ML Engineers who are getting started with Google Cloud. This course introduces the Google Cloud big data and machine learning products and services that support the data-to-AI lifecycle. It explores the processes, challenges, and benefits of building a big data pipeline and machine learning models with Vertex AI on Google Cloud.

Self Paced
Self-Paced
Data Processing with Azure (Coursera) Coursera
LearnQuest

Data Processing with Azure (Coursera)

This Azure training course is designed to equip students with the knowledge need to process, store and analyze data for making informed business decisions. Through this Azure course, the student will understand what big data is along with the importance of big data analytics, which will improve the students mathematical and programming skills. Students will learn the most effective method of using essential analytical tools such as Python, R, and Apache Spark.

Aug 24th 2026
3 Weeks