Analyze Datasets and Train ML Models using AutoML (Coursera)

Offered by DeepLearning.AI, AWS,
Analyze Datasets and Train ML Models using AutoML (Coursera)

In the first course of the Practical Data Science Specialization, you will learn foundational concepts for exploratory data analysis (EDA), automated machine learning (AutoML), and text classification algorithms. With Amazon SageMaker Clarify and Amazon SageMaker Data Wrangler, you will analyze a dataset for statistical bias, transform the dataset into machine-readable features, and select the most important features to train a multi-class text classifier.

Class Deals by MOOC List - Click here and see Coursera's Active Discounts, Deals, and Promo Codes.

You will then perform automated machine learning (AutoML) to automatically train, tune, and deploy the best text-classification algorithm for the given dataset using Amazon SageMaker Autopilot. Next, you will work with Amazon SageMaker BlazingText, a highly optimized and scalable implementation of the popular FastText algorithm, to train a text classifier with very little code.
Practical data science is geared towards handling massive datasets that do not fit in your local hardware and could originate from multiple sources. One of the biggest benefits of developing and running data science projects in the cloud is the agility and elasticity that the cloud offers to scale up and out at a minimum cost.
The Practical Data Science Specialization helps you develop the practical skills to effectively deploy your data science projects and overcome challenges at each step of the ML workflow using Amazon SageMaker. This Specialization is designed for data-focused developers, scientists, and analysts familiar with the Python and SQL programming languages and want to learn how to build, train, and deploy scalable, end-to-end ML pipelines - both automated and human-in-the-loop - in the AWS cloud.
Course 1 of 3 in the Practical Data Science Specialization.

What You Will Learn
Prepare data, detect statistical data biases, and perform feature engineering at scale to train models with pre-built algorithms.

Syllabus

WEEK 1
Explore the Use Case and Analyze the Dataset
Ingest, explore, and visualize a product review data set for multi-class text classification.

WEEK 2
Data Bias and Feature Importance
Determine the most important features in a data set and detect statistical biases.

WEEK 3
Use Automated Machine Learning to train a Text Classifier
Inspect and compare models generated with automated machine learning (AutoML).

WEEK 4
Built-in algorithms
Train a text classifier with BlazingText and deploy the classifier as a real-time inference endpoint to serve predictions.

Go to Class
MOOC List is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Related Courses

Validity and Bias in Epidemiology (Coursera) Coursera
Imperial College Business School

Validity and Bias in Epidemiology (Coursera)

Epidemiological studies can provide valuable insights about the frequency of a disease, its potential causes and the effectiveness of available treatments. Selecting an appropriate study design can take you a long way when trying to answer such a question. However, this is by no means enough. A study can yield biased results for many different reasons. This course offers an introduction to some of these factors and provides guidance on how to deal with bias in epidemiological research.

Oct 12th 2026
4 Weeks
AI Fundamentals for Non-Data Scientists (Coursera) Coursera
University of Pennsylvania

AI Fundamentals for Non-Data Scientists (Coursera)

In this course, you will go in-depth to discover how Machine Learning is used to handle and interpret Big Data. You will get a detailed look at the various ways and methods to create algorithms to incorporate into your business with such tools as Teachable Machine and TensorFlow. You will also learn different ML methods, Deep Learning, as well as the limitations but also how to drive accuracy and use the best training data for your algorithms.

Oct 12th 2026
4 Weeks
Build Better Generative Adversarial Networks (GANs) (Coursera) Coursera
DeepLearning.AI

Build Better Generative Adversarial Networks (GANs) (Coursera)

In this course, you will: Assess the challenges of evaluating GANs and compare different generative models; Use the Fréchet Inception Distance (FID) method to evaluate the fidelity and diversity of GANs; Identify sources of bias and the ways to detect it in GANs; Learn and implement the techniques associated with the state-of-the-art StyleGANs.

Oct 12th 2026
3 Weeks
Intellectual Humility: Science (Coursera) Coursera
University of Edinburgh

Intellectual Humility: Science (Coursera)

It’s clear that the world needs more intellectual humility. But how do we develop this virtue? And why do so many people still end up so arrogant? Do our own biases hold us back from becoming as intellectually humble as we could be—and are there some biases that actually make us more likely to be humble? Which cognitive dispositions and personality traits give people an edge at being more intellectually humble - and are they stable from birth, learned habits, or something in between? And what can contemporary research on the emotions tell us about encouraging intellectual humility in ourselves and others?

Oct 12th 2026
4 Weeks
Practical Steps for Building Fair AI Algorithms (Coursera) Coursera
Fred Hutchinson Cancer Center

Practical Steps for Building Fair AI Algorithms (Coursera)

Algorithms increasingly help make high-stakes decisions in healthcare, criminal justice, hiring, and other important areas. This makes it essential that these algorithms be fair, but recent years have shown the many ways algorithms can have biases by age, gender, nationality, race, and other attributes. This course will teach you ten practical principles for designing fair algorithms. It will emphasize real-world relevance via concrete takeaways from case studies of modern algorithms, including those in criminal justice, healthcare, and large language models like ChatGPT. You will come away with an understanding of the basic rules to follow when trying to design fair algorithms, and assess algorithms for fairness.

Oct 12th 2026
4 Weeks
Python for Data Science (Coursera) Coursera
Fractal Analytics

Python for Data Science (Coursera)

Understanding the importance of Python as a data science tool is crucial for anyone aspiring to leverage data effectively. This course is designed to equip you with the essential skills and knowledge needed to thrive in the field of data science. This course teaches the vital skills to manipulate data using pandas, perform statistical analyses, and create impactful visualizations. Learn to solve real-world business problems and prepare data for machine learning applications.

Oct 12th 2026
5-12 Weeks
Generative AI's Applications in Marketing Analytics (Coursera) Coursera
Edureka

Generative AI's Applications in Marketing Analytics (Coursera)

Welcome to the "Generative AI's Applications in Marketing Analytics" short course, a journey into the innovative fusion of Generative AI and marketing analytics. Throughout this course, you'll embark on an exploration of how Generative AI can revolutionize marketing analytics, offering powerful tools to drive insights, predictions, and strategies.

Oct 26th 2026
1 Week
Inclusive Leadership: The Power of Workplace Diversity (Coursera) Coursera
University of Colorado System

Inclusive Leadership: The Power of Workplace Diversity (Coursera)

This course will equip and empower you to be a highly inclusive leader. You will learn principles, perspectives and practices that help to reap the power of workplace diversity. Workplace diversity has expanded beyond traditional demographics such as gender, race, and ethnicity. Those categories always will matter. In today’s workplace, diversity also encompasses nationality, religious background, sexual orientation, age, ability, experience, and diversity of thought, among other characteristics. Leaders will need to consider all of these aspects of diversity – among all stakeholders. That includes current and prospective employees, customers, clients , students, peers, and anyone else with whom a leader interacts.

Oct 19th 2026
4 Weeks
Big Data Analysis with Scala and Spark (Scala 2 version) (Coursera) Coursera
École Polytechnique Fédérale de Lausanne

Big Data Analysis with Scala and Spark (Scala 2 version) (Coursera)

Manipulating big data distributed over a cluster using functional concepts is rampant in industry, and is arguably one of the first widespread industrial uses of functional ideas. This is evidenced by the popularity of MapReduce and Hadoop, and most recently Apache Spark, a fast, in-memory distributed collections framework written in Scala. In this course, we'll see how the data parallel paradigm can be extended to the distributed case, using Spark throughout.

Oct 12th 2026
4 Weeks
AI Workflow: Feature Engineering and Bias Detection (Coursera) Coursera
IBM

AI Workflow: Feature Engineering and Bias Detection (Coursera)

This is the third course in the IBM AI Enterprise Workflow Certification specialization. You are STRONGLY encouraged to complete these courses in order as they are not individual independent courses, but part of a workflow where each course builds on the previous ones. Course 3 introduces you to the next stage of the workflow for our hypothetical media company. In this stage of work you will learn best practices for feature engineering, handling class imbalances and detecting bias in the data.

Oct 12th 2026
2 Weeks
Data Science with R - Capstone Project (Coursera) Coursera
IBM

Data Science with R - Capstone Project (Coursera)

In this capstone course, you will apply various data science skills and techniques that you have learned as part of the previous courses in the IBM Data Science with R Specialization or IBM Data Analytics with Excel and R Professional Certificate. For this project, you will assume the role of a Data Scientist who has recently joined an organization and be presented with a challenge that requires data collection, analysis, basic hypothesis testing, visualization, and modeling to be performed on real-world datasets.

Oct 12th 2026
5-12 Weeks