EdX

AI skills for Engineers: Supervised Machine Learning (edX)

AI skills for Engineers: Supervised Machine Learning (edX)

Learn the fundamentals of machine learning to help you correctly apply various classification and regression machine learning algorithms to real-life problems using the Python toolbox scikit-learn. Machine learning classification and regression techniques have potential uses in various engineering disciplines. These machine learning models allow you to make predictions for a category (classification) or for a number (regression) given sensor data, and can be used in, for example, predicting properties of objects (such as their weight or shape).

Class Deals by MOOC List - Click here and see EdX's Active Discounts, Deals, and Promo Codes.

Using hands-on and interactive exercises you will get insight into:
Machine learning and its variants, such as supervised learning, semi-supervised learning, unsupervised learning and reinforcement learning.
Regression techniques such as linear regression, K-nearest neighbor regression, how to deal with outliers and evaluation metrics such as the mean squared error (MSE) and mean absolute error (MAE).
Classification techniques such as the histogram method, the nearest mean (or nearest medoid) method and the nearest neighbor classifier. We cover the classification setting and important concepts such as the Bayes classifier and the Bayes error, the optimal classifier in theory.
Training models using (stochastic) gradient descent and its variants, we learn how to tune this optimizer, and how to use it to construct a logistic regression classification model.
Overfitting means a classifier works well on a training set but not on unseen test data. We discuss how to build complex non-linear models, and we analyze how we can understand overfitting using the bias-variance decomposition and the curse of dimensionality. Finally, we discuss how to evaluate fairly and tune machine learning models and estimate how much data they need for an sufficient performance.
Regularization methods can help to mitigate overfitting. We discuss two regularization techniques for estimating the linear regression coefficients: ridge regression and LASSO. The latter can also be used for variable selection.
Classifier evaluation metrics such as the ROC curve and confusion matrix can give more insight into the performance of classifiers. We also discuss what constitutes a “good” accuracy; this is given by so-called dummy-classifiers which are naïve baselines.
Support Vector Machines (SVMs) are more advanced classification models that can provide good performance even in high-dimensional spaces and with little data. We discuss their different variants such as the soft-margin SVM, the hard-margin SVM and the nonlinear kernel SVM.
Decision Trees are simple models that can easily be understood by lay people. They are easy to use and visualize, and instead of a black box they can be easily understood as an interpretable white box model, making them suitable for various applications.
The lectures feature a unique combination of videos mixed with hands-on interaction with machine learning algorithms to stimulate a deeper understanding. In the exercises you apply the algorithms in Python using scikit-learn and in the final project you will further deepen your understanding of the various concepts by building and tuning a machine learning pipeline from start to finish.
This course is part of the AI Skills: Basic and Advanced Techniques in Machine Learning Professional Certificate.

What you'll learn

  • Apply common operations (pre-processing, plotting, etc.) to datasets using Python.
  • Explain the concept of supervised, semi-supervised, unsupervised machine learning and reinforcement learning.
  • Explain how various supervised learning models work and recognize their limitations.
  • Analyze which factors impact the performance of learning algorithms.
  • Apply learning algorithms to datasets using Python and Scikit-learn and evaluate their performance.
  • Optimize a machine learning pipeline using Python and Scikit-learn.

Syllabus

Week 1: Introduction & Regression
This week is an introduction to the course with an overview of the topics. We give a brief introduction to machine learning and its different variants, and we will make a gentle start with regression. In the regression setting, a machine learning model will need to predict a number.
In the introductory part we cover:

  • Why use machine learning?
  • Machine learning basics and terminology
  • The biggest challenge in machine learning
  • Machine learning frameworks: supervised, semi-supervised, unsupervised and reinforcement learning

In the regression part we cover:

  • The regression setting and its assumptions
  • The mean squared error (MSE) and mean absolute error (MAE)
  • Outliers in regression
  • Linear regression and K-nearest neighbour regression

Week 2: Classification & Training Models
This week we discuss the classification setting and how to train models using gradient descent. In the classification setting, a machine learning model will need to predict a category or class. Gradient descent is an iterative procedure to train models, such as logistic regression and neural networks.
In the classification part we cover:

  • Terminology and basics of classification
  • Building classifiers using histograms, nearest mean (nearest medoid) classifier, K-nearest neighbour (KNN) classifier
  • The Bayes classifier and the Bayes error
  • How to use the KNN classifier in practice

In the training models part, we cover:

  • The basics of gradient descent
  • The three variants of gradient descent: batch, mini-batch and stochastic gradient descent (SGD)
  • How to tune gradient descent
  • The basics of logistic regression

Week 3: Overfitting & Regularization
This week focuses on overfitting and regularization. Overfitting is the problem where a machine learning algorithm performs well on the training set but does not perform well on new and unseen data. Regularization covers various techniques that aim to solve this problem.
In the overfitting part we cover:

  • How to use linear models for nonlinear tasks?
  • The bias-variance trade-off and the curse of dimensionality
  • How to use learning curves to estimate the amount of data needed
  • Cross validation, model selection and hyperparameter tuning

In the regularization part we cover:

  • Ridge regression
  • LASSO regularization and how it’s used for variable selection

Week 4: Classifier Evaluation & Support Vector Machines
Classifier evaluation delves deeper into the various evaluation metrics for classifiers. In the second part we cover the support vector machine classifier. The support vector machine is a well-known more advanced classification model.
In the classifier evaluation part, we cover:

  • What a “good” accuracy means (e.g., naïve baselines/dummy classifiers)
  • The confusion matrix (false positive, false negative, costs)
  • ROC-curves

In the support vector machine part (SVM) we cover:

  • Basics of the SVM, the margin and the hard-margin SVM
  • The soft-margin SVM
  • Kernels

Week 5: Decision Trees & Final Project
In this last week we discuss decision trees and you will work on a more in-depth final project. Decision trees are simple and interpretable models that are very user-friendly. The final project will involve building a machine learning pipeline, including hyperparameter tuning and a careful and fair evaluation, to solve a small practical application (MNIST).
In the decision tree part, we cover:

  • Basics of decision trees and their terminology
  • How to train decision trees with CART
  • Overfitting and other pros and cons of decision trees

Week 6: Wrap up
In this week, there will be extra time for questions regarding the earlier weeks and to discuss the final project.

Go to Class
MOOC List is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Related Courses

Data Science: Capstone (edX) EdX
HarvardX,Harvard University

Data Science: Capstone (edX)

Show what you’ve learned from the Professional Certificate Program in Data Science. To become an expert data scientist you need practice and experience. By completing this capstone project you will get an opportunity to apply the knowledge and skills in R data analysis that you have gained throughout the series. This final project will test your skills in data visualization, probability, inference and modeling, data wrangling, data organization, regression, and machine learning.

Self Paced
Self-Paced
Dynamic Programming: Applications In Machine Learning and Genomics (edX) EdX
University of California, San Diego,UC San DiegoX

Dynamic Programming: Applications In Machine Learning and Genomics (edX)

Learn how dynamic programming and Hidden Markov Models can be used to compare genetic strings and uncover evolution. If you look at two genes that serve the same purpose in two different species, how can you rigorously compare these genes in order to see how they have evolved away from each other?

Self Paced
Self-Paced
Digital Marketing Analytics: Tools and Techniques (edX) EdX
University of Maryland, College Park,University System of Maryland - USM,USMx,UMD

Digital Marketing Analytics: Tools and Techniques (edX)

Learn how to leverage leading tools and approaches to digital marketing data analysis. Dive into SEO and SEM strategies including web analytics, machine learning and AI/Big Data applications to strengthen your digital marketing efforts and leverage your resources most effectively.

Self Paced
Self-Paced
Probability and Statistics in Data Science using Python (edX) EdX
University of California, San Diego,UC San DiegoX

Probability and Statistics in Data Science using Python (edX)

Using Python, learn statistical and probabilistic approaches to understand and gain insights from data. The job of a data scientist is to glean knowledge from complex and noisy datasets. Reasoning about uncertainty is inherent in the analysis of noisy data. Probability and Statistics provide the mathematical foundation for such reasoning.

Self Paced
Self-Paced
Impact Evaluation Methods with Applications in Low- and Middle-Income Countries (edX) EdX
Georgetown University,GeorgetownX

Impact Evaluation Methods with Applications in Low- and Middle-Income Countries (edX)

Economic development is about making a difference in the lives of the poor, through interventions in the health, education, microfinance, transport, agriculture, and other sectors. This course will provide you with the experimental and statistical tools you need to measure the impacts you are hoping to see. How do you design and conduct a randomized control trial, and how do you evaluate the data using regression techniques?

Self Paced
Self-Paced
AI in Practice: Preparing for AI (edX) EdX
Delft University of Technology,DelftX

AI in Practice: Preparing for AI (edX)

Learn to recognize and understand the implications of Artificial Intelligence for organizations, and the importance of compliance and ethics when AI is applied in practice. This course is not about difficult algorithms and complex programming; it is a course for anyone interested in learning about the benefits and implications of AI when applied in practical settings.

Self Paced
Self-Paced
Tech for Good: The Role of ICT in Achieving the SDGs (edX) EdX
SDGAcademyX,SDG Academy

Tech for Good: The Role of ICT in Achieving the SDGs (edX)

What opportunities and challenges do digital technologies present for the development of our society? Tech for Good was developed by UNESCO and Cetic, the Brazilian Network Information Center’s Regional Center for Studies on the Development of the Information Society. It brings together thought leaders and changemakers in the fields of information and communication technologies (ICT) and sustainable development to show how digital technologies are empowering billions of people around the world by providing access to education, healthcare, banking, and government services; and how “big data” is being used to inform smarter, evidence-based policies to improve people’s lives in fundamental ways.

Self Paced
Self-Paced
Computing for Data Analysis (edX) EdX
Georgia Institute of Technology,GTx

Computing for Data Analysis (edX)

A hands-on introduction to basic programming principles and practice relevant to modern data analysis, data mining, and machine learning. The modern data analysis pipeline involves collection, preprocessing, storage, analysis, and interactive visualization of data. In the course, you’ll see how computing and mathematics come together.

Aug 24th 2026
13-24 Weeks