Intro to Data Science (Udacity)

Offered by Udacity,
Intro to Data Science (Udacity)

Learn what it takes to become a data scientist. The Introduction to Data Science class will survey the foundational topics in data science, namely: Data Manipulation; Data Analysis with Statistics and Machine Learning; Data Communication with Information Visualization; Data at Scale -- Working with Big Data.

Class Deals by MOOC List - Click here and see Udacity's Active Discounts, Deals, and Promo Codes.

The class will focus on breadth and present the topics briefly instead of focusing on a single topic in depth. This will give you the opportunity to sample and apply the basic techniques of data science.
This course is also a part of our Data Analyst Nanodegree.

What You Will Learn

Lesson 1
Introduction to Data Science

  • Pi-Chaun (Data Scientist @ Google): What is Data Science?
  • Gabor (Data Scientist @ Twitter): What is Data Science?
  • Problems solved by data science.

Lesson 2
Data Wrangling

  • What is Data Wrangling?
  • Acquiring data.
  • Common data formats.

Lesson 3
Data Analysis

  • Statistical rigor.
  • Kurt (Data Scientist @ Twitter) - Why is Stats Useful?
  • Introduction to normal distribution.

Lesson 4
Data Visualization

  • Effective information visualization.
  • An analysis of Napoleon's invasion of Russia!
  • Don (Principal Data Scientist @ AT&T): Communicating Findings.

Lesson 5
MapReduce

  • Introduction to Big Data and MapReduce.
  • Learn the basics of MapReduce.
  • Mapper.

Prerequisites and Requirements
The ideal students for this class are prepared individuals who have:
Strong interest in data science
Background in intro level statistics
Python programming experience
Or understanding of programming concepts such as variables, functions, loops, and basic python data structures like lists and dictionaries
If you need to brush up on your programming, we highly recommend Introduction to Computer Science: Building a Search Engine. If you need a refresher on statistics, enroll in Intro to Descriptive Statistics and Intro to Inferential Statisitics. All three are on Udacity!

Why Take This Course
You will have an opportunity to work through a data science project end to end, from analyzing a dataset to visualizing and communicating your data analysis.
Through working on the class project, you will be exposed to and understand the skills that are needed to become a data scientist yourself.

Go to Class
MOOC List is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Related Courses

Machine Learning (Udacity) Udacity
Georgia Institute of Technology,Udacity

Machine Learning (Udacity)

Supervised, Unsupervised & Reinforcement. Machine Learning is a graduate-level course covering the area of Artificial Intelligence concerned with computer programs that modify and improve their performance through experiences. The first part of the course covers Supervised Learning, a machine learning task that makes it possible for your phone to recognize your voice, your email to filter spam, and for computers to learn a bunch of other cool stuff. In part two, you will learn about Unsupervised Learning. Ever wonder how Netflix can predict what movies you'll like? Or how Amazon knows what you want to buy before you do? Such answers can be found in this section!

Self Paced
Self-Paced
Understanding China, 1700-2000: A Data Analytic Approach, Part 1 (Coursera) Coursera
The Hong Kong University of Science and Technology - HKUST

Understanding China, 1700-2000: A Data Analytic Approach, Part 1 (Coursera)

The purpose of this course is to summarize new directions in Chinese history and social science produced by the creation and analysis of big historical datasets based on newly opened Chinese archival holdings, and to organize this knowledge in a framework that encourages learning about China in comparative perspective. Our course demonstrates how a new scholarship of discovery is redefining what is singular about modern China and modern Chinese history.

Sep 28th 2026
5-12 Weeks
Bioinformatic Methods II (Coursera) Coursera
University of Toronto

Bioinformatic Methods II (Coursera)

Large-scale biology projects such as the sequencing of the human genome and gene expression surveys using RNA-seq, microarrays and other technologies have created a wealth of data for biologists. However, the challenge facing scientists is analyzing and even accessing these data to extract useful information pertaining to the system being studied. This course focuses on employing existing bioinformatic resources – mainly web-based programs and databases – to access the wealth of data to answer questions relevant to the average biologist, and is highly hands-on.

Sep 28th 2026
5-12 Weeks
Encoder-Decoder Architecture with Google Cloud (Udacity) Udacity
Udacity,Google Cloud

Encoder-Decoder Architecture with Google Cloud (Udacity)

Learn about the main components of the encoder-decoder architecture and how to train and serve these models. This course gives you a synopsis of the encoder-decoder architecture, which is a powerful and prevalent machine learning architecture for sequence-to-sequence tasks such as machine translation, text summarization, and question answering.

Self Paced
Self-Paced
Introduction to Machine Learning using Microsoft Azure (Udacity) Udacity
Udacity,Microsoft Azure

Introduction to Machine Learning using Microsoft Azure (Udacity)

Gain a high-level introduction to the field of machine learning and prepare to use Azure Machine Learning Studio to train machine learning models. Plus, learn how to perform a variety of tasks on Azure Machine Learning labs — from data import, transformation and management to training, validating and evaluating models. Access to the Azure Machine Learning Labs will close after a predetermined number of students have completed the course.

Self Paced
Self-Paced
Data Visualization and D3.js (Udacity) Udacity
Udacity,Zipfian Academy

Data Visualization and D3.js (Udacity)

Communicating with Data. Learn the fundamentals of data visualization and practice communicating with data. This course covers how to apply design principles, human perception, color theory, and effective storytelling to data visualization. If you present data to others, aspire to be an analyst or data scientist, or if you’d like to become more technical with visualization tools, then you can grow your skills with this course.

Self Paced
Self-Paced
Segmentation and Clustering (Udacity) Udacity
Udacity

Segmentation and Clustering (Udacity)

Use machine learning to create segments. The Segmentation and Clustering course provides students with the foundational knowledge to build and apply clustering models to develop more sophisticated segmentation in business contexts. In this course, you'll learn how to use an advanced analytical method called clustering to create useful segments for business contexts, whether its stores, customers, geographies, etc. You'll learn this through improving your fluency in Alteryx, a data analytics tool that enables you prepare, blend, and analyze data quickly.

Self Paced
Self-Paced
Real-Time Analytics with Apache Storm (Udacity) Udacity
Udacity,Twitter

Real-Time Analytics with Apache Storm (Udacity)

The world is trending in real time! Learn from Twitter to scalably process tweets, or any big data stream, in real-time to drive d3 visualizations using Apache Storm, the "Hadoop of Real Time." Storm is free, open source, and fun to use! Learn from Karthik Ramasamy, about the distributed, fault-tolerant, and flexible technology used to power Twitter’s real-time data flow pipeline. Twitter open sourced Storm in 2011, and it graduated to a top-level Apache project in September, 2014.

Self Paced
Self-Paced
Big Data Analytics in Healthcare (Udacity) Udacity
Georgia Institute of Technology,Udacity

Big Data Analytics in Healthcare (Udacity)

Data science plays an important role in many industries. In facing massive amount of heterogeneous data, scalable machine learning and data mining algorithms and systems become extremely important for data scientists. The growth of volume, complexity and speed in data drives the need for scalable data analytic algorithms and systems. In this course, we study such algorithms and systems in the context of healthcare applications.

Self Paced
Self-Paced