Introduction to Bayesian Data Analysis (openHPI)

Introduction to Bayesian Data Analysis (openHPI)

Bayesian data analysis is increasingly becoming the tool of choice for many data-analysis problems. This free course on Bayesian data analysis will teach you basic ideas about random variables and probability distributions, Bayes' rule, and its application in simple data analysis problems. You will learn to use the R package brms (which is a front-end for the probabilistic programming language Stan). The focus will be on regression modeling, culminating in a brief introduction to hierarchical models (otherwise known as mixed or multilevel models). This course is appropriate for anyone familiar with the programming language R and for anyone who has done some frequentist data analysis (e.g., linear modeling and/or linear mixed modeling) in the past.

Introduction: Why are Bayesian methods important for data analysts?
Here are some of the advantages of Bayesian methods over the standard frequentist approach used in data analysis:

  • Prior knowledge/expertise can be incorporated into the data analysis
  • Models can be flexibly specified to reflect the assumed generative process
  • The results of the analysis – the posterior distributions of the parameters of interest – have an intuitive interpretation
  • Hypothesis testing can be carried out in a more meaningful manner than the standard used null hypothesis significance testing

Prerequisites: Who is this course for?
We assume the following in this course:

  • Basic familiarity with the programming language R, openHPI offers a free R course for Beginners (in German)
  • Experience with data analysis using linear models
  • It is helpful (but not necessary) to have had some exposure to linear mixed models using the R library lme4
  • High-school mathematics (pre-calculus)
  • Some basic concepts from probability theory (sum and product rule, conditional probability)

This course is not appropriate for participants who don't know R programming and who have no experience at all with data analysis.

Course outcomes: What will you learn from this course?

  • Some basic ideas relating to random variables
  • Some fundamental properties of probability distributions
  • Application of Bayes' rule in data analysis
  • The concept of likelihood and its role in Bayesian statistical modeling
  • Bayesian regression models using brms (a front-end for Stan)
  • How to visualize and interpret prior and posterior distributions
  • How to generate prior and posterior predictive distributions for evaluating models
  • How to interpret the results of simple regression models

After completing this course, you will be in a good position to learn how to use more advanced Bayesian methods, such as hierarchical models, finite mixture models, multinomial processing tree models, measurement error models, etc.

What you'll learn

  • Bayesian statistics
  • Data analysis
  • Bayesian regression models using brms

Course contents

Week 0 - Initial Setup:
Installing R and RStudio, rstan, brms, and other necessary packages in R; Setting up R markdown for reproducible data analyses.

Week 1 - Introduction:
Learn the foundational ideas about random variables and probability distributions; Reading: Chapter 1 of the textbook (excluding the section on bivariate distributions).

Week 2 - Bayesian data analysis:
Understand Bayes' rule, derive the posterior using Bayes' rule; visualize the prior, likelihood, and posterior; distinguish the relationship between the prior, likelihood, and posterior; incorporate prior knowledge into the analysis; Reading: Chapter 2.

Week 3 - Computational Bayesian data analysis:
Derive the posterior through sampling; perform simple regression modeling of a simple button-pressing task using Stan/brms; do prior predictive distributions, sensitivity analysis, and different classes of prior; do posterior predictive distributions; derive the log-normal likelihood; Reading: Chapter 3.

Week 4 - Bayesian regression and hierarchical models:
Perform simple linear regressions using the normal and binomial likelihoods to answer the following research questions: (i) Does attentional load affect pupil size? (ii) Does trial id affect response times? (iii) Does set size affect recall accuracy? Take a brief look-ahead at linear mixed models; Reading: Chapter 4 and up to section 5.3 of chapter 5.

Final Exam:
Final Exam

Go to Class
MOOC List is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Related Courses

Basic Statistics (Coursera) Coursera
University of Amsterdam

Basic Statistics (Coursera)

Understanding statistics is essential to understand research in the social and behavioral sciences. In this course you will learn the basics of statistics; not just how to calculate them, but also how to evaluate them. This course will also prepare you for the next course in the specialization - the course Inferential Statistics. In the first part of the course we will discuss methods of descriptive statistics. You will learn what cases and variables are and how you can compute measures of central tendency (mean, median and mode) and dispersion (standard deviation and variance). Next, we discuss how to assess relationships between variables, and we introduce the concepts correlation and regression.

Sep 28th 2026
5-12 Weeks
Data Science Bootcamp (openHPI) OpenHPI
Hasso-Plattner-Institut

Data Science Bootcamp (openHPI)

The ultimate goal of the bootcamp is to cultivate strong data science skills with an emphasis on machine learning techniques to satisfactorily meet and exceed the requests of the Data science world. In the process, we will develop good habits for operating independently as data scientists and for operating as members of productive data science teams.

Jun 7th 2023
4 Weeks
HI-FIVE: Health Informatics For Innovation, Value & Enrichment (Administrative/IT Perspective) (Coursera) Coursera
Columbia University

HI-FIVE: Health Informatics For Innovation, Value & Enrichment (Administrative/IT Perspective) (Coursera)

HI-FIVE (Health Informatics For Innovation, Value & Enrichment) Training is an approximately 10-hour online course designed by Columbia University in 2016, with sponsorship from the Office of the National Coordinator for Health Information Technology (ONC). The training is role-based and uses case scenarios. No additional hardware or software are required for this course. Our nation’s healthcare system is changing at a rapid pace.

Sep 28th 2026
4 Weeks
Statistical Thinking for Industrial Problem Solving, presented by JMP (Coursera) Coursera
SAS

Statistical Thinking for Industrial Problem Solving, presented by JMP (Coursera)

Statistical Thinking for Industrial Problem Solving is an applied statistics course for scientists and engineers offered by JMP, a division of SAS. By completing this course, students will understand the importance of statistical thinking, and will be able to use data and basic statistical methods to solve many real-world problems.

Sep 21st 2026
5-12 Weeks
Statistics and Data Analysis with Excel, Part 1 (Coursera) Coursera
University of Colorado Boulder

Statistics and Data Analysis with Excel, Part 1 (Coursera)

Designed for students with no prior statistics knowledge, this course will provide a foundation for further study in data science, data analytics, or machine learning. Topics include descriptive statistics, probability, and discrete and continuous probability distributions. Assignments are conducted in Microsoft Excel (Windows or Mac versions). Designed to be taken with the follow-up course, “Statistics and Data Analysis with Excel, Part 2.”

Sep 28th 2026
5-12 Weeks
Probabilistic Graphical Models 3: Learning (Coursera) Coursera
Stanford University

Probabilistic Graphical Models 3: Learning (Coursera)

Probabilistic graphical models (PGMs) are a rich framework for encoding probability distributions over complex domains: joint (multivariate) distributions over large numbers of random variables that interact with each other. These representations sit at the intersection of statistics and computer science, relying on concepts from probability theory, graph algorithms, machine learning, and more. They are the basis for the state-of-the-art methods in a wide variety of applications, such as medical diagnosis, image understanding, speech recognition, natural language processing, and many, many more. They are also a foundational tool in formulating many machine learning problems.

Sep 28th 2026
5-12 Weeks
Doing Economics: Measuring Climate Change (Coursera) Coursera
University of London,University College London,CORE

Doing Economics: Measuring Climate Change (Coursera)

This course will give you practical experience in working with real-world data, with applications to important policy issues in today’s society. Each week, you will learn specific data handling skills in Excel and use these techniques to analyse climate change data, with appropriate readings to provide background information on the data you are working with. You will also learn about the consequences of climate change and how governments can address this issue.

Sep 21st 2026
4 Weeks
Data Visualization for Genome Biology (Coursera) Coursera
University of Toronto

Data Visualization for Genome Biology (Coursera)

The past decade has seen a vast increase in the amount of data available to biologists, driven by the dramatic decrease in cost and concomitant rise in throughput of various next-generation sequencing technologies, such that a project unimaginable 10 years ago was recently proposed, the Earth BioGenomes Project, which aims to sequence the genomes of all eukaryotic species on the planet within the next 10 years. So while data are no longer limiting, accessing and interpreting those data has become a bottleneck. One important aspect of interpreting data is data visualization. This course introduces theoretical topics in data visualization through mini-lectures, and applied aspects in the form of hands-on labs.

Sep 21st 2026
5-12 Weeks
Bioinformatic Methods II (Coursera) Coursera
University of Toronto

Bioinformatic Methods II (Coursera)

Large-scale biology projects such as the sequencing of the human genome and gene expression surveys using RNA-seq, microarrays and other technologies have created a wealth of data for biologists. However, the challenge facing scientists is analyzing and even accessing these data to extract useful information pertaining to the system being studied. This course focuses on employing existing bioinformatic resources – mainly web-based programs and databases – to access the wealth of data to answer questions relevant to the average biologist, and is highly hands-on.

Sep 28th 2026
5-12 Weeks
Data Science for Business Innovation (Coursera) Coursera
Politecnico di Milano,EIT Digital

Data Science for Business Innovation (Coursera)

The course is a compendium of the must-have expertise in data science for executive and middle-management to foster data-driven innovation. It consists of introductory lectures spanning big data, machine learning, data valorization and communication. Topics cover the essential concepts and intuitions on data needs, data analysis, machine learning methods, respective pros and cons, and practical applicability issues.

Sep 21st 2026
4 Weeks