EdX

Principles, Statistical and Computational Tools for Reproducible Science (edX)

Principles, Statistical and Computational Tools for Reproducible Science (edX)

Learn skills and tools that support data science and reproducible research, to ensure you can trust your own research results, reproduce them yourself, and communicate them to others. Today the principles and techniques of reproducible research are more important than ever, across diverse disciplines from astrophysics to political science. No one wants to do research that can’t be reproduced. Thus, this course is really for anyone who is doing any data intensive research. While many of us come from a biomedical background, this course is for a broad audience of data scientists.

Class Deals by MOOC List - Click here and see EdX's Active Discounts, Deals, and Promo Codes.

To meet the needs of the scientific community, this course will examine the fundamentals of methods and tools for reproducible research. Led by experienced faculty from the Harvard T.H. Chan School of Public Health, you will participate in six modules that will include several case studies that illustrate the significant impact of reproducible research methods on scientific discovery.

This course will appeal to students and professionals in biostatistics, computational biology, bioinformatics, and data science. The course content will blend video lectures, case studies, peer-to-peer engagements and use of computational tools and platforms (such as R/RStudio, and Git/Github), culminating in a final presentation of a final reproducible research project.
We’ll cover Fundamentals of Reproducible Science; Case Studies; Data Provenance; Statistical Methods for Reproducible Science; Computational Tools for Reproducible Science; and Reproducible Reporting Science. These concepts are intended to translate to fields throughout the data sciences: physical and life sciences, applied mathematics and statistics, and computing.
Consider this course a survey of best practices: we’d like to make you aware of pitfalls in reproducible data science, some failure - and success - stories in the past, and tools and design patterns that might help make it all easier. But ultimately it’ll be up to you to take the skills you learn from this course to create your own environment in which you can easily carry out reproducible research, and to encourage and integrate with similar environments for your collaborators and colleagues. We look forward to seeing you in this course and the research you do in the future!

What you'll learn

  • Understand a series of concepts, thought patterns, analysis paradigms, and computational and statistical tools, that together support data science and reproducible research.
  • Fundamentals of reproducible science using case studies that illustrate various practices
  • Key elements for ensuring data provenance and reproducible experimental design
  • Statistical methods for reproducible data analysis
  • Computational tools for reproducible data analysis and version control (Git/GitHub, Emacs/RStudio/Spyder), reproducible data (Data repositories/Dataverse) and reproducible dynamic report generation (Rmarkdown/R Notebook/Jupyter/Pandoc), and workflows.
  • How to develop new methods and tools for reproducible research and reporting
  • How to write your own reproducible paper.

Course Syllabus

Module 1: Introduction to Course Overview Introduction to faculty Project assignment
Module 2: Fundamentals of Reproducible Science Why reproducible research matters Definitions and concepts Factors affecting reproducibility
Module 3: Case Studies in Reproducible Research Potti 2006 Baggerly and Coombes 2007 Ioannidis 2009 Reproducible Reporting
Module 4: Data Provenance Project design Journal requirements and mechanisms Repositories Privacy and security
Module 5: Statistical Methods for Reproducible Science Prediction Models Coefficient of determination Brier score AUC Concordance in survival analysis Cross validation Bootstrap
Module 6: Computational Tools for Reproducible Science R and Rstudio Python Git and GitHub Creating a repository Data sources Dynamic report generation Workflows
Course Conclusion Final Project: Write a reproducible report that could be submitted at a peer review journal

Go to Class
MOOC List is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Related Courses

Data Analytics Basics for Everyone (edX) EdX
IBM

Data Analytics Basics for Everyone (edX)

Learn the fundamentals of Data Analytics and gain an understanding of the data ecosystem, the process and lifecycle of data analytics, career opportunities, and the different learning paths you can take to be a Data Analyst. In this course, you will learn about the various components of a modern data ecosystem and the role Data Analysts, Data Scientists, and Data Engineers play in this ecosystem.

Self Paced
Self-Paced
Introduction to Linear Models and Matrix Algebra (edX) EdX
HarvardX,Harvard University

Introduction to Linear Models and Matrix Algebra (edX)

Learn to use R programming to apply linear models to analyze data in life sciences. Matrix Algebra underlies many of the current tools for experimental design and the analysis of high-dimensional data. In this introductory data analysis course, we will use matrix algebra to represent the linear models that commonly used to model differences between experimental units. We perform statistical inference on these differences. Throughout the course we will use the R programming language.

Self Paced
Self-Paced
Big Data Capstone Project (edX) EdX
University of Adelaide,AdelaideX

Big Data Capstone Project (edX)

Further develop your knowledge of big data by applying the skills you have learned to a real-world data science project. This project will give you the opportunity to deepen your learning by giving you valuable experience in evaluating, selecting and applying relevant data science techniques, principles and theory to a data science problem. This project will see you plan and execute a reasonably substantial project and demonstrate autonomy, initiative and accountability.

Self Paced
Self-Paced
Data Science and Agile Systems for Product Management (edX) EdX
University of Maryland, College Park,University System of Maryland - USM,USMx,UMD

Data Science and Agile Systems for Product Management (edX)

Deliver faster, higher quality, and fault-tolerant products regardless of industry using the latest in Agile, DevOps, and Data Science. Modern systems today must be designed for agility in order to outpace the competition. Concepts like Agile, DevOps, and Data Science were once considered only for the technology-based companies. Today that means every company. Because there is no greater currency than timely information for optimizing operations and meeting the needs of customers.

Self Paced
Self-Paced
Advanced Statistical Inference and Modelling Using R (edX) EdX
University of Canterbury,UCx

Advanced Statistical Inference and Modelling Using R (edX)

Extend your knowledge of linear regression to the situations where the response variable is binary, a count, or categorical as well as to hierarchical experimental set-up. Advanced Statistical Inference and Modelling Using R is part two of the Statistical Analysis in R professional certificate. This course is directed at people who are already familiar with basic linear regression and fundamentals of statistical inference. It extends the knowledge of linear regression to the situations where the response variable is binary, a count, or categorical as well as to hierarchical experimental set-up.

Self Paced
Self-Paced
Data Science: Wrangling (edX) EdX
HarvardX,Harvard University

Data Science: Wrangling (edX)

Learn to process and convert raw data into formats needed for analysis. In this course, we cover several standard steps of the data wrangling process like importing data into R, tidying data, string processing, HTML parsing, working with dates and times, and text mining. Rarely are all these wrangling steps necessary in a single analysis, but a data scientist will likely face them all at some point.

Self Paced
Self-Paced
Estadística Aplicada a los Negocios (edX) EdX
Galileo University,GalileoX

Estadística Aplicada a los Negocios (edX)

Aprende las principales herramientas y técnicas de la estadística descriptiva y la estadística inferencial para analizar e interpretar datos desde la perspectiva de negocios facilitando la toma de decisiones. Este curso proporciona una introducción al análisis de datos en base a las principales herramientas estadísticas, enfocándose en la estadística descriptiva y la estadística inferencial.

Self Paced
Self-Paced
Introduction to Data Science (edX) EdX
IBM

Introduction to Data Science (edX)

Learn about the world of data science first-hand from real data scientists. The art of uncovering the insights and trends in data has been around for centuries. The ancient Egyptians applied census data to increase efficiency in tax collection and they accurately predicted the flooding of the Nile river every year.

Self Paced
Self-Paced
Statistics and R (edX) EdX
HarvardX,Harvard University

Statistics and R (edX)

An introduction to basic statistical concepts and R programming skills necessary for analyzing data in the life sciences. We will learn the basics of statistical inference in order to understand and compute p-values and confidence intervals, all while analyzing data with R. We provide R programming examples in a way that will help make the connection between concepts and implementation.

Self Paced
Self-Paced