Data Engineering und Data Science – Klarheit in den Schlagwort-Dschungel (openHPI)

Data Engineering und Data Science – Klarheit in den Schlagwort-Dschungel (openHPI)

Die Schlagwörter Künstliche Intelligenz, Data Science, Data Engineering, und Big Data dominieren seit einigen Jahren nicht nur die IT-Schlagzeilen. In unserem Kurs wollen wir diese Wörter mit grundlegendem Inhalt füllen und die typischen Arbeitsschritte eines Data Scientists nachvollziehen. Insbesondere schauen wir hinter die Kulissen und betrachten den oft mühsamen Weg der Daten bis sie endlich genutzt werden können um z.B. mittels maschinellem Lernen Modelle trainieren zu können. Dazu gehören die Datenbeschaffung, die Datenreinigung, und die Datenintegration. Anschließend lernen wir, wie man aus diesen Daten und auch aus Texten neue Erkenntnisse mittels Data Mining und maschinellem Lernen gewinnt. Der Abschluss bildet eine Diskussion über Ethik und Fairness bei der automatisierten Datenanalyse.

Zielgruppe
Interessierte Öffentlichkeit, PraktikerInnen und Bachelorstudierende

Kursstruktur

Woche 1: Big Data und Data Science
Woche 2: Data Science Anwendungen und Text Mining
Woche 3: Skalierbares Datenmanagement
Woche 4: Datenaufbereitung
Woche 5: Informationsintegration
Woche 6: Statistik, Data Mining, Machine Learning
Woche 7: Klausur

Go to Class
MOOC List is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Related Courses

Using R for Regression and Machine Learning in Investment (Coursera) Coursera
Sungkyunkwan University - SKKU

Using R for Regression and Machine Learning in Investment (Coursera)

In this course, the instructor will discuss various uses of regression in investment problems, and she will extend the discussion to logistic, Lasso, and Ridge regressions. At the same time, the instructor will introduce various concepts of machine learning. You can consider this course as the first step toward using machine learning methodologies in solving investment problems. The course will cover investment analysis topics, but at the same time, make you practice it using R programming. This course's focus is to train you to use various regression methodologies for investment management that you might need to do in your job every day and make you ready for more advanced topics in machine learning.

Sep 21st 2026
2 Weeks
Introduction to Machine Learning (Coursera) Coursera
Duke University

Introduction to Machine Learning (Coursera)

This course will provide you a foundational understanding of machine learning models (logistic regression, multilayer perceptrons, convolutional neural networks, natural language processing, etc.) as well as demonstrate how these models can solve complex problems in a variety of industries, from medical diagnostics to image recognition to text prediction.

Sep 21st 2026
5-12 Weeks
Big Data Analytics (openHPI) OpenHPI
Hasso-Plattner-Institut

Big Data Analytics (openHPI)

In diesem kostenlosen offenen Online-Kurs führen wir Sie in das aktuelle und viel diskutierte Thema Big Data Analytics ein. Der sechswöchige Kurs wird Ihnen verständlich machen, warum Daten der Schatz des 21. Jahrhunderts sind und wie man diesen heben kann. Ob Finanzdienstleister, produzierende Unternehmen, Internet-Dienstleister oder Forschungszentren: In Wirtschaft und Wissenschaft wächst das Interesse, in dem gewaltig anschwellenden Meer an erhobenen und anfallenden Daten diejenigen herauszufischen, die es auf interessante Zusammenhänge hin zu analysieren sowie vernünftig zu strukturieren und zu verknüpfen gilt. So können schneller bessere Erkenntnisse gewonnen, Entscheidungen getroffen und Prognosen erstellt werden. Das Meer an Daten führt dann zu einem Mehr an Wissen!

Self Paced
Self-Paced
clean-IT: Towards Sustainable Digital Technologies (openHPI) OpenHPI
Hasso-Plattner-Institut

clean-IT: Towards Sustainable Digital Technologies (openHPI)

Digitalization is a game changer in the pursuit of a sustainable future. The latest digital technologies and applications like cloud, AI, and mobile devices enable us to achieve the Sustainable Development Goals and reduce carbon emissions in many sectors. Yet computer systems themselves have an immense energy requirement for their countless devices, data centers, applications and global networks. To effectively reduce the carbon footprint of digitalization, it is necessary to apply algorithmic efficiency and sustainability by design as guiding principles in digital engineering.

Mar 31st 2021
5-12 Weeks
Data Science Ethics (Coursera) Coursera
University of Michigan

Data Science Ethics (Coursera)

What are the ethical considerations regarding the privacy and control of consumer information and big data, especially in the aftermath of recent large-scale data breaches? This course provides a framework to analyze these concerns as you examine the ethical and privacy implications of collecting and managing big data. Explore the broader impact of the data science field on modern society and the principles of fairness, accountability and transparency as you gain a deeper understanding of the importance of a shared set of ethical values.

Sep 21st 2026
4 Weeks
Media ethics & governance (Coursera) Coursera
University of Amsterdam

Media ethics & governance (Coursera)

This course explores some of the basic theories, models and concepts in the field of media ethics. We will introduce influential ethical theories and perspectives, explore changing societal demands and expectations of media creation and media use, and we will elaborate on existing ethical norms for media professionals. After following this course, you will be able to reflect on ethical dilemmas and develop a well-substantiated argumentation for ethical decision making in a variety of media-related contexts.

Sep 21st 2026
4 Weeks
Big Data Science with the BD2K-LINCS Data Coordination and Integration Center (Coursera) Coursera
Icahn School of Medicine at Mount Sinai

Big Data Science with the BD2K-LINCS Data Coordination and Integration Center (Coursera)

In this course we briefly introduce the DCIC and the various Centers that collect data for LINCS. We then cover metadata and how metadata is linked to ontologies. We then present data processing and normalization methods to clean and harmonize LINCS data. This follow discussions about how data is served as RESTful APIs. Most importantly, the course covers computational methods including: data clustering, gene-set enrichment analysis, interactive data visualization, and supervised learning. Finally, we introduce crowdsourcing/citizen-science projects where students can work together in teams to extract expression signatures from public databases and then query such collections of signatures against LINCS data for predicting small molecules as potential therapeutics.

Sep 21st 2026
5-12 Weeks
Sistemas difusos (Coursera) Coursera
Universidad Nacional de Colombia

Sistemas difusos (Coursera)

Los sistemas difusos permiten efectuar cálculos cuando hay información con incertidumbre, o cuando se debe combinar información tanto cuantitativa como cualitativa. Se trata de una aproximación matemática para modelar esas situaciones. Este curso está diseñado para ayudar a entender y explicar cómo funcionan dichos sistemas. El curso tiene una aproximación teórica y práctica. Los principios matemáticos son de un nivel bajo y están al alcance de un público muy amplio. El curso cuenta con varios laboratorios para aprender a utilizar las herramientas de software que usan esos principios. Este componente práctico requiere una comprensión mínima de programación.

Sep 21st 2026
4 Weeks
Process Mining: Data science in Action (Coursera) Coursera
Eindhoven University of Technology

Process Mining: Data science in Action (Coursera)

Process mining is the missing link between model-based process analysis and data-oriented analysis techniques. Through concrete data sets and easy to use software the course provides data science knowledge that can be applied directly to analyze and improve processes in a variety of domains. Data science is the profession of the future, because organizations that are unable to use (big) data in a smart way will not survive. It is not sufficient to focus on data storage and data analysis. The data scientist also needs to relate data to process analysis.

Sep 21st 2026
5-12 Weeks
Künstliche Intelligenz und maschinelles Lernen für Einsteiger (openHPI) OpenHPI
Hasso-Plattner-Institut

Künstliche Intelligenz und maschinelles Lernen für Einsteiger (openHPI)

Hier lernen Jugendliche und andere Interessierte ohne Programmier-Erfahrung und technisches Hintergrund-Wissen, die Welt des maschinellen Lernens und der künstlichen Intelligenz zu verstehen. Wir führen Sie dazu in die grundlegenden Konzepte ein. Dabei erfahren Sie, wo die Unterschiede zwischen herkömmlicher Programmierung und der Entwicklung selbstlernender Software liegen. Anhand von Beispielen erfahren Sie, was überwachtes, nicht überwachtes und verstärkendes Lernen sind. Denn diese Konzepte bilden den Kern für die Algorithmen, welche das maschinelle Lernen bewirken. Erleben Sie anhand einer konkreten Anwendung, wie mit einem solchen Lernprozess Muster und Strukturen in großen Datenmengen erkannt werden können. Auch auf ethische Fragen beim Einsatz künstlicher Intelligenz sowie die Begrenzungen der Technologie maschinellen Lernens wird in dem vierwöchigen Gratis-Kurs eingegangen. Geleitet wird er von den Masterstudenten Johannes Hötter und Christian Warmuth.

Self Paced
Self-Paced
Data Science for Business Innovation (Coursera) Coursera
Politecnico di Milano,EIT Digital

Data Science for Business Innovation (Coursera)

The course is a compendium of the must-have expertise in data science for executive and middle-management to foster data-driven innovation. It consists of introductory lectures spanning big data, machine learning, data valorization and communication. Topics cover the essential concepts and intuitions on data needs, data analysis, machine learning methods, respective pros and cons, and practical applicability issues.

Sep 21st 2026
4 Weeks