Scalable Machine Learning (edX)

Start Date

Not Available

Free Course

EdX

University of California, Berkeley

Categories

CS: Systems, Security & Networking

CS: Software Engineering

Statistics & Data Analysis

Effort

Certification

Languages

Programming background; comfort with mathematical and algorithmic reasoning; familiarity with basic machine learning concepts; exposure to algorithms, probability, linear algebra and calculus; experience with Python (or the ability to learn it quickly).

Misc

Learn the underlying principles required to develop scalable machine learning pipelines and gain hands-on experience using Apache Spark. Machine learning aims to extract knowledge from data, relying on fundamental concepts in computer science, statistics, probability and optimization.

A newer version of this course is available here:
Distributed Machine Learning with Apache Spark

Learning algorithms enable a wide range of applications, from everyday tasks such as product recommendations and spam filtering to bleeding edge applications like self-driving cars and personalized medicine. In the age of ‘Big Data,’ with datasets rapidly growing in size and complexity and cloud computing becoming more pervasive, machine learning techniques are fast becoming a core component of large-scale data processing pipelines.

This course introduces the underlying statistical and algorithmic principles required to develop scalable real-world machine learning pipelines. We present an integrated view of data processing by highlighting the various components of these pipelines, including exploratory data analysis, feature extraction, supervised learning, and model evaluation. You will gain hands-on experience applying these principles using Apache Spark, a cluster computing system well-suited for large-scale machine learning tasks. You will implement scalable algorithms for fundamental statistical models (linear regression, logistic regression, matrix factorization, principal component analysis) while tackling key problems from various domains: online advertising, personalized recommendation, and cognitive neuroscience.

What you'll learn:

- The underlying statistical and algorithmic principles required to develop scalable real-world machine learning pipelines

- Exploratory data analysis, feature extraction, supervised learning, and model evaluation

- Application of these principles using Apache Spark

- How to implement scalable algorithms for fundamental statistical models