EdX

Big Data Analytics (edX)

Big Data Analytics (edX)

Learn key technologies and techniques, including R and Apache Spark, to analyse large-scale data sets to uncover valuable business information. Gain essential skills in today’s digital age to store, process and analyse data to inform business decisions.

Class Deals by MOOC List - Click here and see EdX's Active Discounts, Deals, and Promo Codes.

In this course, part of the Big Data MicroMasters program, you will develop your knowledge of big data analytics and enhance your programming and mathematical skills. You will learn to use essential analytic tools such as Apache Spark and R.
Topics covered in this course include:

  • cloud-based big data analysis;
  • predictive analytics, including probabilistic and statistical models;
  • application of large-scale data analysis;
  • analysis of problem space and data needs.

By the end of this course, you will be able to approach large-scale data science problems with creativity and initiative.
This course is part of the Big Data MicroMasters.

What you'll learn

  • How to develop algorithms for the statistical analysis of big data;
  • Knowledge of big data applications;
  • How to use fundamental principles used in predictive analytics;
  • Evaluate and apply appropriate principles, techniques and theories to large-scale data science problems.

Prerequisites
Candidates pursuing the MicroMasters program are advised to complete Programming for Data Science, Computational Thinking and Big Data & Big Data Fundamentals before undertaking this course.

Course Syllabus

Section 1: Simple linear regression
Fit a simple linear regression between two variables in R; Interpret output from R; Use models to predict a response variable; Validate the assumptions of the model.

Section 2: Modelling data
Adapt the simple linear regression model in R to deal with multiple variables; Incorporate continuous and categorical variables in their models; Select the best-fitting model by inspecting the R output.

Section 3: Many models
Manipulate nested dataframes in R; Use R to apply simultaneous linear models to large data frames by stratifying the data; Interpret the output of learner models.

Section 4: Classification
Adapt linear models to take into account when the response is a categorical variable; Implement Logistic regression (LR) in R; Implement Generalised linear models (GLMs) in R; Implement Linear discriminant analysis (LDA) in R.

Section 5: Prediction using models
Implement the principles of building a model to do prediction using classification; Split data into training and test sets, perform cross validation and model evaluation metrics; Use model selection for explaining data with models; Analyse the overfitting and bias-variance trade-off in prediction problems.

Section 6: Getting bigger
Set up and apply sparklyr; Use logical verbs in R by applying native sparklyr versions of the verbs.

Section 7: Supervised machine learning with sparklyr
Apply sparklyr to machine learning regression and classification models; Use machine learning models for prediction; Illustrate how distributed computing techniques can be used for “bigger” problems.

Section 8: Deep learning
Use massive amounts of data to train multi-layer networks for classification; Understand some of the guiding principles behind training deep networks, including the use of autoencoders, dropout, regularization, and early termination; Use sparklyr and H2O to train deep networks.

Section 9: Deep learning applications and scaling up
Understand some of the ways in which massive amounts of unlabelled data, and partially labelled data, is used to train neural network models; Leverage existing trained networks for targeting new applications; Implement architectures for object classification and object detection and assess their effectiveness.

Section 10: Bringing it all together
Consolidate your understanding of relationships between the methodologies presented in this course, theirrelative strengths, weaknesses and range of applicability of these methods.

Go to Class
MOOC List is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Related Courses

Big Data Capstone Project (edX) EdX
University of Adelaide,AdelaideX

Big Data Capstone Project (edX)

Further develop your knowledge of big data by applying the skills you have learned to a real-world data science project. This project will give you the opportunity to deepen your learning by giving you valuable experience in evaluating, selecting and applying relevant data science techniques, principles and theory to a data science problem. This project will see you plan and execute a reasonably substantial project and demonstrate autonomy, initiative and accountability.

Self Paced
Self-Paced
Python for Data Science (edX) EdX
University of California, San Diego,UC San DiegoX

Python for Data Science (edX)

Learn to use powerful, open-source, Python tools, including Pandas, Git and Matplotlib, to manipulate, analyze, and visualize complex datasets. In the information age, data is all around us. Within this data are answers to compelling questions across many societal domains (politics, business, science, etc.). But if you had access to a large dataset, would you be able to find the answers you seek?

Self Paced
Self-Paced
Basics of Statistical Inference and Modelling Using R (edX) EdX
University of Canterbury,UCx

Basics of Statistical Inference and Modelling Using R (edX)

Learn why a statistical method works, how to implement it using R and when to apply it and where to look if the particular statistical method is not applicable in the specific situation. Basics of Statistical Inference and Modelling Using R is part one of the Statistical Analysis in R professional certificate.

Self Paced
Self-Paced
Unix Tools: Data, Software and Production Engineering (edX) EdX
Delft University of Technology,DelftX

Unix Tools: Data, Software and Production Engineering (edX)

Grow from being a Unix novice to Unix wizard status! Process big data, analyze software code, run DevOps tasks and excel in your everyday job through the amazing power of the Unix shell and command-line tools. Processing information is the hallmark of all modern organizations, which are increasingly digital: absorbing, processing and generating information is a key element of their business.

Self Paced
Self-Paced
Computational Thinking and Big Data (edX) EdX
University of Adelaide,AdelaideX

Computational Thinking and Big Data (edX)

Learn the core concepts of computational thinking and how to collect, clean and consolidate large-scale datasets. Computational thinking is an invaluable skill that can be used across every industry, as it allows you to formulate a problem and express a solution in such a way that a computer can effectively carry it out.

Self Paced
Self-Paced
Minería de Datos: Análisis de la Canasta de Compra (edX) EdX
Universidad Anáhuac,AnahuacX

Minería de Datos: Análisis de la Canasta de Compra (edX)

¿Conoces realmente qué productos de la canasta de mercado compran tus clientes o te dejas llevar por lo que aparenta a simple vista? En este curso aprenderás a construir modelos basados en técnicas de data mining o minería de datos, que te permitirán conocer información relevante de tus clientes y descubrir patrones de comportamiento para definir estrategias de marketing de acuerdo a la compra de productos.

Self Paced
Self-Paced
Big Data Computing with Spark (edX) EdX
The Hong Kong University of Science and Technology - HKUST,HKUSTx

Big Data Computing with Spark (edX)

Learn the theory and gain hands-on experience of big data systems, using Spark as the exemplary platform. Big data systems such as Hadoop and Spark emerge as enabling technologies in managing massive amounts of data across hundreds or even thousands of computing nodes. Meanwhile, cloud computing platforms have made these technologies easily accessible to individuals as well as large enterprises.

Self Paced
Self-Paced
Visualización de Datos y Storytelling (edX) EdX
Tecnológico de Monterrey,TecdeMonterreyX

Visualización de Datos y Storytelling (edX)

Aprende en este curso en línea que es la visualización de datos, sus usos; los elementos que la conforman y la forma de poder utilizarla para el apoyo en la toma de las mejores decisiones para las empresas basadas en el análisis de datos. Digamos que necesitas comprender big data; miles o incluso millones de filas de datos, y tienes poco tiempo para hacerlo. los datos pueden provenir de tu equipo, en cuyo caso tal vez ya estés familiarizado con lo que estás midiendo y de los resultados que se esperan. O puede provenir de otro equipo, o tal vez de varios equipos a la vez, y estar completamente familiarizado.

Self Paced
Self-Paced
Predictive Analytics (edX) EdX
Indian Institute of Management, Bangalore,IIMBx

Predictive Analytics (edX)

Master the tools of predictive analytics in this statistics based analytics course. Decision makers often struggle with questions such as: What should be the right price for a product? Which customer is likely to default in his/her loan repayment? Which products should be recommended to an existing customer? Finding right answers to these questions can be challenging yet rewarding.

Self Paced
5-12 Weeks