Login Sign Up

Scalable Machine Learning on Big Data using Apache Spark

IBM via Coursera

Coursera based on 1 ratings

Share

0

Overview

This course will empower you with the skills to scale data science and machine learning (ML) tasks on Big Data sets using Apache Spark. Most real world machine learning work involves very large data sets that go beyond the CPU, memory and storage limitations of a single computer. Apache Spark is an open source framework that leverages cluster computing and distributed storage to process extremely large data sets in an efficient and cost effective manner. Therefore an applied knowledge of...

Syllabus

  • Week 1: Introduction
    • This is an introduction to Apache Spark. You'll learn how Apache Spark internally works and how to use it for data processing. RDD, the low level API is introduced in conjunction with parallel programming / functional programming. Then, different types of data storage solutions are contrasted. Finally, Apache Spark SQL and the optimizer Tungsten and Catalyst are explained.
  • Week 2: Scaling Math for Statistics on Apache Spark
    • Applying basic statistical calculations using the Apache Spark RDD API in order to experience how parallelization in Apache Spark works
  • Week 3: Introduction to Apache SparkML
    • Understand the concept of machine learning pipelines in order to understand how Apache SparkML works programmatically
  • Week 4: Supervised and Unsupervised learning with SparkML
    • Apply Supervised and Unsupervised Machine Learning tasks using SparkML
Scalable Machine Learning on Big Data using Apache Spark
Go to Class

IBM via Coursera

6 hours 38 minutes

Paid Certificate Available

English

On-Demand

Intermediate

Instructor

Romeo Kienzler

Reviews

No reviews yet. Be the first to review!