Course Outline
Introduction
Grasping the Concept of Big Data
Spark Framework Overview
Python Language Overview
Introduction to PySpark
- Distributing Data via the Resilient Distributed Datasets (RDD) Framework
- Distributing Computation through Spark API Operators
Configuring Python with Spark
Installing and Setting Up PySpark
Utilizing Amazon Web Services (AWS) EC2 Instances for Spark
Setting Up Databricks
Configuring the AWS EMR Cluster
Foundations of Python Programming
- Initiating Python Development
- Utilizing the Jupyter Notebook
- Managing Variables and Basic Data Types
- Handling Lists
- Implementing Conditional Logic (if Statements)
- Processing User Inputs
- Utilizing While Loops
- Creating and Using Functions
- Designing Classes
- Managing Files and Exception Handling
- Interacting with Projects, Data, and APIs
Essentials of Spark DataFrames
- Introduction to Spark DataFrames
- Performing Fundamental Operations with Spark
- Applying GroupBy and Aggregation Functions
- Handling Timestamps and Date Data
Practical Exercise: Spark DataFrame Project
Exploring Machine Learning with MLlib
Applying MLlib, Spark, and Python for Machine Learning Tasks
Introduction to Regression Models
- Understanding Linear Regression Theory
- Writing Code for Regression Evaluation
- Practical Exercise: Linear Regression
- Understanding Logistic Regression Theory
- Implementing Logistic Regression Code
- Practical Exercise: Logistic Regression
Random Forests and Decision Trees
- Theoretical Foundations of Tree-based Methods
- Implementing Code for Decision Trees and Random Forests
- Practical Exercise: Random Forest Classification
Implementation of K-means Clustering
- Theory behind K-means Clustering
- Coding a K-means Clustering Algorithm
- Practical Exercise: Clustering Task
Developing Recommender Systems
Implementing Natural Language Processing
- Concepts of Natural Language Processing (NLP)
- Survey of NLP Tools
- Practical Exercise: NLP Task
Streaming Data with Spark and Python
- Overview of Spark Streaming
- Practical Exercise: Spark Streaming
Requirements
- Foundational programming proficiency
Target Audience
- Software Developers
- IT Professionals
- Data Scientists
Testimonials (6)
I liked that it was practical. Loved to apply the theoretical knowledge with practical examples.
Aurelia-Adriana - Allianz Services Romania
Course - Python and Spark for Big Data (PySpark)
The course was about a series of very complex related topics & Pablo has in-depth expertise of each of them. Sometimes nuances were lost in communication and/or due to time pressures and possibly expectations were not quite met due to this. Also there were some UHG/Azure Databricks setup issues however Pablo / UHG resolved these quickly once they became apparent - this to me showed a high level of understanding and professionalism between UHG & Pablo,
Michael Monks - Tech NorthWest Skillnet
Course - Python and Spark for Big Data (PySpark)
Individual attention.
ARCHANA ANILKUMAR - PPL
Course - Python and Spark for Big Data (PySpark)
Hands on Training..
Abraham Thomas - PPL
Course - Python and Spark for Big Data (PySpark)
The lessons were taught in a Jupyter notebook. The topics were structured with a logical sequence and naturally helped develop the session from the easier parts to the more complex. I'm already an advanced user of Python with background in Machine Learning, so found the course easier to follow than, possibly, some of my classmates that took the training course. I appreciate that some of the most elementary concepts were skipped and that he focused on the most substantial matters.
Angela DeLaMora - ADT, LLC
Course - Python and Spark for Big Data (PySpark)
practice tasks