Get in Touch

Course Outline

Introduction

This module offers a foundational overview of machine learning, discussing when to apply it, key considerations, and its implications, including advantages and limitations. It covers data types (structured, unstructured, static, and streamed), data validity and volume, the distinction between data-driven and user-driven analytics, and the comparison between statistical models and machine learning models. Additionally, it addresses the challenges of unsupervised learning, the bias-variance trade-off, iterative evaluation, cross-validation methods, and the paradigms of supervised, unsupervised, and reinforcement learning.

KEY TOPICS

1. Grasping Naive Bayes

  • Foundational concepts of Bayesian methods
  • Probability theory
  • Joint probability
  • Conditional probability via Bayes' theorem
  • The Naive Bayes algorithm
  • Naive Bayes classification
  • Laplace smoothing
  • Incorporating numerical features into Naive Bayes

2. Mastering Decision Trees

  • The divide-and-conquer strategy
  • The C5.0 decision tree algorithm
  • Selecting optimal splits
  • Pruning decision trees

3. Exploring Neural Networks

  • Transition from biological to artificial neurons
  • Activation functions
  • Network architecture
  • Determining the number of layers
  • Direction of information flow
  • Configuring nodes per layer
  • Training networks using backpropagation
  • Deep Learning fundamentals

4. Comprehending Support Vector Machines

  • Classification using hyperplanes
  • Maximizing the margin
  • Handling linearly separable data
  • Handling non-linearly separable data
  • Applying kernels for non-linear spaces

5. Understanding Clustering

  • Clustering as a machine learning task
  • The k-means clustering algorithm
  • Using distance metrics for cluster assignment and updates
  • Determining the optimal number of clusters

6. Assessing Classification Performance

  • Working with classification prediction datasets
  • Analyzing confusion matrices in detail
  • Utilizing confusion matrices for performance measurement
  • Metrics beyond accuracy
  • The kappa statistic
  • Sensitivity and specificity
  • Precision and recall
  • The F-measure
  • Visualizing performance trade-offs
  • ROC curves
  • Predicting future performance
  • The holdout method
  • Cross-validation
  • Bootstrap sampling

7. Optimizing Standard Models for Enhanced Performance

  • Leveraging caret for automated parameter tuning
  • Constructing basic tuned models
  • Customizing the tuning workflow
  • Enhancing performance through meta-learning
  • Concepts of ensembles
  • Bagging
  • Boosting
  • Random forests
  • Training random forest models
  • Evaluating random forest efficacy

SECONDARY TOPICS

8. Classification via Nearest Neighbors

  • The kNN algorithm
  • Distance calculation
  • Selecting an appropriate k value
  • Data preparation for kNN
  • The lazy nature of the kNN algorithm

9. Classification Rules

  • The separate-and-conquer approach
  • The One Rule algorithm
  • The RIPPER algorithm
  • Deriving rules from decision trees

10. Regression Fundamentals

  • Simple linear regression
  • Ordinary least squares estimation
  • Correlation analysis
  • Multiple linear regression

11. Regression Trees and Model Trees

  • Integrating regression into tree structures

12. Association Rules

  • The Apriori algorithm for association rule mining
  • Measuring rule significance via support and confidence
  • Constructing rule sets using the Apriori principle

Additional Content

  • Spark, PySpark, MLlib, and Multi-armed bandits

Requirements

Proficiency in Python

 21 Hours

Number of participants


Price per participant

Testimonials (7)

Upcoming Courses

Related Categories