Get in Touch
 Duration 14 hours

Course Outline

Readying Machine Learning Models for Production

  • Containerizing models using Docker
  • Model export processes for TensorFlow and PyTorch
  • Best practices for version control and storage management

Serving Models on Kubernetes

  • Introduction to key inference server technologies
  • Deployment strategies for TensorFlow Serving and TorchServe
  • Configuration of model endpoints

Optimizing Inference Performance

  • Effective batching methodologies
  • Managing concurrent request loads
  • Tuning for low latency and high throughput

Automating ML Workload Scaling

  • Implementing Horizontal Pod Autoscaler (HPA)
  • Applying Vertical Pod Autoscaler (VPA)
  • Leveraging Kubernetes Event-Driven Autoscaling (KEDA)

Managing GPU Resources and Allocation

  • Setup and configuration of GPU-enabled nodes
  • Overview of the NVIDIA device plugin
  • Defining resource requests and limits for ML tasks

Strategic Model Release and Rollouts

  • Implementing blue/green deployment patterns
  • Utilizing canary release strategies
  • Conducting A/B tests for model validation

Production ML Monitoring and Observability

  • Tracking essential inference metrics
  • Best practices for logging and distributed tracing
  • Building dashboards and setting up alerts

Ensuring Security and System Resilience

  • Protecting model endpoints
  • Configuring network policies and access controls
  • Architecting for high availability

Wrap-up and Future Recommendations

Requirements

  • A solid grasp of containerized application lifecycle management
  • Practical experience with Python-based machine learning models
  • A foundational understanding of Kubernetes concepts

Target Audience

  • ML Engineers
  • DevOps Engineers
  • Platform Engineering Teams

Number of participants


Price per participant

Testimonials (4)

Upcoming Courses

Related Categories