Get in Touch

Course Outline

Performance Fundamentals and Key Metrics

  • Analysis of latency, throughput, power consumption, and resource usage
  • Distinguishing between system-level and model-level constraints
  • Profiling techniques for inference versus training phases

Profiling Workflows on Huawei Ascend

  • Leveraging CANN Profiler and MindInsight tools
  • Diagnosing kernel and operator behavior
  • Managing offload patterns and memory mapping

Performance Analysis on Biren GPU

  • Utilizing Biren SDK for performance monitoring
  • Techniques for kernel fusion, memory alignment, and execution queue management
  • Profiling with awareness of power and thermal constraints

Profiling on Cambricon MLU

  • Employing BANGPy and Neuware performance utilities
  • Gaining kernel-level insight and interpreting logs
  • Integrating MLU profilers with deployment frameworks

Optimization at the Graph and Model Level

  • Strategies for graph pruning and quantization
  • Operator fusion and reorganizing computational graphs
  • Standardizing input sizes and tuning batch parameters

Memory and Kernel Efficiency

  • Refining memory layout and data reuse patterns
  • Managing buffers effectively across different chipsets
  • Platform-specific kernel tuning methods

Best Practices for Cross-Platform Optimization

  • Ensuring performance portability through abstraction strategies
  • Developing unified tuning pipelines for multi-chip setups
  • Case study: optimizing an object detection model across Ascend, Biren, and MLU

Conclusion and Future Directions

Requirements

  • Practical experience in AI model training or deployment pipelines
  • Comprehension of GPU/MLU computing principles and model optimization strategies
  • Foundational knowledge of performance profiling tools and key metrics

Target Audience

  • Performance Engineers
  • Machine Learning Infrastructure Teams
  • AI System Architects
 21 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories