Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Performance Fundamentals and Key Metrics
- Analysis of latency, throughput, power consumption, and resource usage
- Distinguishing between system-level and model-level constraints
- Profiling techniques for inference versus training phases
Profiling Workflows on Huawei Ascend
- Leveraging CANN Profiler and MindInsight tools
- Diagnosing kernel and operator behavior
- Managing offload patterns and memory mapping
Performance Analysis on Biren GPU
- Utilizing Biren SDK for performance monitoring
- Techniques for kernel fusion, memory alignment, and execution queue management
- Profiling with awareness of power and thermal constraints
Profiling on Cambricon MLU
- Employing BANGPy and Neuware performance utilities
- Gaining kernel-level insight and interpreting logs
- Integrating MLU profilers with deployment frameworks
Optimization at the Graph and Model Level
- Strategies for graph pruning and quantization
- Operator fusion and reorganizing computational graphs
- Standardizing input sizes and tuning batch parameters
Memory and Kernel Efficiency
- Refining memory layout and data reuse patterns
- Managing buffers effectively across different chipsets
- Platform-specific kernel tuning methods
Best Practices for Cross-Platform Optimization
- Ensuring performance portability through abstraction strategies
- Developing unified tuning pipelines for multi-chip setups
- Case study: optimizing an object detection model across Ascend, Biren, and MLU
Conclusion and Future Directions
Requirements
- Practical experience in AI model training or deployment pipelines
- Comprehension of GPU/MLU computing principles and model optimization strategies
- Foundational knowledge of performance profiling tools and key metrics
Target Audience
- Performance Engineers
- Machine Learning Infrastructure Teams
- AI System Architects
21 Hours