Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
The Chinese AI GPU Ecosystem Landscape
- Contrasting Huawei Ascend, Biren, and Cambricon MLU architectures
- Differences between CUDA and CANN, Biren SDK, or BANGPy programming models
- Market trends and vendor ecosystem dynamics
Readiness for Migration
- Evaluating the complexity of your existing CUDA codebase
- Defining target platforms and selecting appropriate SDK versions
- Configuring toolchains and preparing development environments
Strategies for Code Translation
- Adapting CUDA memory management and kernel logic
- Mapping compute grid and thread structures
- Evaluating automated tools versus manual translation approaches
Implementation on Specific Platforms
- Leveraging Huawei CANN operators and developing custom kernels
- Utilising the Biren SDK conversion pipeline
- Reconstructing models using BANGPy (Cambricon)
Cross-Platform Validation and Tuning
- Profiling execution metrics on each target architecture
- Comparing memory optimisation and parallel execution strategies
- Monitoring performance and iterative improvement
Operationalising Mixed GPU Environments
- Designing hybrid deployments across multiple architectures
- Implementing fallback mechanisms and device discovery
- Utilising abstraction layers to enhance code maintainability
Case Studies and Industry Standards
- Porting vision and NLP models to Ascend or Cambricon platforms
- Adapting inference workflows for Biren clusters
- Resolving version incompatibilities and API discrepancies
Recap and Future Roadmap
Requirements
- Proficiency in CUDA programming or GPU-based application development
- Solid grasp of GPU memory hierarchies and compute kernels
- Knowledge of AI model deployment or acceleration processes
Target Audience
- GPU Developers
- System Architects
- Software Porting Specialists
21 Hours