Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 21 hours
Course Outline
Insights into the Chinese AI GPU Ecosystem
- Analyzing Huawei Ascend, Biren, and Cambricon MLU architectures.
- Contrasting CUDA with CANN, Biren SDK, and BANGPy frameworks.
- Market trends and the landscape of vendor ecosystems.
Readiness for Migration
- Conducting a review of your current CUDA codebase.
- Defining target platforms and required SDK versions.
- Setting up toolchains and configuring development environments.
Methodologies for Code Translation
- Adapting CUDA memory access patterns and kernel logic.
- Aligning compute grid and thread structures.
- Evaluating automated versus manual translation approaches.
Implementation on Specific Platforms
- Leveraging Huawei CANN operators and developing custom kernels.
- Utilizing the Biren SDK conversion workflow.
- Reconstructing models using BANGPy (Cambricon).
Testing and Optimization Across Platforms
- Profiling execution performance on each designated platform.
- Adjusting memory usage and comparing parallel execution strategies.
- Monitoring performance metrics and iterating on solutions.
Oversight of Mixed GPU Setups
- Implementing hybrid deployments across multiple architectures.
- Developing fallback mechanisms and device detection logic.
- Creating abstraction layers to ensure long-term code maintainability.
Practical Examples and Industry Standards
- Porting vision and NLP models to Ascend or Cambricon.
- Adapting inference pipelines for Biren clusters.
- Mitigating version discrepancies and API limitations.
Conclusion and Future Directions
Requirements
- Prior experience in programming with CUDA or GPU-accelerated applications.
- A solid grasp of GPU memory hierarchies and compute kernel design.
- Familiarity with workflows for deploying or accelerating AI models.
Target Audience
- GPU developers.
- System architects.
- Specialists in code porting.