Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 21 hours
Course Outline
Foundational Performance Metrics and Concepts
- Analysis of latency, throughput, power consumption, and resource utilization
- Distinguishing between system-wide and model-specific bottlenecks
- Profiling methodologies for inference versus training phases
Profiling Techniques for Huawei Ascend
- Leveraging CANN Profiler and MindInsight tools
- Diagnostic analysis of kernels and operators
- Examining offload patterns and memory mapping strategies
Performance Analysis on Biren GPU
- Utilizing Biren SDK monitoring capabilities
- Exploring kernel fusion, memory alignment, and execution queue management
- Implementing power and temperature-aware profiling
Optimizing Cambricon MLU Performance
- Using BANGPy and Neuware performance utilities
- Gaining kernel-level visibility and interpreting diagnostic logs
- Integrating the MLU profiler with deployment frameworks
Graph and Model Architecture Optimization
- Strategies for graph pruning and quantization
- Operator fusion and restructuring computational graphs
- Standardizing input sizes and optimizing batch settings
Memory and Kernel-Level Enhancements
- Refining memory layout and data reuse patterns
- Managing buffers efficiently across different chipsets
- Applying platform-specific kernel tuning techniques
Best Practices for Cross-Platform Development
- Achieving performance portability through abstraction strategies
- Constructing shared tuning pipelines for multi-chip environments
- Case study: Tuning an object detection model across Ascend, Biren, and MLU architectures
Conclusion and Future Recommendations
Requirements
- Practical experience with AI model training or deployment pipelines
- Comprehensive grasp of GPU/MLU compute principles and model optimization techniques
- Familiarity with fundamental performance profiling tools and key metrics
Target Audience
- Performance engineers
- Machine learning infrastructure teams
- AI system architects