Get in Touch
 Duration 21 hours

Course Outline

Overview of Biren GPU Architecture

  • Introduction to Biren and its primary use cases
  • Hardware structure: examining cores, memory, and compute clusters
  • Benchmarking against NVIDIA and AMD GPU capabilities

Preparing the Biren Development Environment

  • Installing the Biren SDK and associated runtime components
  • Grasping the toolchain and compiler logic
  • Exploring basic project structures and build workflows

Programming with the Biren Software Stack

  • Managing thread and block models
  • Handling memory management and data transfer operations
  • Developing kernels and defining launch patterns

Migrating Code from CUDA to Biren

  • Strategies for translating CUDA codebases
  • Mapping common APIs and necessary adaptations
  • Practical labs focused on code conversion

Debugging and Performance Profiling

  • Leveraging Biren’s integrated debugger and profiler
  • Pinpointing performance bottlenecks
  • Analyzing memory access patterns for optimization

Advanced Optimization Strategies

  • Refining thread scheduling and instruction pipelining
  • Utilizing loop unrolling and shared memory effectively
  • Fine-tuning kernels for maximum throughput

Practical Case Studies and Applications

  • Training machine learning models using Biren accelerators
  • Porting and profiling vision or NLP model workloads
  • Evaluating performance in comparison to CUDA/NVIDIA ecosystems

Recap and Future Directions

Requirements

  • A solid grasp of GPU architecture and parallel processing concepts
  • Hands-on experience with CUDA, OpenCL, or comparable GPU programming frameworks
  • Proficiency with deep learning frameworks like PyTorch or TensorFlow

Target Audience

  • High-Performance Computing (HPC) developers
  • AI infrastructure engineers
  • Performance optimization specialists

Number of participants


Price per participant

Testimonials (2)

Upcoming Courses

Related Categories