Search by job, company or skills

AI20P Compiler Engineer

AI20P Compiler Engineer

qpisemi
2-4 Years
Not Disclosed
  • Posted 21 days ago
  • Be among the first 10 applicants

Job Description


About the Role

We are seeking a highly skilled AI20P Compiler Engineer to join our accelerated computing division. You will develop and optimize the parallelizing compiler that partitions, schedules, and executes large-scale machine learning and generative AI workloads across clusters of AI20P accelerators. You will bridge the gap between ML frameworks (like PyTorch or TensorFlow) and custom silicon hardware, ensuring peak performance, scaling, and efficiency.

Key Responsibilities

  • Compiler Development: Write production-level code (primarily in C++ and Python) to design and implement compiler toolchains and Intermediate Representations (IR) for machine learning accelerators.
  • Performance Optimization: Develop and enhance parallelization, memory management, and scheduling algorithms to minimize compute and data movement costs.
  • Hardware Co-Design: Collaborate with hardware architects to shape the Hardware/Software (HW/SW) interface for next-generation AI20P architectures.
  • Scaling & Distribution: Scale large ML models across distributed training and inference clusters.
  • Toolchain & Infrastructure: Modernize compiler development infrastructure, build/test fixtures, and debugging tools to unblock developer productivity.

Qualifications

  • Education: Bachelor's degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience.
  • Experience: 2+ years of experience with software development, coding in data structures, algorithms, and low-level programming.
  • Programming: Strong proficiency in C++ and Python.
  • Systems: Experience with compiler construction, parallel computing, or low-level hardware interaction.

Preferred Qualification

  • Education: Master's degree or PhD in Computer Science.
  • Domain Expertise: Previous experience working with ML compilation frameworks (e.g., XLA, LLVM, MLIR, TVM) or ML accelerators (AI20Ps, GPUs, DSPs, or Vector machines).
  • Parallel Computing: Demonstrated success in mapping complex algorithms (especially generative AI or large-language models) onto parallel computing architectures.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

ML compilation frameworks

LLVM

low-level hardware interaction

MLIR

TVM

XLA

About Company