Greeting from TEKNIKOZ
Location- Remote
Experience- 10-12yrs or +
- We are looking for a Compiler & Performance Engineer to develop the compiler stack for our custom AI accelerators and optimize modern Transformer/LLM workloads on our hardware.
- Responsibilities
- Develop compiler infrastructure, IRs, lowering passes, code generation, and hardware-specific optimizations.
- Optimize Transformer/LLM workloads including GEMM, attention, normalization, KV-cache, and data movement.
- Develop and optimize low-level kernels for our accelerator architecture.
- Profile workloads and identify bottlenecks across compute, memory, and communication.
- Improve latency, throughput, memory efficiency, and accelerator utilization.
- Work closely with hardware, ML, runtime, and systems teams on hardware-software co-design.
- Enable new ML models and operators on the accelerator.
- Requirements
- Strong C++ and Python programming skills.
- Strong understanding of compiler fundamentals and computer architecture.
- Experience with performance optimization, profiling, and low-level systems.
- Familiarity with ML compilers such as LLVM/MLIR, TVM, XLA, or Triton is a plus.
- Experience with Transformers/LLMs, GPU/NPU programming, or AI accelerators is highly desirable.
- Strong analytical and debugging skills.