Search Jobs

Search by job, company or skills

Senior NPU Compiler Engineer

Senior NPU Compiler Engineer

InCommon
7-9 Years
  • Posted 8 hours ago
  • Be among the first 10 applicants

Job Description

Role: Senior NPU Compiler Engineer

Years: 7+ | Remote

Build the compiler that brings Processing-in-Memory AI to life.

InCommon is hiring on behalf of a California based company that is building ultra-low-power neural processing technology for edge AI — chips that run advanced AI workloads within tight power and memory budgets. The Processing-in-Memory (PIM) architecture computes directly inside memory, breaking the traditional bottleneck of shuttling data back and forth.

We're looking for a Senior NPU Compiler Engineer to build the compiler stack — largely from the ground up — that takes models from TFLite, ONNX, and PyTorch and turns them into instructions our NPU can execute.

What you'll do:

Design and build the full compilation pipeline — graph import, operator lowering, scheduling, memory planning, and instruction generation

Map neural network graphs onto a novel dataflow / Processing-in-Memory architecture

Optimize for execution efficiency, memory bandwidth, and power

Build validation tools to compare compiler output against simulator and RTL results

Work directly with architecture, RTL, firmware, and ML teams to shape the hardware/software interface

What you bring:

Strong experience building compilers, graph compilers, or backend toolchains

Solid C++ and Python

Deep understanding of compiler fundamentals — IRs, graph transforms, lowering, scheduling, codegen

Experience with ML model formats (ONNX, TFLite, PyTorch) and neural network operators/tensors

Background building software for NPUs, DSPs, GPUs, or other specialized accelerators

Nice to have:

Experience with MLIR, LLVM, TVM, XLA, Glow, or IREE

Familiarity with dataflow or Processing-in-Memory architectures

Experience with binary instruction streams, DMA engines, or on-chip SRAM

Why join us:

This is a rare chance to build a compiler stack from scratch for genuinely novel silicon — with real architectural influence, not just downstream implementation work. You'll shape decisions that affect the hardware itself, not just work around it.

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

Processing-in-Memory architectures

IREE

DMA engines

GPUs

tensors

binary instruction streams

TVM

MLIR

on-chip SRAM

Glow

XLA

ONNX

DSPs

TFLite

graph compilers

ML model formats

LLVM

compiler fundamentals

backend toolchains

software for NPUs

neural network operators

About Company