Job Description
Role: Senior NPU Compiler Engineer
Years: 7+ | Remote
Build the compiler that brings Processing-in-Memory AI to life.
InCommon is hiring on behalf of a California based company that is building ultra-low-power neural processing technology for edge AI — chips that run advanced AI workloads within tight power and memory budgets. The Processing-in-Memory (PIM) architecture computes directly inside memory, breaking the traditional bottleneck of shuttling data back and forth.
We're looking for a Senior NPU Compiler Engineer to build the compiler stack — largely from the ground up — that takes models from TFLite, ONNX, and PyTorch and turns them into instructions our NPU can execute.
What you'll do:
Design and build the full compilation pipeline — graph import, operator lowering, scheduling, memory planning, and instruction generation
Map neural network graphs onto a novel dataflow / Processing-in-Memory architecture
Optimize for execution efficiency, memory bandwidth, and power
Build validation tools to compare compiler output against simulator and RTL results
Work directly with architecture, RTL, firmware, and ML teams to shape the hardware/software interface
What you bring:
Strong experience building compilers, graph compilers, or backend toolchains
Solid C++ and Python
Deep understanding of compiler fundamentals — IRs, graph transforms, lowering, scheduling, codegen
Experience with ML model formats (ONNX, TFLite, PyTorch) and neural network operators/tensors
Background building software for NPUs, DSPs, GPUs, or other specialized accelerators
Nice to have:
Experience with MLIR, LLVM, TVM, XLA, Glow, or IREE
Familiarity with dataflow or Processing-in-Memory architectures
Experience with binary instruction streams, DMA engines, or on-chip SRAM
Why join us:
This is a rare chance to build a compiler stack from scratch for genuinely novel silicon — with real architectural influence, not just downstream implementation work. You'll shape decisions that affect the hardware itself, not just work around it.
