Search Jobs

Search by job, company or skills

AI Infrastructure System Performance Architect

AI Infrastructure System Performance Architect

Randstad Sourceright
Early Applicant
  • Posted 5 days ago
  • Be among the first 10 applicants

Job Description

Position: AI Infrastructure System Performance Architect (IRT815ST RM 4407)

Position Summary

Analyze system-level performance, scalability, and bottlenecks of the company's chiplet-based AI infrastructure using virtual platforms and performance models.

Education : B.E. / B.Tech / M.E. / M.Tech in Electronics, Electrical, Computer Engineering, or related discipline.

Key Responsibilities

  • Memory / Data-Movement Modelling: Analyze memory bandwidth and latency requirements, model XPU-to-memory and chiplet-to-memory traffic, and identify data-movement bottlenecks.
  • Multi-Chiplet / Multi-XPU Scalability: Analyze scaling from single-XPU to multi-XPU systems, evaluate scale-up/scale-out architectures, and study bandwidth, latency, congestion, and resource utilization.
  • AI Infrastructure System Modelling: Develop system-level performance models covering compute, interconnect, memory, and I/O interactions, and develop representative AI/HPC traffic patterns.
  • System Performance & Architecture Exploration: Perform architecture trade-offs and analyze latency, bandwidth, throughput, utilization, queue depth, congestion, and bottlenecks; generate performance reports and architecture recommendations.

Required Technical Skills

  • 8–15+ years of system/performance architecture experience.
  • AI/HPC infrastructure and performance modelling.
  • SystemC/TLM.
  • C++ / Python.
  • PCIe / CXL / UCIe.
  • HBM/DDR.
  • System-level architecture.

Preferred Skills

  • XPU/GPU/NPU architecture.
  • AI workload characterization.
  • UALink and networking.
  • Cluster/rack-scale architecture.
  • Power-performance analysis.

Project Experience: Demonstrated system-level performance analysis (latency, bandwidth, throughput, utilization, congestion) on a chiplet-based or multi-XPU AI system.

Soft Skills

  • Strong analytical rigor in translating raw performance data into architecture recommendations.
  • Ability to collaborate across connectivity, memory, and power architecture teams to build a coherent system view.

Good to Have

  • Experience characterizing real AI/HPC workloads for use in performance models.
  • Exposure to cluster or rack-scale system architecture.

Expected Deliverables

  • System-level performance models covering compute, interconnect, memory, and I/O.
  • Architecture trade-off studies and performance reports with recommendations.
  • Representative AI/HPC traffic pattern definitions for use across the team's models.

Growth Path

  • Progress into Principal Performance Architect, System Architecture Lead, or Practice Lead (Performance Engineering) roles.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

SystemC TLM

HBM

CXL

System-level architecture

UCIe

AI HPC infrastructure and performance modelling