Search Jobs

Search by job, company or skills

Principal Software Engineer

Principal Software Engineer

Microsoft Corp
  • Posted 9 hours ago
  • Be among the first 10 applicants

Job Description

Overview

Microsoft's AI infrastructure is evolving rapidly to support the next generation of large-scale AI training and inference. The AI Frameworks (AIFx) Networking & Systems Tools (NeST) organisation develops foundational system software that enables Microsoft's Maia accelerator platforms across pre-silicon development, hardware bring-up, cloud integration and production cloud infrastructure that enables AI accelerators to operate reliably and efficiently at cloud scale.

Within NeST, the India Development Centre is building end-to-end engineering competency across the Maia system software stack. Our work spans the boundary between distributed cloud systems and low-level accelerator software, including control-plane services, host and device management software, accelerator virtualisation, Kubernetes-based infrastructure, Developer/Debugger Infrastructure tools, hardware lifecycle management, reliability, telemetry and diagnostics.

We are looking for a Principal Software Engineer with deep expertise in systems software and distributed systems to help architect and build the software infrastructure for large-scale AI accelerator platforms.

In this role, you will work across cloud services, operating systems, host agents, device interfaces and accelerator infrastructure. You will solve challenging problems involving hardware lifecycle management, orchestration, resource management, fault detection and recovery, virtualisation, observability and infrastructure reliability.

This is an opportunity to influence the architecture of foundational AI infrastructure and build systems that operate across a diverse and evolving ecosystem of GPUs and AI accelerators.

Microsoft's mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Responsibilities

  • Architect and develop distributed control-plane services for provisioning, orchestration and lifecycle management of AI accelerator infrastructure.
  • Design scalable systems for state management, health monitoring, reconciliation and fault recovery.
  • Build host and device management software connecting cloud infrastructure with operating systems, drivers, firmware and accelerator devices.
  • Develop infrastructure supporting accelerator virtualisation, device assignment, isolation and resource management.
  • Integrate accelerator infrastructure with Kubernetes and cloud-native platforms.
  • Design robust APIs and abstractions across cloud services, host software and device interfaces.
  • Drive improvements in reliability, security, observability, diagnostics, testing and operational readiness.
  • Champion AI-assisted engineering practices across architecture, design, coding, testing, debugging, code reviews and documentation to improve engineering velocity and software quality.
  • Apply AI-assisted workflows to accelerate code comprehension, root-cause analysis, test development, design exploration and engineering automation.
  • Identify opportunities to integrate AI into engineering workflows to reduce repetitive work, shorten development feedback loops and improve developer effectiveness.
  • Use AI-assisted learning together with engineering fundamentals to accelerate development of systems, hardware and AI infrastructure domain competency.
  • Establish reusable engineering practices and mentor engineers on the responsible and effective application of AI throughout the software development lifecycle.
  • Diagnose complex system issues spanning distributed services, operating systems, host software, drivers, firmware and hardware.
  • Lead architecture and design reviews, solve complex cross-layer system issues and mentor engineers.
  • Collaborate across hardware, firmware, OS, cloud infrastructure and AI platform teams to deliver end-to-end solutions.
  • Partner with hardware, firmware, operating system, cloud infrastructure and AI platform teams to deliver end-to-end solutions.
  • Mentor engineers and contribute to raising the technical capabilities and engineering standards of the broader organisation.

Qualifications

Required Qualifications:

  • Bachelor's Degree in Computer Science, Computer Engineering or related technical discipline, or equivalent experience with at least 12+ years of relevant industry experience.
  • Software engineering experience in C++, C, C#, Rust, Go or similar languages.
  • Experience designing and developing distributed systems, cloud infrastructure or systems software.
  • Understanding of concurrency, state management, asynchronous programming and failure recovery.
  • Experience building reliable production software with strong debugging and problem-solving skills.
  • Demonstrated technical leadership across complex, multi-team engineering projects.
  • Ability to leverage AI-assisted engineering tools and workflows, combined with strong engineering judgement, to improve software development effectiveness, quality and technical problem solving.
  • Ability to rapidly develop expertise in complex technical domains and help raise engineering competency across the team through mentoring, technical guidance and knowledge sharing.

Preferred Qualifications

  • Experience with cloud control planes, resource orchestration or infrastructure management systems.
  • Experience with Kubernetes, containers and cloud-native infrastructure.
  • Experience with GPU, AI accelerator or heterogeneous compute infrastructure.
  • Knowledge of PCIe, SR-IOV, PF/VF, device virtualisation or device passthrough.
  • Experience with host/device agents, Linux systems software, device drivers or firmware interfaces.
  • Experience with hardware lifecycle management, health monitoring and fault recovery.
  • Experience with telemetry, observability and debugging large-scale production systems.
  • Experience applying AI-assisted software engineering to design, development, code comprehension, debugging, testing, documentation or engineering automation.
  • Experience identifying and developing AI-enabled engineering workflows that improve developer productivity and reduce repetitive engineering effort.
  • Experience using AI-assisted approaches to accelerate understanding of large and complex codebases, system architectures and cross-layer technical issues.
  • Experience establishing engineering practices that help teams adopt AI effectively while maintaining code quality, security, correctness and engineering accountability.

#AIINFRA

This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.

Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

device virtualisation

Linux systems software

AI-assisted engineering tools

fault recovery

firmware interfaces

observability

About Company

Similar Jobs

12-14 yrs
Bengaluru, India
Skills:
Azure Cosmos DB, Apache Kafka, Spring Boot, Kubernetes, Java 8, Azure Service Bus, Azure AKS, Event streaming platforms, Azure Monitor, gRPC APIs, Microservices architecture
12-14 yrs
Bengaluru, India
Skills:
Amazon Web Services, Rust, React, Typescript, Terraform, Nodejs, Pulumi, Azure, GitHub Actions, Infrastructure-as-code solutions, Go
12-14 yrs
Bengaluru, India
Skills:
Data Architecture, API design, Distributed Systems, System Architecture, Fault Tolerance, service-oriented systems, data integrity, AI-enabled development workflows, event-driven architectures, cloud-native technologies, platform engineering, asynchronous processing, system integrations, Operational Excellence, observability