Search by job, company or skills

AI Infrastructure Architect (AWS AI Services | AWS Cloud Architecture)

AI Infrastructure Architect (AWS AI Services | AWS Cloud Architecture)

LTM
12-14 Years
Not Disclosed
Early Applicant
  • Posted 6 days ago
  • Be among the first 10 applicants

Job Description

Role Description

Role Overview

We are seeking a highly experienced AI Infrastructure Architect to design, build, and govern scalable, secure, resilient, and cost-efficient AI platforms across AWS, Microsoft Azure, and Google Cloud Platform. The role will lead end-to-end architecture for AI/ML workloads, including data platforms, model training, fine-tuning, inference, MLOps, GPU infrastructure, observability, security, and governance for enterprise and production-grade use cases.

The ideal candidate combines deep cloud infrastructure expertise, hands-on AI/ML platform knowledge, and strong enterprise architecture leadership across multi-cloud environments.

Key Responsibilities AI/ML Infrastructure Architecture

  • Lead the design of end-to-end AI infrastructure for model experimentation, training, fine-tuning, inference, deployment, monitoring, and ongoing operations.
  • Architect scalable platforms for batch and real-time ML workloads, LLM-based solutions, Generative AI pipelines, and enterprise AI applications.
  • Define standards for model experimentation, versioning, registry, promotion, lifecycle management, and retirement.
  • Create reusable reference architectures, blueprints, guardrails, and design patterns for AI workloads.
  • Ensure platforms meet scalability, availability, performance, disaster recovery, and operational requirements.

Multi-Cloud Platform Design: AWS, Azure & GCP

  • Architect cloud-native and cloud-agnostic AI platforms across AWS SageMaker, EKS, EC2 GPU, S3 and IAM; Azure Machine Learning, AKS, Azure OpenAI and GPU VM series; and GCP Vertex AI, GKE and TPU/GPU infrastructure.
  • Define workload placement principles based on capability, security, latency, resilience, portability, cost, and strategic vendor alignment.
  • Enable workload portability and standardized operating practices across cloud environments.
  • Define hybrid and multi-cloud AI operating models, including connectivity, identity, observability, governance, and disaster recovery.

MLOps, DevOps & Platform Engineering

  • Establish MLOps frameworks for CI/CD and continuous training of models, pipelines, features, and AI applications.
  • Design automation for model lifecycle management, experiment tracking, model registry, validation, deployment, rollback, and retraining.
  • Implement monitoring for service health, model performance, drift, data quality, latency, throughput, reliability, and cost.
  • Integrate AI delivery pipelines with enterprise DevOps, platform engineering, security, and change-management standards.

Data & Compute Architecture

  • Design scalable data ingestion, feature stores, training datasets, data lakes, and data access patterns for AI/ML workloads.
  • Architect accelerator strategies using NVIDIA GPUs, TPUs, and fit-for-purpose inference compute.
  • Optimize utilization, performance, scheduling, capacity, and cost for training and inference workloads.
  • Define storage, networking, caching, and distributed-compute patterns for large-scale AI platforms.

Security, Governance & Compliance

  • Define AI security architecture covering identity, privileged access, data access, network isolation, secrets and key management, encryption, supply-chain security, and tenant/workload isolation.
  • Implement governance controls for model usage, data privacy, lineage, approvals, responsible AI, risk management, and compliance.
  • Align AI platforms with enterprise security architecture, regulatory obligations, audit requirements, and internal governance frameworks.
  • Embed security-by-design, policy-as-code, traceability, and evidence collection into platform workflows.

Leadership & Advisory

  • Act as the technical authority for AI infrastructure and platform architecture decisions.
  • Guide cloud architects, platform engineers, data engineers, ML engineers, security teams, and application teams.
  • Support AI platform roadmaps, cloud strategy, capability assessments, investment decisions, and architecture reviews.
  • Communicate architectural choices, trade-offs, risks, and recommendations to business, engineering, and leadership stakeholders.
  • Mentor teams and promote reusable engineering practices and architecture standards.

Core Technical Skills Cloud Platforms & Architecture

  • Advanced architecture expertise across AWS, Microsoft Azure, and GCP.
  • Strong experience in cloud networking, IAM, security architecture, landing zones, resilience, and multi-cloud governance.

AI/ML Platforms

  • Hands-on experience designing and deploying enterprise AI/ML infrastructure.
  • Expertise with Azure Machine Learning, AWS SageMaker, and GCP Vertex AI.
  • Experience with Generative AI and LLM platforms supporting training, fine-tuning, evaluation, inference, and monitoring.

Infrastructure & Platform Engineering

  • Kubernetes expertise across EKS, AKS, and GKE.
  • GPU/accelerator infrastructure architecture, scheduling, performance tuning, capacity management, and cost optimization.
  • Infrastructure as Code using Terraform, ARM/Bicep and/or CloudFormation.
  • Containerization, platform automation, service mesh, networking, and observability.

MLOps & Automation

  • CI/CD and continuous training for ML pipelines and AI applications.
  • Model registry, experiment tracking, feature/pipeline versioning, deployment automation, and inference scaling.
  • Monitoring, logging, ing, drift detection, reliability engineering, and performance tuning.

Data Systems

  • Large-scale data platforms for AI/ML workloads, including batch and streaming architectures.
  • Feature stores, data ingestion, data quality, lineage, governance, and secure data-access patterns.
  • Strong understanding of distributed systems and high-performance computing concepts.

Preferred Qualifications

  • Experience designing and governing enterprise AI platforms at scale.
  • Exposure to Responsible AI frameworks, model risk management, and AI governance operating models.
  • Strong background in cost optimization and FinOps for GPU-intensive AI workloads.
  • Consulting, client-facing advisory, or architecture review experience.
  • Experience supporting regulated industries.
  • Relevant certifications in AWS, Azure, GCP, Kubernetes, enterprise architecture, security, or AI/ML.

Education & Experience

  • Bachelor's or Master's degree in Computer Science, Engineering, Information Technology, or a related field.
  • 12+ years of overall experience in infrastructure, cloud, platform engineering, or enterprise architecture.
  • Proven experience leading AI/ML infrastructure architecture initiatives from strategy through production adoption.
  • Strong architectural judgement, written and verbal communication, and stakeholder-management capability.

Key Competencies

  • Enterprise architecture leadership
  • Strategic thinking and decision-making
  • Multi-cloud architecture and governance
  • AI platform engineering and MLOps
  • Security, resilience, compliance, and cost management
  • Executive communication and stakeholder influence
  • Technical mentorship and cross-functional collaboration

Job Location: Coimbatore

More Info

Job Type:
Industry:
Employment Type:

Key Skills

feature stores

GKE

GPU infrastructure

data platforms

AKS

CI CD

EKS

GCP Vertex AI

AWS SageMaker

data ingestion

Bicep

AI ML infrastructure

About Company