DevOps Engineer - AWS & AI
Location: Pune, India
Company Type: Product-Based
About the Opportunity
We are hiring for a fast-growing product-based company based in Pune, looking for a skilled DevOps Engineer I to design, build, and scale cloud infrastructure on Amazon Web Services (AWS). This role carries a strong focus on AI workloads, LLMOps, and hands-on involvement in high-impact AI-driven product initiatives. Candidates with demonstrated experience working in product-based companies and delivering AI projects at scale are strongly preferred.
Role Overview
You will be embedded in a cross-functional engineering team responsible for building, maintaining, and scaling infrastructure that supports both production workloads and internal AI/ML systems. This role requires strong ownership, sound architectural judgment, and a relentless focus on reliability, security, and cost efficiency across AWS-native services.
Key Responsibilities
- Infrastructure Management: Own and maintain production AWS infrastructure with high availability, fault tolerance, and proactive monitoring.
- IaC & Automation: Build and manage scalable cloud infrastructure using Terraform and Ansible on AWS.
- AI/LLM Support: Deploy and manage AI/LLM workloads, vector databases, and model inference pipelines.
- CI/CD Pipelines: Design and maintain secure CI/CD pipelines for microservices and AI/ML systems.
- Security & Compliance: Implement security best practices including IAM policies, secrets management, and vulnerability scanning.
- Observability: Set up monitoring for both system and AI metrics (latency, throughput, cost, usage).
- Cost Optimization: Optimize cloud costs, particularly for GPU-backed and compute-intensive AI workloads.
- Networking: Manage VPC networking, DNS, load balancing, and secure connectivity across AWS environments.
- Cross-Functional Collaboration: Collaborate closely with AI/ML engineers and product teams to ensure infrastructure supports fast experimentation and reliable deployment.
Required Qualifications
- Experience: 3–5 years of hands-on experience in DevOps, SRE, or cloud infrastructure roles.
- Background: Proven experience working at a product-based company with fast-paced engineering teams.
- AWS Expertise: Strong hands-on expertise with core AWS services: EKS, ECS/Fargate, EC2, S3, RDS, SageMaker, Lambda, VPC, and CloudFront.
- AI Stack: Experience deploying and managing AI/LLM workloads, vector databases (e.g., Pinecone, Weaviate, pgvector), and LLM APIs (OpenAI, Bedrock, etc.).
- Containerization & IaC: Solid expertise in Docker, Kubernetes (EKS), Terraform, and Ansible.
- Automation: Proficiency in Python scripting for automation, tooling, and infrastructure workflows (Mandatory Requirement).
- OS & Networking: Strong understanding of Linux/Unix systems (preferably Ubuntu), networking fundamentals, security architecture, and IAM principles on AWS.
- Track Record: Demonstrated experience delivering AI projects in production—not just experimental or PoC environments.
Preferred Skills
- LLMOps: Experience with prompt versioning, model monitoring, RAG pipelines, and inference optimization.
- ML Tooling: Familiarity with AI/ML frameworks (SageMaker Pipelines, MLflow, Kubeflow) and model serving patterns.
- Monitoring Stack: Hands-on experience with observability tools such as Prometheus, Grafana, CloudWatch, or OpenTelemetry.
- AWS AI Services: Exposure to AWS Bedrock, Rekognition, Comprehend, or other managed AI/ML services.
- Problem Solving: Strong debugging, performance tuning, and root cause analysis skills with an automation-first mindset.
Soft Skills
- Excellent verbal and written communication skills with a habit of thorough documentation.
- Strong analytical thinking and structured problem-solving approach.
- Quick learner with the ability to adapt to evolving AI tooling and cloud-native ecosystems.
- Collaborative team player with the ability to mentor peers and contribute to engineering culture.