Project Role : Operations Engineer
Project Role Description : Support the operations and/or manage delivery for production systems and services based on operational requirements and service agreement.
Must have skills : Amazon Web Services (AWS)
Good to have skills : NA
Minimum 12 Year(s) Of Experience Is Required
Educational Qualification : 15 years full time education
Summary:
A Principal Engineer — Cloud leads the hands-on design, build, and operation of large-scale cloud infrastructure across AWS and GCP. The role is accountable for engineering secure, resilient, and cost-optimized cloud platforms — building real infrastructure, writing real code, and owning outcomes end to end alongside product, security, and application engineering teams.
This individual operates as a recognized cloud engineering authority — setting platform standards, driving infrastructure automation, and solving complex distributed cloud Infrastructure challenges including having good understanding of the insurance industry. In a financial services context, engineering decisions must satisfy stringent regulatory requirements including PCI DSS, SOX, and DORA, alongside enterprise-grade reliability and availability commitments.
Roles & Responsibilities:
Cloud Platform Engineering
– Design, build, and operate large-scale, cloud-native infrastructure across AWS and GCP — owning platform reliability, security posture, and cost efficiency
– Implement and maintain infrastructure-as-code across both cloud platforms using Terraform and native tooling — enforcing standards, modularity, and drift detection
– Build and maintain CI/CD pipelines, GitOps workflows, and deployment automation for cloud infrastructure and platform services
– Engineer landing zones, account vending, and governance guardrails on AWS (Control Tower, SCPs) and GCP (Resource Manager, Org Policies)
AI Orchestration: High Level understanding of LangChain, LlamaIndex, CrewAI — agentic workflow design, multi-agent orchestration, tool calling, human-in-the-loop workflows
– Design and implement cloud networking — AWS VPC, Transit Gateway, Direct Connect and GCP VPC, Shared VPC, Cloud Interconnect — for hybrid and multi-cloud connectivity
– Implement security controls across IAM, network policy, encryption, secrets management, and audit logging on both platforms
– Drive platform observability — metrics, logging, distributed tracing, and alerting using CloudWatch, GCP Operations Suite, and third-party tooling
– Perform capacity planning, cost analysis, and right-sizing to maintain efficient cloud spend across AWS and GCP estates
LLM & AI Tooling: High Level Understanding of OpenAI, Anthropic Claude, Google Gemini — prompt engineering, RAG architecture, vector databases (Pinecone, Weaviate, ChromaDB), LLMOps
Application & Container Platform Engineering
– Engineer and operate Kubernetes-based container platforms — EKS on AWS and GKE on GCP — including cluster lifecycle, node management, networking, and security hardening
– Define and enforce standards for microservices, APIs, event-driven architectures, and data platform integrations across cloud environments
– Collaborate with application engineering teams on cloud-native design patterns — serverless, service mesh, and distributed data stores – Build and maintain developer platform tooling — internal developer portals, self-service infrastructure, and platform APIs — to accelerate engineering delivery
– Drive observability and reliability engineering practices — SLIs, SLOs, error budgets, and chaos engineering — across cloud-hosted services
– Support secure software delivery — SAST, DAST, container image scanning, and supply chain security in CI/CD pipelines
Technical Leadership & Engineering Excellence
– Set cloud engineering standards, patterns, and reference architectures — codified in reusable Terraform modules, runbooks, and design guides
– Lead proof-of-concept and spike work for emerging cloud technologies — making build-vs-buy recommendations grounded in engineering evidence
– Conduct technical design reviews and production readiness assessments for cloud-hosted services
– Mentor senior and mid-level engineers — providing hands-on guidance on IaC, cloud-native patterns, and production operations
– Drive continuous improvement of the cloud platform — reducing toil through automation, improving reliability, and eliminating technical debt
– Contribute engineering patterns and reusable modules back to internal platform teams and cloud centre of excellence
Professional & Technical Skills:
Must-Have Technical Skills
Infrastructure as Code: Terraform (advanced) — modular design, remote state, workspace strategy, and policy-as-code (Sentinel/OPA) across AWS and GCP
Containers & Kubernetes: EKS and GKE — cluster lifecycle, CNI, RBAC, admission controllers, network policies, and production-grade operations
Cloud Networking: Deep expertise in hybrid connectivity, multi-cloud routing, DNS strategy, and network security across both platforms
Security Engineering: IAM design, zero-trust principles, encryption at rest/in transit, secrets management (Vault, AWS Secrets Manager, GCP Secret Manager), and CSPM
CI/CD & Automation: GitHub Actions, ArgoCD, Terraform Cloud, or equivalent — GitOps workflows, pipeline security, and infrastructure deployment automation
AI Orchestration: High Level Understanding of LangChain, LlamaIndex, CrewAI — agentic workflow design, multi-agent orchestration, tool calling, human-in-the-loop workflows
Observability: Metrics, logging, and distributed tracing at scale — SLO engineering, alerting design, and incident response tooling
Scripting & Development: Python or Go — automation scripting, SDK usage for AWS and GCP, and platform tooling development
LLM & AI Tooling: High Level understanding of OpenAI, Anthropic Claude, Google Gemini — prompt engineering, RAG architecture, vector databases (Pinecone, Weaviate, ChromaDB), LLMOps
Incident Management: Production incident ownership, structured RCA, and reliability engineering practices in high-availability FSI environments
Cost Engineering: FinOps practices — reserved instance/committed use optimisation, tagging strategy, and cost allocation across AWS and GCP estates
Preferred / Advantageous
– Experience with service mesh (Istio or Linkerd) for microservices networking and observability
– Familiarity with data platform engineering — Kafka, Spark, or cloud-native analytics services on AWS or GCP
– Exposure to FinOps tooling — CloudHealth, Apptio Cloudability, or native cloud cost management platforms
– Background in platform engineering, SRE, or cloud infrastructure consulting within financial services
Additional Information:
Certifications
— Professional GCP Professional Cloud Architect AWS DevOps Engineer
— Professional or GCP DevOps Engineer
– Cloud platforms on AWS and GCP are secure, observable, and reliably operated — meeting FSI SLA and regulatory requirements consistently
– Infrastructure-as-code coverage is comprehensive — environments are reproducible, drift-free, and deployable through automated pipelines without manual intervention
– Engineering teams across the organisation adopt platform standards, reusable modules, and cloud-native patterns — reducing bespoke build effort and improving consistency – Platform reliability improves measurably quarter-on-quarter — MTTR reduces, toil decreases, and error budgets are maintained across cloud-hosted services
– Cross-functional stakeholders — security, application, data, and operations teams — regard the Principal Engineer as a technically authoritative, delivery-focused partner who unblocks rather than bottlenecks
- The candidate should have minimum 12 years of experience in Amazon Web Services (AWS).
- This position is based at our Bengaluru office.
- A 15 years full time education is required.