Search by job, company or skills

Operations Engineer

  • Posted 9 hours ago
  • Be among the first 10 applicants

Job Description

Project Role : Operations Engineer

Project Role Description : Support the operations and/or manage delivery for production systems and services based on operational requirements and service agreement.

Must have skills : Amazon Web Services (AWS)

Good to have skills : NA

Minimum 12 Year(s) Of Experience Is Required

Educational Qualification : 15 years full time education

Summary:

A Principal Engineer — Cloud leads the hands-on design, build, and operation of large-scale cloud infrastructure across AWS and GCP. The role is accountable for engineering secure, resilient, and cost-optimized cloud platforms — building real infrastructure, writing real code, and owning outcomes end to end alongside product, security, and application engineering teams.

This individual operates as a recognized cloud engineering authority — setting platform standards, driving infrastructure automation, and solving complex distributed cloud Infrastructure challenges including having good understanding of the insurance industry. In a financial services context, engineering decisions must satisfy stringent regulatory requirements including PCI DSS, SOX, and DORA, alongside enterprise-grade reliability and availability commitments.

Roles & Responsibilities:

Cloud Platform Engineering

– Design, build, and operate large-scale, cloud-native infrastructure across AWS and GCP — owning platform reliability, security posture, and cost efficiency

– Implement and maintain infrastructure-as-code across both cloud platforms using Terraform and native tooling — enforcing standards, modularity, and drift detection

– Build and maintain CI/CD pipelines, GitOps workflows, and deployment automation for cloud infrastructure and platform services

– Engineer landing zones, account vending, and governance guardrails on AWS (Control Tower, SCPs) and GCP (Resource Manager, Org Policies)

AI Orchestration: High Level understanding of LangChain, LlamaIndex, CrewAI — agentic workflow design, multi-agent orchestration, tool calling, human-in-the-loop workflows

– Design and implement cloud networking — AWS VPC, Transit Gateway, Direct Connect and GCP VPC, Shared VPC, Cloud Interconnect — for hybrid and multi-cloud connectivity

– Implement security controls across IAM, network policy, encryption, secrets management, and audit logging on both platforms

– Drive platform observability — metrics, logging, distributed tracing, and alerting using CloudWatch, GCP Operations Suite, and third-party tooling

– Perform capacity planning, cost analysis, and right-sizing to maintain efficient cloud spend across AWS and GCP estates

LLM & AI Tooling: High Level Understanding of OpenAI, Anthropic Claude, Google Gemini — prompt engineering, RAG architecture, vector databases (Pinecone, Weaviate, ChromaDB), LLMOps

Application & Container Platform Engineering

– Engineer and operate Kubernetes-based container platforms — EKS on AWS and GKE on GCP — including cluster lifecycle, node management, networking, and security hardening

– Define and enforce standards for microservices, APIs, event-driven architectures, and data platform integrations across cloud environments

– Collaborate with application engineering teams on cloud-native design patterns — serverless, service mesh, and distributed data stores – Build and maintain developer platform tooling — internal developer portals, self-service infrastructure, and platform APIs — to accelerate engineering delivery

– Drive observability and reliability engineering practices — SLIs, SLOs, error budgets, and chaos engineering — across cloud-hosted services

– Support secure software delivery — SAST, DAST, container image scanning, and supply chain security in CI/CD pipelines

Technical Leadership & Engineering Excellence

– Set cloud engineering standards, patterns, and reference architectures — codified in reusable Terraform modules, runbooks, and design guides

– Lead proof-of-concept and spike work for emerging cloud technologies — making build-vs-buy recommendations grounded in engineering evidence

– Conduct technical design reviews and production readiness assessments for cloud-hosted services

– Mentor senior and mid-level engineers — providing hands-on guidance on IaC, cloud-native patterns, and production operations

– Drive continuous improvement of the cloud platform — reducing toil through automation, improving reliability, and eliminating technical debt

– Contribute engineering patterns and reusable modules back to internal platform teams and cloud centre of excellence

Professional & Technical Skills:

Must-Have Technical Skills

Infrastructure as Code: Terraform (advanced) — modular design, remote state, workspace strategy, and policy-as-code (Sentinel/OPA) across AWS and GCP

Containers & Kubernetes: EKS and GKE — cluster lifecycle, CNI, RBAC, admission controllers, network policies, and production-grade operations

Cloud Networking: Deep expertise in hybrid connectivity, multi-cloud routing, DNS strategy, and network security across both platforms

Security Engineering: IAM design, zero-trust principles, encryption at rest/in transit, secrets management (Vault, AWS Secrets Manager, GCP Secret Manager), and CSPM

CI/CD & Automation: GitHub Actions, ArgoCD, Terraform Cloud, or equivalent — GitOps workflows, pipeline security, and infrastructure deployment automation

AI Orchestration: High Level Understanding of LangChain, LlamaIndex, CrewAI — agentic workflow design, multi-agent orchestration, tool calling, human-in-the-loop workflows

Observability: Metrics, logging, and distributed tracing at scale — SLO engineering, alerting design, and incident response tooling

Scripting & Development: Python or Go — automation scripting, SDK usage for AWS and GCP, and platform tooling development

LLM & AI Tooling: High Level understanding of OpenAI, Anthropic Claude, Google Gemini — prompt engineering, RAG architecture, vector databases (Pinecone, Weaviate, ChromaDB), LLMOps

Incident Management: Production incident ownership, structured RCA, and reliability engineering practices in high-availability FSI environments

Cost Engineering: FinOps practices — reserved instance/committed use optimisation, tagging strategy, and cost allocation across AWS and GCP estates

Preferred / Advantageous

– Experience with service mesh (Istio or Linkerd) for microservices networking and observability

– Familiarity with data platform engineering — Kafka, Spark, or cloud-native analytics services on AWS or GCP

– Exposure to FinOps tooling — CloudHealth, Apptio Cloudability, or native cloud cost management platforms

– Background in platform engineering, SRE, or cloud infrastructure consulting within financial services

Additional Information:

Certifications

  • AWS Solutions Architect

— Professional GCP Professional Cloud Architect AWS DevOps Engineer

— Professional or GCP DevOps Engineer

– Cloud platforms on AWS and GCP are secure, observable, and reliably operated — meeting FSI SLA and regulatory requirements consistently

– Infrastructure-as-code coverage is comprehensive — environments are reproducible, drift-free, and deployable through automated pipelines without manual intervention

– Engineering teams across the organisation adopt platform standards, reusable modules, and cloud-native patterns — reducing bespoke build effort and improving consistency – Platform reliability improves measurably quarter-on-quarter — MTTR reduces, toil decreases, and error budgets are maintained across cloud-hosted services

– Cross-functional stakeholders — security, application, data, and operations teams — regard the Principal Engineer as a technically authoritative, delivery-focused partner who unblocks rather than bottlenecks

  • The candidate should have minimum 12 years of experience in Amazon Web Services (AWS).
  • This position is based at our Bengaluru office.
  • A 15 years full time education is required.


More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152473009

Similar Jobs

Bengaluru, India

Skills:

Ibm MqPostgreSQLKafkaMongoDBApi GatewayIaC AutomationService MeshOracle EngineeringMuleSoft Anypoint PlatformCloud DB ServicesSQL Server EngineeringObservabilityFSI Compliance

Bengaluru, India

Skills:

data engineering snowflake ApisPower BiRpachange managementInformaticaData QualityAzure Data FactoryPythonplatform reliability engineeringETL ELT technologiesAI GenAI toolsincident problem managementCI CDLookerRcaDataOpsmetadata lineageFivetranobservabilityGovernance

Bengaluru

Skills:

TerraformCloudformationCDK

Beware of Scammers

We don’t charge money for job offers