Search by job, company or skills

Operations Engineer

  • Posted 19 hours ago
  • Be among the first 10 applicants

Job Description

Project Role : Operations Engineer

Project Role Description : Support the operations and/or manage delivery for production systems and services based on operational requirements and service agreement.

Must have skills : Splunk Administration, Splunk Enterprise Observability, ITSI and Event management with AIOPS

Good to have skills : Splunk Security Information and Event Management (SIEM), Splunk Enterprise Architecture and Design

Minimum 7.5 Year(s) Of Experience Is Required

Educational Qualification : 15 years full time education

Summary:

A Principal Engineer — Tools & Platforms provides expert technical leadership across the infrastructure tooling ecosystem — spanning observability, automation, infrastructure-as-code, ITSM platforms, internal developer portals, and AI-augmented operations. The role is accountable for building, governing, and continuously improving the engineering platforms that enable Cloud, Network, Security, Database, and Voice towers to operate at scale with reliability. A key differentiator of this role is the ability to embed AI-augmented capabilities — LLM-driven runbook automation, agentic ITSM workflows, and intelligent alert correlation — directly into infrastructure platforms, accelerating toil reduction and operational intelligence across the organization.

Observability

– ELK/Splunk

– OpenTelemetry

– Jaeger/Zipkin

Automation & IaC

– Terraform/Pulumi

– Ansible/Chef

– GitHub Actions/ArgoCD

– HashiCorp Vault

ITSM & DevOps

– ServiceNow

– Jira/Confluence

– Xmaters

AI-Augmented Ops

– LLM runbook automation

– Agentic ITSM workflows

– AI alert correlation

– RAG knowledge bases

– LLMOps pipelines

Roles & Responsibilities:

– Design, build, and operate enterprise observability platforms — Splunk — providing unified metrics, logging, and tracing across all infrastructure tiers with OpenTelemetry instrumentation standards

– Engineer and govern the IaC platform — Terraform or Pulumi — including module libraries, state management, policy-as-code (Sentinel/OPA), drift detection, and multi-cloud coverage across AWS and GCP

– Build and maintain CI/CD and GitOps pipeline infrastructure — GitHub Actions, ArgoCD — for infrastructure and application platform delivery operate the Internal Developer Portal (Backstage) for self-service provisioning

– Own ITSM platform engineering — ServiceNow workflow automation, CMDB reconciliation, PagerDuty/OpsGenie integrations, and automated incident creation pipelines across all infrastructure towers

– Engineer LLM-powered runbook automation — integrating OpenAI, Anthropic Claude, or Google Gemini with infrastructure tooling to automate incident triage, resolution workflows, and knowledge base queries

– Build agentic ITSM workflows — multi-agent orchestration (LangChain, LlamaIndex, CrewAI) that autonomously handle incident classification, runbook execution, and stakeholder communication with human-in-the-loop controls

– Implement AI-powered alert correlation and noise reduction — LLM-based reasoning to group related alerts, suppress false positives, and surface actionable incident context to on-call engineers

– Design RAG architectures for infrastructure knowledge bases — enabling engineers to query runbooks, operational docs, and incident history through natural language interfaces backed by vector databases

– Engineer secret management infrastructure — HashiCorp Vault — dynamic secrets, PKI automation, and Kubernetes auth across multi-cloud and on-premises environments

– Mentor engineers across all towers set platform engineering standards, drive tooling roadmaps, and act as the highest technical escalation point for platform and AI tooling issues

Professional & Technical Skills:

Mandatory Certifications

  • Terraform Associate or Professional- Splunk Professional -AWS DevOps Engineer Pro or GCP DevOps Engineer
  • Certified Kubernetes Administrator (CKA)- HashiCorp Vault Associate
  • Observability: Splunk — SLO engineering, OpenTelemetry, ELK/Splunk log pipelines, distributed tracing
  • IaC & Automation: Terraform (advanced) — modular design, policy-as-code, drift detection Ansible/Chef for configuration management
  • CI/CD & GitOps: GitHub Actions, ArgoCD — pipeline engineering, GitOps workflows, policy enforcement, infrastructure deployment automation
  • ITSM Engineering: ServiceNow — workflow automation, CMDB integration, incident/problem/change API integration, PagerDuty/OpsGenie connectivity
  • Internal Dev Portal: service catalogue, golden path templates, self-service provisioning, and plugin development
  • Secret Management: HashiCorp Vault — dynamic secrets, PKI automation, Kubernetes auth, multi-cloud secrets lifecycle
  • LLM & AI Tooling: OpenAI, Anthropic Claude, Google Gemini — prompt engineering, RAG architecture, vector databases (Pinecone, Weaviate, ChromaDB), LLMOps
  • AI Orchestration: LangChain, LlamaIndex, CrewAI — agentic workflow design, multi-agent orchestration, tool calling, human-in-the-loop workflows
  • Scripting & Dev: Python (advanced) — platform tooling development, LLM SDK integration, infrastructure automation, API development
  • FSI Compliance: PCI DSS, SOX, DORA — platform access controls, audit logging, policy-as-code enforcement, and regulatory evidence production

Preferred / Advantageous


– Crossplane for Kubernetes-native cloud resource provisioning

– eBPF-based observability — Cilium, Pixie, or Hubble — for deep infrastructure telemetry

– AI/ML infrastructure — GPU node management, model serving platforms (Triton, vLLM), or MLflow

– Platform engineering or SRE consulting background within financial services

Additional Information:

Infrastructure Platform Engineering & Tooling

– Observability, IaC, CI/CD, ITSM, and IDP platforms are reliable, well-governed, and actively adopted across all infrastructure towers — with measurable reductions in operational toil

– AI-augmented operations are deployed in production — runbook automation, agentic ITSM workflows, and LLM-powered alert correlation reduce MTTR and free engineers from repetitive work – IaC coverage is comprehensive and drift-free — environments reproducible, policy-compliant, and pipeline-deployable without manual intervention

– Engineering teams across all towers regard the Principal Engineer as the authoritative platform expert who removes blockers, sets the standard, and makes hard problems tractable

  • 12–15+ years in platform, tools, or infrastructure engineering
  • The candidate should have minimum 7.5 years of experience in Agentic Orchestration.
  • This position is based at our Bengaluru office.
  • A 15 years full time education is required.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152534801

Similar Jobs

Bengaluru, India

Skills:

ServicenowElkConfluenceTerraformPythonPulumiJiraAnsibleSplunkHashiCorp VaultChefLlmRAG knowledge basesGitHub ActionsOpenTelemetryLLM runbook automationAgentic OrchestrationAI OrchestrationZipkinAI alert correlationJaegerLLMOps pipelinesAgentic ITSM workflowsXmatersAgentic AIArgoCD

Bengaluru, India

Skills:

security automation Cloud SecurityNetworkingBashDatadogIncident ResponseLog AnalysisThreat HuntingLinuxSiemPythonAWSDetection EngineeringElastic

Beware of Scammers

We don’t charge money for job offers