Search by job, company or skills

Operations Engineer

Early Applicant
  • Posted 19 hours ago
  • Be among the first 10 applicants

Job Description

Project Role : Operations Engineer

Project Role Description : Support the operations and/or manage delivery for production systems and services based on operational requirements and service agreement.

Must have skills : Agentic Orchestration, Agentic AI, AI Orchestration &LLM

Good to have skills : NA

Minimum 3 Year(s) Of Experience Is Required

Educational Qualification : 15 years full time education

Summary:

A Tools & Platforms Service Engineer is responsible for the day-to-day operation, monitoring, support, and maintenance of the infrastructure engineering tooling estate — spanning observability platforms, infrastructure-as-code tooling, CI/CD pipelines, ITSM platforms, internal developer portals, secret management, and AI-augmented operations tooling. The role ensures that the platforms enabling Cloud, Network, Security, Database, and Voice teams to operate at scale remain available, performant, and fit for purpose.

At Level 9 / 10, this individual forms the primary operational execution layer within the tools and platforms tower — a proficient hands-on engineer who keeps platform tooling running reliably, resolves incidents and service requests efficiently, executes changes cleanly, and supports the adoption of AI-augmented operational capabilities across the infrastructure engineering function.

PLATFORMS IN SCOPE

Observability:

– ELK/Splunk

– OpenTelemetry

– Alertmanager

Automation & IaC:

– Terraform (consumer)

– Ansible playbooks

– GitHub Actions

– HashiCorp Vault

ITSM & DevOps:

– ServiceNow

– Jira/Confluence

– Xmatters

– CMDB tooling

AI-Augmented Ops:

– LLM runbook support

– AI alert triage

– Chatbot ops interfaces

– Prompt execution

– LLMOps monitoring

Roles & Responsibilities:

Observability & Monitoring Platform Operations

– Monitor and maintain observability platforms — Splunk/ELK — ensuring dashboards, alert rules, and data pipelines remain operational and accurate across all infrastructure tiers

– Triage and resolve observability platform incidents — broken scrapers, missing metrics, failed log ingestion, dashboard errors — within defined SLA commitments

– Manage alert configuration — creating, updating, and tuning alert rules in Alertmanager, or equivalent to reduce noise and maintain actionable alerting across infrastructure teams

– Operate ELK stack or Splunk log pipelines — managing index lifecycle, ingest pipeline health, and log source connectivity for infrastructure platforms

– Execute observability platform changes — dashboard deployments, alert rule updates, retention policy changes — following approved change management procedures

– Maintain runbooks and operational documentation for all observability platforms — keeping procedures current, accurate, and audit-ready

ITSM, DevOps Tooling & AI-Augmented Operations

– Manage Xmatters — maintaining escalation policies, on-call schedules, alert routing rules, and integration health with upstream observability platforms

– Support Jira and Confluence platform operations — managing project configurations, workflow rules, automation triggers, and user access across engineering teams

– Monitor and support AI-augmented operations tooling — LLM runbook automation pipelines, agentic ITSM workflow health, and AI alert correlation services — ensuring AI tools are operational and escalating failures to Principal Engineers

– Support engineering teams using AI-augmented operations capabilities — triaging issues with LLM-powered runbook execution, RAG knowledge base queries, and chatbot operational interfaces

– Monitor LLMOps pipelines — tracking prompt execution logs, model API health, token usage, and cost anomalies — escalating degradation or unexpected behaviour promptly

– Perform root cause analysis (RCA) for recurring platform incidents across observability, IaC, CI/CD, ITSM, and AI tooling — documenting findings and driving preventive actions

– Support audit, regulatory, and compliance activities — producing platform access logs, configuration exports, and change evidence for PCI DSS, SOX, and FSI regulatory review

Professional & Technical Skills:

Must-Have Technical Skills

Observability Operations: Splunk — dashboard management, alert rule configuration, metric scraper troubleshooting, and log pipeline operations (ELK/Splunk)

Terraform Operations: Plan/apply execution, state file management, drift alert response, module consumption support, and workspace administration across AWS and GCP environments

CI/CD Operations: GitHub Actions and ArgoCD — pipeline health monitoring, runner management, deployment failure triage, and GitOps workflow support for infrastructure and application teams

HashiCorp Vault: Secret lease monitoring, auth method health, cluster status monitoring, and escalation-ready diagnosis of secrets management failures

Xmatters: On-call schedule management, escalation policy configuration, alert routing, and integration health with upstream monitoring platforms

AI Tooling Operations: LLM runbook pipeline monitoring, agentic ITSM workflow health, LLMOps dashboard operation, token/cost anomaly alerting, and AI tool incident triage

Scripting & Automation: Python or Bash — operational scripting for platform task automation, API-driven tooling management, and alert-driven remediation workflows

Incident & Change Management: Structured P1–P3 platform incident handling, ITIL-aligned change execution, post-change verification, and RCA documentation in FSI environments

CMDB & Asset Management: Automated discovery tool operations, asset record accuracy monitoring, lifecycle tracking, and CMDB reconciliation for infrastructure tooling assets

Preferred / Advantageous

– Familiarity with LLM APIs (OpenAI, Anthropic, Google Gemini)

— sufficient to triage AI tooling incidents and support engineering teams with prompt execution issues

– Experience with Ansible or Chef for configuration management task execution and playbook troubleshooting – Exposure to Kubernetes operations (EKS or GKE)

— sufficient to support platform tooling deployed on container platforms

– Familiarity with observability-as-code practices — dashboard and alert rules managed as version-controlled configurations

Additional Information:

– Observability, IaC, CI/CD, ITSM, and AI tooling platforms are available and performing — incidents are resolved within SLA with accurate RCA and preventive actions that reduce recurrence

– Engineering teams across Cloud, Network, Security, Database, and Voice towers experience reliable, responsive platform tooling — support requests are resolved promptly and platform outages are communicated clearly

– Changes across the tooling estate are executed cleanly — validated, implemented without unplanned impact, and verified post-completion with documented evidence

– AI-augmented operations tooling is operational and monitored — LLM pipelines, agentic workflows, and alert correlation services are healthy, and degradation is escalated before it impacts engineering teams

– Platform access, audit logs, and CMDB records are consistently accurate — supporting PCI DSS, SOX, and regulatory review without remediation effort

  • The candidate should have minimum 3 years of experience in Agentic Orchestration.
  • This position is based at our Bengaluru office.
  • A 15 years full time education is required.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152534803

Similar Jobs

Bengaluru, India

Skills:

ServicenowElkConfluenceTerraformPythonPulumiJiraAnsibleSplunkHashiCorp VaultChefLlmRAG knowledge basesGitHub ActionsOpenTelemetryLLM runbook automationAgentic OrchestrationAI OrchestrationZipkinAI alert correlationJaegerLLMOps pipelinesAgentic ITSM workflowsXmatersAgentic AIArgoCD

Beware of Scammers

We don’t charge money for job offers