Search by job, company or skills

Operations Engineer

8-10 Years
  • Posted 7 hours ago
  • Be among the first 10 applicants

Job Description

Project Role : Operations Engineer

Project Role Description : Support the operations and/or manage delivery for production systems and services based on operational requirements and service agreement.

Must have skills : Site Reliability Engineering

Good to have skills : NA

Minimum 7.5 Year(s) Of Experience Is Required

Educational Qualification : 15 years full time education

Summary:

As an Operations Engineer, a typical day involves overseeing the smooth functioning of production systems and services, ensuring they meet operational requirements and service agreements. This role requires continuous monitoring, managing delivery processes, and promptly addressing any issues that arise to maintain system reliability and performance. The position demands coordination with various teams to uphold service standards and support operational excellence throughout the production environment.

Role Title: Site Reliability Engineer (SRE)

Role Summary:

The Site Reliability Engineer (SRE) is a hands-on engineer responsible for improving the reliability, availability, performance, and operational efficiency of the assigned technology tower. The role blends deep tower-specific technical expertise with strong automation capability (Python and Ansible) and modern SRE practices. The SRE identifies high-value automation opportunities, eliminates repetitive manual effort (toil), and drives use cases through feasibility, build, test, deployment, and hyper-care, while partnering with operations and engineering teams to meet service-level objectives.

Key Responsibilities:

Identify, groom, and prioritize automation and SRE use cases with stakeholders across the tower.

Build and manage automation pipelines and reusable Ansible Playbooks / Python modules to reduce toil and improve operational efficiency.

Define, measure, and improve reliability objectives (SLI/SLO/SLA) and error budgets for the tower.

Drive incident, problem, and change management participate in bridge / RCA calls and lead root-cause analysis.

Perform data analysis on incidents and manual effort, converting recurring patterns into automation opportunities.

Develop workflow diagrams, process documentation, runbooks, and obtain stakeholder approvals.

Collaborate with cross-functional teams to drive automation adoption and standardization.

Support implementation using Python, Ansible, and DevOps / CI-CD practices.

Track benefits realization (FTE savings, productivity gains) and report to leadership on a weekly / fortnightly basis.

Ensure adherence to operational standards, governance, security, and reliability objectives.

Required Skills:

Automation: Strong hands-on Python scripting and Ansible automation for provisioning, configuration, remediation, and orchestration.

SRE concepts: Reliability engineering, toil reduction, SLI/SLO/error budgets, observability, and capacity / performance management.

DevOps & CI/CD: Working knowledge of Git / GitHub / GitLab, CI-CD pipelines (Jenkins / GitLab CI), and version control practices.

ITSM: Change, Request, and Incident Management using ServiceNow delivering within SLA under tight timelines.

Analytical: Strong troubleshooting, data-analysis, and problem-solving skills.

Communication: Strong written and verbal communication able to create presentations and process / workflow documentation.

GenAI & Agentic AI:

Exposure GenAI-assisted operations and GenWizard / GenAI solution development.

Understanding Agentic AI concepts, prompt engineering, and AI-powered workflow automation.

1.5 Experience, Education & Professional Attributes

Experience: 8 – 10 years of relevant tower experience, including hands-on automation.

Education: Bachelor's degree in IT, Computer Science, or a related discipline.

Foundational certification: ITIL v3 / v4 Foundation certified (preferred).

High technical aptitude and commitment to continuous skill development.

Ability to operate independently, take initiative, and make decisions with minimal supervision.

Strong team player, adaptable to cross-skilling and flexible to on-call / 24x7 support.

Shift Details:

  • UK shift (B shift)
  • RTO as per project guidelines

Additional Information:

  • The candidate should have minimum 7.5 years of experience in Site Reliability Engineering.
  • This position is based at our Pune office.
  • A 15 years full time education is required.


More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152476983

Similar Jobs

Pune, India

Skills:

ServicenowPrometheusDatadogTerraformAutomation FrameworksSprint PlanningApisAgile MethodologiesJiraVersion ControlAnsiblePython developmentRest ApisKubernetesCRDsGenAIPrompt engineeringToken-routing logicCI CDWorkflow orchestrationIntegration developmentAgentic AI workflowsBacklog managementDevOps practicesAutomation lifecycle managementAgentic AI platformAgentic Frameworkrbac

Pune, India

Skills:

Vmware EsxiPythonBashRed Hat Enterprise LinuxAnsibleKvmMicrosoft Windows ServerDell server hardwareRed Hat SatelliteActive Directory

Pune

Skills:

YamlpythonAws CloudIncident ManagementDns ManagementBashteam collaboration

Pune

Skills:

YAML.BashPythonAws

Pune, India

Skills:

google apps script JavascriptBi ToolsCSSHTMLPythonAi

Beware of Scammers

We don’t charge money for job offers