Search by job, company or skills

  • Posted 20 hours ago
  • Be among the first 10 applicants

Job Description

SRE Manager – Reliability Engineering, Automation & Operational Excellence

Role Overview

Step into a technically hands-on SRE leadership role where you will engineer reliability – not just manage it. Own architecture for monitoring and automation, review infrastructure-as-code with rigor, author operational designs, and personally dive into complex incidents.

Bring an engineering-first mindset: be the leader who sees beyond what is currently working, identifies risks before they become incidents, champions AIOps, and actively expands the team into adjacent platform engineering.

You will lead a 15+ member team supporting large-scale cloud platform reliability across multiple workstreams – driving operational excellence, cloud cost efficiency, and engineering innovation. This is a high-visibility role as the single point of accountability between the vendor team and client engineering leads.

Key Responsibilities

Architecture & Technical Leadership

  • Own reliability architecture: drive decisions for monitoring, alerting, automation, and infrastructure tooling; conduct code reviews on IaC, scripts, and configs
  • Personally engage in complex incident debugging using KQL, telemetry, and distributed traces to unblock the team and drive rapid root cause resolution
  • Author runbook designs, define operational architecture, standardize procedures, and continuously raise the reliability bar

Delivery & Execution

  • Drive delivery in Agile/Scrum against contractual commitments: sprint progress, capacity management, delivery forecasting
  • Manage on-call rotation with health tracking (burnout, incident volume, team well-being)
  • Drive incident management excellence: SLA targets, post-incident reviews, MTTD/MTTR tracking, corrective actions, systemic improvements
  • Drive deployment operations: staged rollouts, Blue/Green deployments, deployment validation gates, safe deployment practices across cluster-based infrastructure

Cloud Cost & Operational Efficiency

  • Own cloud cost management: Azure spend monitoring, optimization, efficiency practices, burn rate reporting against budget
  • Drive AI and automation: AIOps adoption, automated incident detection, intelligent alerting, auto-remediation, Copilot-assisted engineering

Innovation & Growth

  • Think beyond current operations: identify reliability risks before incidents, propose innovative automation, expand into adjacent platform engineering
  • Solution for new engagements: scoping, estimation, technical proposals, and delivery model shaping for new workstreams

People & Stakeholders

  • Drive multi-stakeholder coordination: primary technical interface with client leads, priority translation, platform roadmap awareness, proactive capacity positioning
  • Take full ownership: end-to-end accountability for team outcomes, proactive issue resolution, hold team to SLA and operational standards
  • Drive recruitment, professional development, and culture: technical interviews, team connects, mentor/coach across levels (intern to principal)


Problem Solvin

  • gDeep infrastructure debugging across distributed systems and cloud service degradatio
  • nSees beyond current operations to identify risks and propose innovative automatio
  • nSystems thinking for reliability architecture trade-off
  • sData-driven incident analysis for staffing, process, and tooling decision
  • sFinancial awareness for cloud spend management and team value demonstratio

n
Technic

  • alSRE/DevOps leadership, operational architecture, distributed systems (Service Fabric, Kubernetes/AK
  • S)Incident management, capacity planning, cloud cost optimizati
  • onAzure DevOps, Agile/Scrum, infrastructure-as-code (ARM, Bicep, Terrafor
  • m)Monitoring/observability (KQL, telemetry, distributed tracing), AIOps, automation tooli
  • ngOn-call rotation design and health monitori
  • ngDelivery metrics (MTTD, MTTR, deployment frequency, SLA complianc

e)
Required Skills & Qualificati

  • onsHands-on SRE or platform engineering management leading 15+ engineers across geograph
  • iesProven reliability architecture ownership, code review rigor, and complex infrastructure debugg
  • ingTrack record of proactive improvements, solutioning for new engagements, and delivery growth beyond sc
  • opeExperience managing contract-based delivery with SLA accountability and operational commitme
  • ntsStrong design reviews, incident coordination, executive reporting, and stakeholder communicat
  • ionOwnership mindset delivering operational excellence beyond expectati

ons
Preferred Qualificat

  • ionsImproving MTTD/MTTR through technical and process cha
  • ngesCloud cost optimization and large-scale SaaS operat
  • ionsAIOps adoption (auto-remediation, intelligent alerting) in SRE t
  • eamsOn-call health management and sustainable support mo
  • delsSolutioning and estimation for new reliability engineering engagem
  • entsOperational architecture documents, runbook designs, reliability stand
  • ardsService Fabric, Kubernetes, or large-scale cluster-based platform operat

ions
Why You'll Love This

  • RoleEngineer reliability at scale – own architecture decisions that directly impact platform uptime and perfor
  • manceChampion AIOps, automation, and innovation while leading a high-performing team across multiple workst
  • reamsShape operational excellence as the trusted engineering partner to client leadership with direct business i

mpact

More Info

Job Type:
Industry:
Employment Type:

Job ID: 152528673

Similar Jobs

Hyderabad, India

Skills:

API designMicroservicesScalabilityTypescriptDistributed SystemsApisReactjsNextJSBFFAWS cloud servicesClaudeAI coding assistantsdata layer conceptscomposable architecturesDevOps practicesRAG patternsCode CursorMessagingCQRSCAP trade-offsGitHub CopilotCachingMCP Model Context ProtocolCypressCI CD pipelinesmicro-frontends

Hyderabad, India

Skills:

snowflake JavaCassandraCloudformationSpring BootKafkaNew RelicJmeterDockerTerraformMySQLAnsibleReactjsKubernetesAws S3OverOps

Hyderabad, India

Skills:

data engineering Data ScienceCloud TechnologiesAzureAWSGenerative AIReal Time Data AnalyticsCloud FrameworksAiPrompt Flow Agent EngineeringML Model Development

Hyderabad, India

Skills:

JavaGolangUNIXGcpDockerLinuxNetworking ProtocolsNetwork ProgrammingAzurePythonKubernetesAWS

Hyderabad, India

Skills:

JavaOauth2Node.jsJwtHttpMicroservicesScalabilityRESTSwaggerOpenshiftRestful ApisKubernetesPythonmutual TLSGoreusabilitybackend programming languagesJSON Web TokensAPI-first developmentclean architectureCI CD pipelinesOpenAPIAPI security protocols

Beware of Scammers

We don’t charge money for job offers