Search Jobs

Search by job, company or skills

SRE / Production Engineering

SRE / Production Engineering

Infosys
  • Posted 11 hours ago
  • Be among the first 10 applicants

Job Description

SRE, Production engineering,Kubernetes, Terraform, CI/CD pipelines, Observability (Prometheus/Grafana)

Key Responsibilities: Reliability & Production Ownership

  • Lead production engineering practices to ensure high availability, scalability, and performance across services and platforms.
  • Define and drive SLOs/SLIs, error budgets, capacity planning, and reliability roadmaps aligned to business priorities.
  • Partner with engineering teams to design resilient architectures and reduce operational risk through proactive improvements. Incident Management & Operational Excellence
  • Own incident response processes (on-call readiness, triage, escalation, communication) and lead major incident bridges when needed.
  • Drive blameless postmortems, root-cause analysis, and corrective/preventive actions to prevent recurrence.
  • Establish operational runbooks, playbooks, and production readiness reviews for new releases and changes. Cloud Operations & Automation
  • Lead cloud operations to ensure secure, cost-effective, and reliable environments across regions/accounts/subscriptions.
  • Identify toil and implement automation to improve deployment safety, recovery time, and operational efficiency.
  • Standardize operational tooling and workflows to improve service health, change success rate, and MTTR. Leadership & Collaboration
  • Mentor engineers and influence cross-functional teams to adopt reliability engineering best practices.
  • Provide technical leadership in prioritization, execution planning, and stakeholder communication for reliability initiatives. Minimum Qualifications:
  • BTECH, MTECH, MCA, or MSC in Computer Science, IT, or a related field (or equivalent practical experience).
  • 12–14 years of experience in SRE, Production Engineering, or Reliability Engineering roles supporting large-scale systems.
  • Strong hands-on experience in cloud operations, incident management, and production support for critical services.
  • Proven ability to drive automation initiatives that reduce manual effort and improve system reliability.
  • Demonstrated experience leading operational processes such as on-call, postmortems, and production readiness practices. Preferred Qualifications:
  • Experience designing and implementing SLO/SLI frameworks, error budgets, and reliability KPIs across multiple teams.
  • Strong background in observability practices (monitoring, alerting, logging, tracing) and building actionable operational dashboards.
  • Expertise in release/change management practices that improve deployment safety and reduce production incidents.
  • Experience leading cross-team reliability programs, influencing stakeholders, and driving measurable improvements in uptime and MTTR.
  • Track record of mentoring engineers and setting engineering standards for operational excellence and automation at scale.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

About Company