I
SRE / Production Engineering
I
SRE / Production Engineering
Infosys12-14 Years
- Posted 11 hours ago
- Be among the first 10 applicants
Job Description
SRE, Production engineering,Kubernetes, Terraform, CI/CD pipelines, Observability (Prometheus/Grafana)
Key Responsibilities: Reliability & Production Ownership
Key Responsibilities: Reliability & Production Ownership
- Lead production engineering practices to ensure high availability, scalability, and performance across services and platforms.
- Define and drive SLOs/SLIs, error budgets, capacity planning, and reliability roadmaps aligned to business priorities.
- Partner with engineering teams to design resilient architectures and reduce operational risk through proactive improvements. Incident Management & Operational Excellence
- Own incident response processes (on-call readiness, triage, escalation, communication) and lead major incident bridges when needed.
- Drive blameless postmortems, root-cause analysis, and corrective/preventive actions to prevent recurrence.
- Establish operational runbooks, playbooks, and production readiness reviews for new releases and changes. Cloud Operations & Automation
- Lead cloud operations to ensure secure, cost-effective, and reliable environments across regions/accounts/subscriptions.
- Identify toil and implement automation to improve deployment safety, recovery time, and operational efficiency.
- Standardize operational tooling and workflows to improve service health, change success rate, and MTTR. Leadership & Collaboration
- Mentor engineers and influence cross-functional teams to adopt reliability engineering best practices.
- Provide technical leadership in prioritization, execution planning, and stakeholder communication for reliability initiatives. Minimum Qualifications:
- BTECH, MTECH, MCA, or MSC in Computer Science, IT, or a related field (or equivalent practical experience).
- 12–14 years of experience in SRE, Production Engineering, or Reliability Engineering roles supporting large-scale systems.
- Strong hands-on experience in cloud operations, incident management, and production support for critical services.
- Proven ability to drive automation initiatives that reduce manual effort and improve system reliability.
- Demonstrated experience leading operational processes such as on-call, postmortems, and production readiness practices. Preferred Qualifications:
- Experience designing and implementing SLO/SLI frameworks, error budgets, and reliability KPIs across multiple teams.
- Strong background in observability practices (monitoring, alerting, logging, tracing) and building actionable operational dashboards.
- Expertise in release/change management practices that improve deployment safety and reduce production incidents.
- Experience leading cross-team reliability programs, influencing stakeholders, and driving measurable improvements in uptime and MTTR.
- Track record of mentoring engineers and setting engineering standards for operational excellence and automation at scale.
More Info
Key Skills
CI CD pipelines
Observability
