Search by job, company or skills

Senior Site Reliability Engineer

Senior Site Reliability Engineer

Infosys Limited
Fresher
Not Disclosed
Early Applicant
  • Posted a month ago
  • Be among the first 10 applicants

Job Description

Job Description:

  • Reliability Engineering
  • Design build and maintain highly available and fault tolerant production systems
  • Define and monitor SLIs SLOs and SLAs for critical services
  • Drive reliability improvements through automation and proactive engineering
  • Conduct capacity planning and performance optimization activities
  • Production Support Operations
  • Manage production environments and ensure service uptime
  • Lead incident response troubleshooting and root cause analysis RCA
  • Develop runbooks operational playbooks and disaster recovery procedures
  • Participate in on call rotations and major incident management processes

Key Responsibilities:

  • Cloud Infrastructure
  • Deploy and manage cloud native infrastructure across AWS Azure or GCP
  • Automate infrastructure provisioning using Infrastructure as Code IaC
  • Implement scalable and secure infrastructure solutions
  • Support Kubernetes based platforms and containerized workloads

Technical Requirements:

  • Observability Monitoring
  • Build monitoring logging tracing and alerting solutions
  • Implement observability frameworks using industry standard tools
  • Monitor application health performance metrics and infrastructure utilization
  • Drive continuous improvements in platform visibility and diagnostics
  • Automation DevOps
  • Automate deployments infrastructure management and operational workflows
  • Improve CI CD pipelines and release processes
  • Implement self healing auto scaling and operational automation solutions
  • Promote DevOps and SRE best practices across engineering teams

Additional Responsibilities:

  • Security Compliance
  • Ensure production environments meet security and compliance requirements
  • Manage secrets access controls and vulnerability remediation
  • Partner with security teams to implement security best practices
  • Leadership Mentoring
  • Mentor junior SREs DevOps Engineers and Production Support Engineers
  • Lead incident reviews and reliability improvement initiatives
  • Conduct knowledge sharing sessions and technical workshops
  • Drive operational excellence and engineering best practices

Preferred Skills:

Technology->DevOps->Site Reliability Engineering(SRE)

More Info

Job Type:
Industry:
Employment Type:

About Company

Similar Jobs

Bengaluru, India
Skills:
snowflake , Java, CSS, PostgreSQL, Kafka, Dynatrace, Microservices, HTML, Angular, Javascript, Datadog, Etl
Bengaluru, India
Skills:
Cassandra, PostgreSQL, Redis, Gcp, Terraform, MySQL, Helm, Python, Kubernetes, AWS, Golang, GitOps, OpenSearch, Observability
Bengaluru, India
Skills:
Gcp, Azure, Kubernetes, Logging, AWS, Alerting, tracing, Monitoring
Bengaluru, India
Skills:
PowerShell, Prometheus, Bash, Grafana, Datadog, New Relic, Jenkins, Terraform, Splunk, Python, Azure DevOps, GitHub Actions
Bengaluru, India
Skills:
containerization , Cloud security, Software Development Lifecycle, Configuration management, Incident Management, Docker, Azure, Kubernetes, AWS, Networking, Prometheus, Grafana, Automation, Datadog, Continuous Delivery, New Relic, Storage, Iam, Distributed System Design, Service orchestration, GCP services, GKE, Secrets management, AKS, EKS, Stackdriver, Compute