Search Jobs

Search by job, company or skills

Site Reliability Engineer (SRE)

Site Reliability Engineer (SRE)

Bahwan Cybertek
Early Applicant
  • Posted 7 days ago
  • Be among the first 20 applicants

Job Description

Role Description

The Site Reliability Engineer (SRE) is responsible for improving the reliability, availability, performance, and operability of PAH-supported software systems. This role combines software engineering and IT operations to automate operational work, monitor system performance, and reduce toil. The SRE establishes and manages monitoring, ing, incident response, and problem management practices to ensure applications remain available and performant during updates and failures. The role partners with engineering, architecture, and product teams to define reliability standards and production readiness requirements. SRE is a practical implementation of DevOps focused on maintaining software quality in fast-paced development environments.

Define, implement, and maintain observability (monitoring, logging, tracing) and actionable ing aligned to service health.

Drive incident management: on-call readiness, triage, incident command support, communications, and post-incident reviews (RCA)

Reduce operational toil through automation (runbooks-to-automation, self-healing, deployment/rollback automation).

Establish reliability standards: SLOs/SLIs, error budgets, production readiness reviews, and release risk controls

Performance and reliability engineering: capacity planning, load/performance analysis, resilience testing, and failure-mode mitigation

Partner with engineering teams to improve operational hygiene (deployability, rollback strategy, configuration, secrets, dependency management)

More Info

Job Type:
Industry:
Employment Type:

Key Skills

load performance analysis

failure-mode mitigation

deployment rollback

resilience testing

observability

About Company