Summary We are looking for an experienced Site Reliability Engineer (SRE) to join Client's Privileged Access Management (PAM) team. The role is focused on production support, incident management, platform reliability, monitoring, troubleshooting, and operational excellence for mission-critical applications. This opportunity is ideal for candidates with strong Linux, scripting, SQL troubleshooting, ITSM, and application support experience who enjoy working in a high-availability production environment.
Responsibilities
Monitor and support critical production applications and infrastructure.
Investigate and resolve incidents, service requests, and production issues.
Perform root cause analysis (RCA) and drive permanent resolutions.
Participate in on-call support and major incident management (P1/P2 incidents).
Troubleshoot application, middleware, and database issues.
Execute change requests (CRQs) following approved implementation plans and MOPs.
Support patching activities and certificate renewals.
Work with cross-functional teams to ensure service availability and reliability.
Track and coordinate issue resolution across dependent teams.
Maintain operational documentation, runbooks, and support procedures.
Mandatory Skills
5 to 7 years of relevant experience.
Strong Linux/Unix administration experience.
Hands-on Shell/Bash scripting or PowerShell scripting.
Experience troubleshooting Oracle or PostgreSQL databases.
Good understanding of SQL queries for debugging and issue analysis.
Strong ITSM knowledge:
Incident Management
Problem Management
Change Management
SLA/SLO concepts
Root Cause Analysis (RCA)
Experience supporting production environments.
Strong communication and stakeholder management skills.