Job Summary
We are seeking an experienced Linux SME to provide expert-level support, governance, and optimization of enterprise Linux environments within a Managed Services framework. The role involves ensuring SLA-driven operations, high availability, security, and compliance while driving automation and continual service improvement.
Key Responsibilities
- Act as the Linux Subject Matter Expert within the Managed Services team for escalations, operations, and technical advisory.
- Provide L3/L4 support for complex Linux-related incidents, problems, and changes across multiple distributions (RHEL, CentOS, Ubuntu, SUSE).
- Manage day-to-day Linux operations including monitoring, patching, OS upgrades, performance tuning, and troubleshooting.
- Ensure security hardening, compliance (CIS, PCI, SOX, ISO), and adherence to managed services baselines.
- Automate recurring administration tasks (patching, user management, monitoring) using scripting (Bash, Python) and tools (Ansible, Puppet, Chef).
- Support backup, DR, and high availability strategies for critical Linux workloads.
- Collaborate with AWS, Windows, Network, and Security SMEs to deliver integrated managed services.
- Conduct root cause analysis (RCA) for high-severity issues and present service improvement recommendations.
- Contribute to Service Improvement Plans (SIPs) and Continuous Service Improvement (CSI) initiatives.
- Provide documentation, knowledge base articles, and operational runbooks for L1/L2 teams.
- Mentor and train junior engineers to strengthen Linux operational capability in the Managed Services team.
Required Skills & Qualifications
- 8–12 years of IT experience with at least 5+ years in Linux administration in enterprise environments.
- Deep expertise in Linux OS (RHEL, CentOS, Ubuntu, SUSE) and system internals.
- Strong experience in day-to-day managed services operations (incident, problem, change, patch management).
- Hands-on automation experience using Ansible, Puppet, Chef, or similar.
- Proficiency in scripting (Bash, Python, Perl).
- Strong troubleshooting skills in multi-tier, distributed, and hybrid environments.
- Experience with monitoring and observability (Nagios, Zabbix, Prometheus, Grafana, ELK, Splunk).
- Knowledge of virtualization (VMware, KVM) and Linux on cloud platforms (AWS, Azure, GCP).
- Familiarity with ITIL processes and managed services delivery.
- Excellent communication and customer-handling skills for escalations and service reviews.
Preferred Qualifications
- RHCE / RHCSA / LFCS certifications.
- Knowledge of containers (Docker, Podman) and orchestration (Kubernetes, OpenShift).
- Experience in hybrid cloud Linux operations and multi-tenant managed environments.
- Exposure to security/compliance frameworks (PCI, SOX, ISO 27001).
- Experience with service governance, audits, and RCA preparation for customers