Search by job, company or skills

  • Posted a month ago
  • Over 50 applicants have applied

Job Description

Role Overview

Provide second-level operational support for a containerized OpenStack Private Cloud deployed using Kolla-Ansible. Focus on advanced troubleshooting, service restoration, root cause analysis, platform validation, and coordination with engineering and infrastructure teams to ensure high availability and operational stability.

Key Responsibilities

• Provide second-level support for incidents and service requests escalated by the L1 Operations

team.

• Troubleshoot OpenStack services including Nova, Neutron, Keystone, Glance, Cinder, Horizon,

Placement, Heat, and Octavia.

• Diagnose and resolve issues related to compute, networking, storage, authentication, scheduling,

and API services.

• Troubleshoot Docker containers deployed through Kolla-Ansible and perform service recovery

following operational procedures.

• Perform advanced Linux administration, including troubleshooting CPU, memory, disk, filesystem,

networking, processes, and service failures.

• Analyze applications, containers, and system logs to identify root causes and restore services.

• Validate VM lifecycle operations, image management, volume operations, floating IP connectivity,

and OpenStack API availability.

• Monitor platform health using Grafana, Prometheus, and Alert manager, and investigate

infrastructure alerts.

• Validate the health of controllers, compute, storage, and network nodes after maintenance

activities.

• Perform troubleshooting of Cinder and Ceph storage services, validate volume creation,

attachment, snapshot operations, storage health, and collect diagnostic information for incident

analysis.

• Monitor storage capacity and backend storage health, coordinate with Infrastructure teams for

storage-related incidents, and validate storage services after maintenance activities.

• Perform hardware health validation using server management interfaces such as iDRAC, iLO, or

IPMI, and correlate hardware events with platform issues.

• Collect hardware diagnostic information and coordinate with Infrastructure, Data Center, and

OEM teams for hardware-related incidents.

• Support planned maintenance activities, platform validation, and change implementations.

• Prepare Root Cause Analysis (RCA) reports and update operational documentation and

knowledge base articles.

• Mentor L1 engineers and provide technical guidance during incident resolution.

Technical Skills

• Strong Linux administration (RHEL/Rocky Linux/Ubuntu)

• Good understanding of OpenStack architecture and core services

• Hands-on experience with Docker container operations

• Working knowledge of Kolla-Ansible-based OpenStack environments

• Experience with Grafana, Prometheus, and Alert manager

• Strong understanding of TCP/IP networking, VLANs, VXLAN, DNS, and SSH

• Experience with OpenStack CLI and basic API troubleshooting

• Knowledge of virtualization technologies (KVM/QEMU/libvirt)

• Basic understanding of HA Proxy, MariaDB/Galera, RabbitMQ, and Ceph

• Experience with ITSM tools such as JIRA or ServiceNow

• Basic shell scripting (Bash/Python) is an added advantage

Soft Skills

• Strong analytical and troubleshooting skills

• Excellent verbal and written communication

• Customer-focused approach

• Ability to work independently in a 24×7 rotational shift environment

• Good documentation and reporting practices

• Strong ownership and accountability

• Team collaboration and mentoring skills

• Effective incident management and escalation discipline

Preferred Certifications

• RHCE (Red Hat Certified Engineer)

• RHCSA (Red Hat Certified System Administrator)

• LFCS (Linux Foundation Certified System Administrator)

• OpenStack Foundation Certification (Preferred)

• Docker Certified Associate (Preferred)

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151459009

Beware of Scammers

We don’t charge money for job offers