Search by job, company or skills

Sr. HPC ENGINEER

Sr. HPC ENGINEER

Cognizant Consulting
Fresher
Not Disclosed
Early Applicant
  • Posted a month ago
  • Be among the first 30 applicants

Job Description

Role Overview

We are seeking seasoned professionals with deep expertise in operating and managing High-Performance Computing (HPC) platforms. The ideal candidate will have hands-on experience in designing, deploying, and maintaining HPC clusters, storage systems, and networking infrastructure, leveraging industry-leading tools and technologies.

Key Responsibilities

  • HPC Infrastructure Management
  • Operate and maintain HPC clusters based on CentOS, RHEL, and hardware platforms like HPE and NVIDIA DGX.
  • Ensure optimal performance, scalability, and reliability of compute resources.
  • Storage Administration
  • Manage large-scale storage systems including Dell Isilon, VAST Storage, Lustre, and GPFS.
  • Implement data lifecycle management and optimize storage performance for HPC workloads.
  • Networking
  • Configure and maintain InfiniBand-based networking for low-latency, high-bandwidth communication.
  • Troubleshoot network performance issues and ensure secure connectivity.
  • Cluster and Job Scheduling
  • Administer cluster management tools such as Bright Cluster Manager, Altair Grid Manager, and IBM LSF.
  • Optimize job scheduling and resource allocation for diverse workloads.
  • Monitoring and Automation
  • Implement monitoring solutions using Zabbix, Grafana, and ELK Stack.
  • Automate provisioning and configuration using Cobbler, Chef, Ansible, and AWS ParallelCluster.
  • Performance Tuning & Troubleshooting
  • Conduct performance benchmarking and tuning for HPC workloads.
  • Diagnose and resolve hardware/software issues across compute, storage, and network layers.
  • Security & Compliance
  • Ensure HPC environment adheres to security best practices and compliance standards.

Required Skills & Qualifications

  • Technical Expertise
  • Strong knowledge of Linux OS (CentOS, RHEL) and HPC hardware platforms (HPE, NVIDIA DGX).
  • Hands-on experience with parallel file systems (Lustre, GPFS) and enterprise storage solutions.
  • Proficiency in InfiniBand networking and high-speed interconnects.
  • Familiarity with job schedulers and cluster management tools (IBM LSF, Bright Cluster Manager, Altair Grid Manager).
  • Automation & Scripting
  • Expertise in Ansible, Chef, Cobbler, and scripting languages (Bash, Python).
  • Experience with AWS ParallelCluster or similar cloud-based HPC solutions.
  • Monitoring & Logging
  • Practical experience with Zabbix, Grafana, and ELK Stack for system health and performance monitoring.
  • Soft Skills
  • Strong problem-solving and analytical skills.
  • Ability to work in a fast-paced environment and lead technical teams.
  • Excellent communication and documentation skills.

Preferred Qualifications

  • Exposure to AI/ML workloads on HPC clusters.
  • Experience with containerization (Docker, Singularity) in HPC environments.
  • Knowledge of security hardening for HPC systems.

Education

  • Bachelor's or Master's degree in Computer Science, Engineering, or related field.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

NVIDIA DGX

AWS ParallelCluster

Bright Cluster Manager

IBM LSF

VAST Storage

Dell Isilon

Altair Grid Manager

HPC Infrastructure Management

HPE