

Search by job, company or skills

Qualification & Experience:
8+ years of working experience in IT working on Linux environments (RHEL) with development
experience. ● Strong Linux/RedHat system administration skills for Installing, managing Scientific
Computing applications and Performance Analysis and debugging real time application execution.
Storage Infrastructure & Data Lifecycle Management
● Design & Engineering: Architectural planning, deployment, and administration of parallel file
systems (e.g., Lustre, IBM Spectrum Scale, BeeGFS) and scalable object storage solutions.
● Performance Engineering: Continuous monitoring of storage I/O performance; identification and
mitigation of I/O bottlenecks to ensure throughput requirements for compute workloads are met.
● Data Lifecycle Strategy: Preparing for the future implementation of tiered storage concepts
(HSM/ILM) to enable automated data migration across different performance tiers (hot-to-cold
archiving).
Network Architecture & Troubleshooting
● Advanced Diagnostics: Analyzing complex interdependencies between storage backends, file
systems, and fabric architectures.
● Protocol Analysis: Utilizing deep packet inspection tools (e.g., tcpdump, Wireshark) and
performance profiling suites (e.g., fio, perf, iostat) to isolate packet loss, latency jitter, or metadata
lock contention.
● Interconnect Optimization: Performance tuning of low-latency network configurations
(InfiniBand/Ethernet) with a specific focus on RDMA protocols (RoCE/iWARP).
Access Control, Security & Compliance
● Identity & Access Governance: Managing NFS (v3/v4) and SMB/CIFS exports, taking into account
multi-tenant structures and cross-cluster access.
● Hardening & Compliance: Enforcing network-level export policies (CIDR-based) and implementing
POSIX-compliant ACLs as well as POSIX group mappings via LDAP/NIS.
● Data Security: Ensuring data-at-rest and data-in-transit encryption in compliance with regulatory
standards (GDPR, Export Control) and institutional security policies.
Workload & User Enablement
● Data Movement: Orchestrating large-scale data transfers using specialized tools (e.g., XCP,
Aspera, Globus).
● Consulting & Optimization: Advising users on I/O-efficient application development and best
practices for highly scalable workflows.
Candidate Profile (Skills)
● Expertise: Deep knowledge in administering HPC storage stacks and distributed file systems
(POSIX-compliant).
● Network Stack: Expert knowledge of TCP/IP networking, InfiniBand fabrics, routing & switching,
and RDMA technologies.
● Automation: Experience in automating provisioning and configuration management (Ansible,
Puppet, Bash/Python).
● Containers & Orchestration: Solid understanding of integrating storage solutions into
containerized environments (Singularity/Apptainer, Kubernetes).
● Good understanding of ITIL processes (Incident, Change, Problem, Service Management);
certification preferred.
● Analytical mindset with strong problem-solving skills, able to diagnose, reproduce, and resolve
complex issues.
● Experience in stakeholder management and user support, with ability to adapt to evolving
business processes.
Good understanding of server architecture and Engineering applications.
● Knowledge on AI tools (Copilot, Gemini Coding Assistant) for Quality improvement, Automation
and Value generation.
● Understanding and experience in Cloud technologies (preferably AWS) and contributing to building
cloud ready solutions. (Bonus points if Certified).
● Excellent communication skills to work in a globally distributed team
Responsibilities:
● Operate, maintain, and optimize Scientific Computing environments (Linux/RedHat, HPC
clusters, scheduling systems).
● Develop and maintain legacy codebases, new codebases and automation scripts.
● Troubleshoot and resolve issues related to Scientific Computing Applications, its installation,
dependency management and performance.
● Apply ITIL processes in daily operations (Incident/Change/Problem Management).
● Proactively challenge existing workflows and propose innovative improvements to enhance
efficiency and user experience.
● Collaborating with business users and relevant stakeholders to define project requirements,
scope and deliverables.
● Provide day-to-day support for production processing by solving incidents and requests raised by
the user community and an ongoing event and alert management
● Analyze and propose/implement to Improve incident resolution quality.
● Perform incident categorisation/classification & find out critical issues/most popular incidents and
do root cause analysis
● Perform root cause analysis for critical incidents & trend analysis for proactive
measures
● Participate in Agile Development activities and deliver features for fix/improvement of the
service Apply DevOps and AI tools, culture and mindset for all your activities on a daily
basis
● Excellent communication skills to work in a globally distributed team
● Contribute to continuous improvement initiatives on demand.
● Engage and motivate the people (involved and stakeholders / customers) to promote Self-
reflection, Self-management, Communication and Teamwork.
● Manage Conflicts and Negotiation to achieve optimized results for the business.
● Define and develop the objectives hierarchy of the project or product using appropriate strategies,
methods and tools.
● Engage stakeholders in a governance that ensures effective use of time, budget and achievement
of quality.
● Integrate processes and people from other functions (matrix organizations).
● Ensure effective integration of lessons learnt. Ensure active risk and opportunity and contract
(internal or external) management.
● Ensure effective management of partners and suppliers.
Job ID: 153753591
Skills:
Ansible, Puppet, Smb, Itil, Nis, Ethernet, Linux Redhat, Git, Iostat, Xcp, Ldap, AWS, Nfs, Kubernetes, Python, Bash, Debugging, Cloud Technologies, lustre, Globus, infiniband, RDMA, AI tools, HPC scheduling, Performance Analysis, IBM Spectrum Scale, Apptainer, Aspera, Singularity, BeeGFS, PERF, fio