Search Jobs

Search by job, company or skills

Senior Principal Infrastructure Engineer

Senior Principal Infrastructure Engineer

Mphasis
  • Posted 2 hours ago
  • Be among the first 10 applicants

Job Description

Job Summary:

We are seeking a highly skilled Senior Principal Infrastructure Engineer to administer and optimize our ClearML Server and associated infrastructure. The ideal candidate will have a strong background in MLOps platforms, HPC execution environments, and containerized solutions, with a focus on supporting GPU-based AI/ML workloads. This role requires a proactive approach to managing resources, ensuring security, and automating processes to enhance operational efficiency.

Responsibilities:

  • Administer ClearML Server, including management of agents, execution queues, projects, users, roles, experiment tracking, pipelines, datasets, artifacts, and model registry.
  • Configure ClearML Agents on CPU and GPU worker nodes, integrating with HPC execution platforms such as Slurm, PBS Professional, or Kubernetes.
  • Support GPU-based AI/ML workloads utilizing NVIDIA drivers, CUDA, NCCL, UCX, and containerized environments.
  • Maintain secure container execution using technologies like Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes.
  • Implement confidential-computing controls leveraging AMD SEV-SNP, Intel TDX, NVIDIA Confidential Computing, Secure Boot, TPM, and remote attestation.
  • Integrate authentication, RBAC, TLS certificates, secrets management, and audit controls for ClearML and confidential workloads.
  • Monitor ClearML services, agents, queues, GPU utilization, task failures, scheduler integration, and overall platform health.
  • Automate deployment, configuration, monitoring, and troubleshooting processes using Python, Bash, Ansible, and Git.

Mandatory Skills:

  • Strong Linux administration skills, particularly with RHEL, SLES, Rocky Linux, or Ubuntu.
  • Hands-on experience with ClearML administration or a comparable MLOps platform.
  • Familiarity with HPC schedulers such as Slurm, PBS Professional, or LSF.
  • Knowledge of GPU platforms, including NVIDIA drivers, CUDA, and distributed training basics.
  • Proficiency in container technologies: Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes.
  • Strong scripting and automation skills in Python, Bash, Ansible, and Git.
  • Understanding of security fundamentals, including RBAC, IAM, TLS, certificates, secrets management, secure boot, and audit logging.
  • Familiarity with confidential-computing concepts such as TEE, encrypted memory, TPM, remote attestation, AMD SEV-SNP, Intel TDX, or equivalent technologies.

Preferred Skills:

  • Experience with the Nvidia Nemo stack on Cloud.
  • Knowledge of advanced monitoring and logging tools for infrastructure management.
  • Familiarity with cloud platforms and services related to AI/ML workloads.

Qualifications:

A degree in Computer Science, Information Technology, Engineering, or a related field is preferred. Relevant certifications in cloud computing, MLOps, or infrastructure management will be considered an advantage.

About Mphasis:

Mphasis applies to next-generation technology to help enterprises transform businesses globally. Customer centricity is foundational to Mphasis and is reflected in the Mphasis Front2Back™ Transformation approach. Front2Back™ uses the exponential power of cloud and cognitive to provide hyper-personalized (C=X2C2TM=1) digital experience to clients and their end customers. Mphasis Service Transformation approach helps shrink the core through the application of digital technologies across legacy environments within an enterprise, enabling businesses to stay ahead in a changing world. Mphasis core reference architectures and tools, speed and innovation with domain expertise and specialization are key to building strong relationships with marquee clients.

Key Skills

Monitoring and logging tools

Cloud platforms and services related to AI/ML workloads

Nvidia Nemo stack

About Company