Search by job, company or skills

Senior ML Engineer - AI Labs

Early Applicant
  • Posted 2 days ago
  • Be among the first 10 applicants

Job Description

Job Requirements

About the Role

As a Senior Machine Learning Engineer within the Data & Analytics team, you will be responsible for managing and optimizing the training infrastructure for Large Language Models (LLMs). This role demands a deep understanding of GPU architecture, machine learning principles, and distributed computing. You will lead Gen AI initiatives in a cross-functional setup, ensuring efficient resource utilization and timely delivery of large-scale ML projects.

Key Responsibilities

Primary Responsibilities

  • Lead Generative AI projects in a cross-functional team environment.
  • Apply advanced machine learning principles and algorithms, particularly for LLMs such as GPT-4, BERT, and Transformers.
  • Utilize deep learning frameworks like TensorFlow, PyTorch, and Keras for model training.
  • Maximize GPU utilization and efficiency through deep knowledge of computer architecture.
  • Manage and optimize cloud-based resources (AWS, Azure, GCP) for deep learning model training.
  • Implement containerization and orchestration using Docker and Kubernetes.
  • Apply parallel and distributed computing principles for scalable model training.
  • Integrate big data technologies like Hadoop and Spark into ML workflows.
  • Adopt MLOps practices and tools to manage the end-to-end ML lifecycle.

Secondary Responsibilities

  • Manage infrastructure for multiple ML projects, especially those involving deep learning models.
  • Optimize performance and resource allocation for large-scale ML tasks.
  • Handle GPU resource management both on-premises and in the cloud.
  • Address challenges in training large models, including memory management, data loading optimization, and hardware troubleshooting.
  • Collaborate closely with data scientists and ML engineers to understand infrastructure needs and deliver efficient solutions.

What We Are Looking For

Education

  • Graduation in BSC or BCA or B.Tech.

Experience

  • 4+ years of relevant experience in managing infrastructure for training large-scale ML models.
  • Hands-on experience with LLMs and deep learning frameworks.
  • Experience in cloud computing, containerization, and distributed systems.
  • Prior involvement in Gen AI projects and cross-functional team collaboration.

Skills and Attributes

  • Strong understanding of GPU architecture and optimization techniques.
  • Proficiency in TensorFlow, PyTorch, Keras, Docker, Kubernetes, and cloud platforms.
  • Knowledge of distributed computing frameworks like Hadoop and Spark.
  • Familiarity with MLOps tools and practices.
  • Excellent problem-solving and troubleshooting skills.
  • Ability to lead technical aspects of projects and ensure error-free, timely deliverables.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 153393399

Beware of Scammers

We don’t charge money for job offers