Search by job, company or skills

Chief manager - application

Chief manager - application

HDB Financial Services
5-8 Years
Not Disclosed
  • Posted a month ago
  • Be among the first 10 applicants

Job Description

The Senior MLOps Engineer will be responsible for designing, implementing, and managing the end-to-end operational lifecycle of Machine Learning and Generative AI models. This includes model packaging, deployment, CI/CD automation, model monitoring, infrastructure management, governance, observability, and continuous optimization.

The role requires strong expertise in cloud infrastructure, DevOps, Kubernetes, ML platforms, model lifecycle management, and enterprise-scale AI deployment.

The ideal candidate will work closely with AI Engineers, Data Scientists, ML Engineers, AI Platform Engineers, Enterprise Architecture, Infrastructure, Information Security, and Digital Engineering teams to operationalize AI across the enterprise.

## MLOps Platform Strategy

  • Design and implement the enterprise MLOps platform.
  • Define the operating model for AI and ML deployments.
  • Establish standards for model lifecycle management.
  • Develop reusable deployment templates and automation frameworks.
  • Build scalable AI deployment capabilities across cloud and on-premise environments.
  • Define engineering best practices for AI operations.

## Machine Learning Lifecycle Management

Manage The Complete Lifecycle Of AI/ML Models Including

  • Model Registration
  • Version Control
  • Model Packaging
  • Model Validation
  • Model Deployment
  • Model Promotion
  • Model Rollback
  • Model Retirement
  • Model Archiving

Implement automated governance throughout the ML lifecycle.

## CI/CD for AI

Develop Automated CI/CD Pipelines For

  • Machine Learning Models
  • Generative AI Applications
  • AI APIs
  • Feature Engineering Pipelines
  • Data Validation Pipelines
  • Model Testing
  • Prompt Evaluation
  • AI Agent Deployments

Integrate deployment pipelines with enterprise DevOps processes.

## Production Deployment

Deploy AI Solutions Across Enterprise Platforms Including

  • Loan Origination Systems
  • CRM
  • Mobile Applications
  • Customer Portals
  • Contact Centre Platforms
  • Collections Platforms
  • Enterprise Data Lake
  • Marketing Automation Platforms
  • Digital Payment Platforms
  • API Gateway

Ensure high availability, resilience, scalability, and fault tolerance.

## Model Monitoring & Observability

Build Monitoring Capabilities For

  • Model Accuracy
  • Data Drift
  • Concept Drift
  • Prompt Performance
  • Latency
  • API Response Time
  • Infrastructure Utilization
  • GPU Utilization
  • Business KPIs
  • Cost Optimization

Develop dashboards and automated alerts for production AI systems.

## Infrastructure Automation

Design And Manage AI Infrastructure Using

  • Infrastructure as Code (IaC)
  • Containerization
  • Kubernetes Orchestration
  • Auto Scaling
  • GPU Resource Management
  • High Availability Architecture
  • Disaster Recovery
  • Backup & Restore
  • Capacity Planning

## Feature Store & Model Registry

Develop And Maintain Enterprise

  • Feature Store
  • Model Registry
  • Experiment Tracking
  • Artifact Repository
  • Metadata Repository

Ensure consistency and reusability across AI initiatives.

## Generative AI Operations

Operationalize Enterprise GenAI Applications Including

  • Large Language Models (LLMs)
  • Retrieval-Augmented Generation (RAG)
  • AI Agents
  • Prompt Libraries
  • Vector Databases
  • Knowledge Bases
  • AI Search
  • Enterprise AI Copilots

Manage prompt versioning, model evaluation, and deployment pipelines.

## Security & Compliance

Implement Security Best Practices For AI Deployments Including

  • Role-Based Access Control (RBAC)
  • Secrets Management
  • API Security
  • Identity & Access Management
  • Encryption
  • Vulnerability Management
  • Audit Logging
  • Secure Model Deployment
  • Compliance with enterprise security standards

Collaborate with Information Security and AI Governance teams.

# Educational Qualifications

  • Bachelor's Degree in Computer Science, Information Technology, Engineering, Artificial Intelligence, Data Science, or related discipline.

# Experience

  • 5-8 years of experience in DevOps, Cloud Engineering, MLOps, or AI Platform Engineering.
  • Minimum 3 years of experience operationalizing machine learning models in production.
  • Experience managing cloud-native AI platforms.
  • Experience in Banking, NBFC, Financial Services, or FinTech is preferred.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

ML Platforms

Model Lifecycle Management

GPU Resource Management

Enterprise-scale AI Deployment

Audit Logging

Security Best Practices

Secrets Management

Observability

Similar Jobs

6-10 yrs
Mumbai, India
Skills:
Cloud Services, AI gateways, AI infrastructure, GPU environments, vector databases, integration with enterprise platforms, model serving, developer tooling