Search by job, company or skills

Senior Data Engineer

2-4 Years
Early Applicant
  • Posted a day ago
  • Be among the first 10 applicants

Job Description

Location Name: Pune Corporate Office - Mantri

Job Purpose

The MLOps Engineer will own the end-to-end operationalisation of machine learning, large language model (LLM), and agentic AI workloads on the Bajaj Finance Enterprise Data Platform — a 5PB+ medallion lakehouse built on Azure Databricks and Unity Catalog. This role sits at the intersection of data engineering, model lifecycle management, and AI governance, ensuring that every model — from classical ML to RAG pipelines and autonomous agents — is reproducible, explainable, observable, and production-safe.

The incumbent will architect and implement the MLOps and LLMOps platform on Databricks, leveraging Agentbricks (Databricks Agent Framework), Databricks Apps, MLflow, Feature Store, Model Serving, and Mosaic AI — embedding rigorous CI/CD, drift monitoring, cost governance, and responsible-AI guardrails across the full lifecycle. This is a high-impact, high-visibility role critical to delivering Bajaj Finance's AI-first data strategy at scale across 120M+ customer interactions.

Duties And Responsibilities

  • MLOps Platform Engineering
  •  Design, build, and maintain the end-to-end MLOps platform on Azure Databricks — covering experiment tracking (MLflow), model registry, Feature Store, batch and real-time model serving, and automated retraining pipelines.
  •  Implement CI/CD pipelines for ML code (Databricks Asset Bundles / DABs, Azure DevOps, GitHub Actions) ensuring reproducible model builds, automated testing, and zero-downtime deployments.
  •  Govern the full model lifecycle: versioning, lineage tracking via Unity Catalog, promotion workflows (Dev Staging Production), and model archival with audit trails.
  •  Establish and maintain Feature Store — curated, reusable feature sets across credit risk, fraud, customer propensity, and collections models — ensuring data freshness, SLA adherence, and lineage traceability.
  •  Operationalise Databricks Model Serving (serverless + provisioned endpoints) and Mosaic AI for scalable, low-latency inference across batch and online serving patterns.
  • LLMOps — Large Language Model Lifecycle
  •  Design and implement LLMOps pipelines for RAG-based applications on Databricks: document ingestion chunking embedding generation vector indexing (Mosaic AI Vector Search / Unity Catalog Volumes) retrieval LLM serving.
  •  Implement prompt versioning, prompt evaluation frameworks (MLflow LLM Evaluate, Mosaic AI Evaluation), and automated hallucination / faithfulness / relevance scoring using LLM-as-a-Judge patterns.
  •  Manage LLM fine-tuning workflows: curate supervised fine-tuning datasets, run PEFT/LoRA jobs on Databricks GPU clusters, register and serve fine-tuned models via MLflow Model Registry.
  •  Build token-cost monitoring, latency tracking, and model quality dashboards; implement automated rollback triggers when LLM quality KPIs degrade beyond defined thresholds.
  •  Enforce LLM governance: input/output guardrails, PII redaction, jailbreak detection, and compliance logging aligned with RBI and DPDP Act requirements.
  • Agentbricks & Agentic AI Operationalisation
  •  Deploy and operationalise autonomous AI agents using Databricks Agentbricks (Agent Framework) — including tool-calling agents, multi-agent orchestration, and human-in-the-loop review gates.
  •  Implement agent observability: trace logging (MLflow Traces), latency profiling, tool-call auditing, and failure mode analysis for production agents such as FinOps Sentinel and Governed Analytics.
  •  Build agent evaluation harnesses — synthetic scenario libraries, adversarial test suites, and regression benchmarks — to validate agent behaviour before and after model updates.
  •  Manage agent state and memory persistence using Databricks-native storage (Delta Lake, Unity Catalog) and integrate with external databases (CosmosDB, Neo4j) as required by agent workflows.
  •  Collaborate with AI Engineers on agent architecture decisions and ensure all agentic workloads meet latency SLAs, cost budgets, and safety standards.|D. Databricks Apps & Self-Serve AI
  •  Develop and deploy internal AI-powered applications using Databricks Apps — enabling business users to interact with ML models, RAG systems, and analytics agents through governed, self-serve interfaces.
  •  Integrate Databricks Apps with Unity Catalog row/column-level security, ensuring data access controls are enforced transparently without requiring users to understand the underlying platform.
  •  Build reusable application templates and deployment blueprints for common use cases (credit decisioning dashboards, collections intelligence tools, KYC automation) to accelerate delivery across business units.
  • Monitoring, Observability & Governance
  •  Implement comprehensive model monitoring: data drift (population stability index, KS-statistic), concept drift, prediction drift, and feature distribution shifts — with automated alerts and retraining triggers via Databricks Workflows.
  •  Build model performance dashboards in Databricks SQL / Power BI tracking accuracy, F1, AUC, RMSE, and business KPIs (approval rate, delinquency lift) across all production models.
  •  Enforce Unity Catalog-based data and model lineage — every model must have traceable lineage from raw source data through features to predictions, satisfying RBI Model Risk Management guidelines and BCBS 239.
  •  Conduct regular model validation and bias audits; document model cards and maintain model risk registers in collaboration with Risk and Compliance teams.
  •  Implement cost governance: cluster auto-scaling policies, spot-instance strategies, DBU budget alerts, and compute right-sizing recommendations to optimise the platform spend within approved budgets.
  • Collaboration & Engineering Excellence
  •  Partner with Data Scientists, AI Engineers, Data Engineers, and Business stakeholders to productionise models rapidly without sacrificing quality or compliance.
  •  Establish and evangelise MLOps best practices, coding standards, and platform conventions through documentation, internal training sessions, and code reviews.
  •  Contribute to the EDIL technical roadmap — evaluating emerging Databricks capabilities (Delta Live Tables, Lakeflow, Genie Spaces, AI/BI Dashboards) and proposing adoption plans with clear ROI justification.

Key Decisions / Dimensions


  •  Selection of MLOps tooling and pipeline patterns within the approved Databricks platform stack.
  •  Model promotion from Staging to Production for models below defined risk thresholds after successful evaluation.
  •  Compute cluster configurations, auto-scaling policies, and spot-instance strategies for ML workloads.
  •  Drift alert thresholds and automated retraining triggers for registered models.
  •  Agent trace sampling rates, logging retention policies, and observability dashboard design.

Major Challenges


  •  Balancing velocity and rigour: delivering fast model deployments across 50+ source systems and 120M+ customer records while maintaining strict audit trails demanded by RBI Model Risk Management frameworks.
  •  LLM non-determinism in production: managing hallucination risk, prompt sensitivity, and output variability in customer-facing AI applications where errors have direct financial and regulatory consequences.
  •  Agentic AI safety: ensuring autonomous agents operating on live financial data remain within sanctioned boundaries — particularly for high-stakes decisions such as credit line adjustments, fraud flags, and collections prioritisation.
  •  Scale and latency: serving real-time inference (sub-100ms) while maintaining model quality and managing compute costs within budget.
  •  Cross-functional alignment: coordinating model deployment gates across Data Science, Risk, Compliance, IT Security, and Business teams — each with different timelines, priorities, and risk appetites.
  •  Keeping pace with the Databricks roadmap: the platform evolves rapidly (Agentbricks, Mosaic AI, Lakeflow); the incumbent must continuously evaluate and integrate new capabilities without destabilising production workloads.

Required Qualifications And Experience


  •  B.Tech / B.E. / M.Tech / M.S. in Computer Science, Information Technology, Data Science, Electrical Engineering, or a related quantitative discipline.

Graduates from IITs, NITs, BITS Pilani, or other Tier-1 institutions preferred; exceptional candidates from other institutions with demonstrable production AI/ML experience will be considered.

  • Work Experience: 2–4 years of total experience in data/AI engineering, with a minimum of 1 years of hands-on MLOps or LLMOps experience in a production environment.
  •  Demonstrated experience deploying and monitoring ML models in production at scale — not prototypes or PoCs alone; evidence of model lifecycle ownership from training through retirement.
  •  Proven experience with Azure Databricks in an enterprise context — ideally within BFSI (banking, financial services, insurance), e-commerce, or a large-scale consumer data platform.
  •  Experience with LLM-based applications in production (RAG pipelines, LLM serving, prompt evaluation) is a strong differentiator.
  •  Exposure to agentic AI frameworks (LangGraph, Agentbricks, CrewAI) in production or advanced PoC settings is highly valued.
  •  Track record of building scalable, observable, and cost-efficient ML infrastructure — demonstrated through measurable outcomes (latency, accuracy, cost, reliability improvements).
  •  Experience working in regulated industries (BFSI preferred) with exposure to model validation, audit documentation, and compliance frameworks is an advantage.


More Info

Job Type:
Industry:
Employment Type:

Job ID: 151637849

Similar Jobs

Pune, India

Skills:

GithubPerformance TuningPysparkKafkaData ModelingAzure DatabricksSqlGitAzure Data FactoryDatabase MirroringPythonAzure DevOpsSpark Structured StreamingOptimizationDevOps pipelinesData Lake Storage

Pune, India

Skills:

snowflake Data ModellingPower BiCsvTableauJsonRedshiftSqlDatabase DesignELTAzure Data FactoryPostgresDatabricksPythonAws S3EtlParquetLambda functions

Pune, India

Skills:

S3RDSApache SparkEmrAWS CloudWatchNosqlLambdaTerraformQuicksightPostgresSparkData LakePythonAWSMWAAIcebergGithub actionsAirFlowGlueTime series data

Pune, India

Skills:

DevopsGitPower BiPysparkSqlAWSDatabricks Lakehouse PlatformDataOpsDelta LakeUnity Catalog

Pune, India

Skills:

semantic modeling SqlDatabricksELTAutomation ToolsEtlMySQLData QualityPythonAzureApisPostgreSQLMongoDBData ModelingDelta LakeSource ControlMetadata Curation