Search by job, company or skills

Data Engineer

8-10 Years

This job is no longer accepting applications

Job Description

Key Responsibilities

Databricks Platform Engineering

  • Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod environments.
  • Configure and manage Databricks Unity Catalog for data governance, access control, fine-grained permissions, and data lineage.
  • Optimize cluster configurations — instance types, auto-scaling policies, spot/preemptible nodes — for cost and performance.
  • Implement workspace-level best practices: folder structures, access controls, secret management (Databricks Secrets / Azure Key Vault / AWS Secrets Manager).
  • Manage Databricks jobs, workflows, and multi-task job orchestration with dependency management.

Delta Lake & Lakehouse Architecture

  • Design and implement Delta Lake tables with appropriate partitioning, Z-ordering, and file compaction (OPTIMIZE / VACUUM).
  • Build Medallion Architecture (Bronze / Silver / Gold) layers for structured data lake organization.
  • Implement Delta Live Tables (DLT) pipelines for declarative, reliable ETL/ELT with built-in data quality expectations.
  • Manage schema evolution, table versioning, time travel, and Change Data Feed (CDF) for incremental processing.
  • Design data lakehouse patterns integrating Delta Lake with external systems (Kafka, ADLS, S3, GCS).

Data Pipeline Development (PySpark / SQL)

  • Develop scalable batch and streaming data pipelines using PySpark, Spark SQL, and Delta Lake.
  • Build structured streaming pipelines for real-time ingestion from Kafka, Event Hubs, and Kinesis into Delta tables.
  • Write optimized PySpark transformations leveraging broadcast joins, adaptive query execution (AQE), and dynamic partition pruning.
  • Create reusable transformation libraries, utility frameworks, and pipeline templates for team productivity.
  • Implement robust error handling, retry logic, and dead-letter queue patterns in production pipelines.

MLflow & AI/ML Workloads

  • Set up and manage MLflow tracking servers, experiment registries, and model lifecycle management on Databricks.
  • Support data scientists and ML engineers in deploying model training and inference workloads on Databricks clusters and GPU instances.
  • Build feature engineering pipelines using Databricks Feature Store for reusable, versioned ML features.
  • Enable GenAI workloads — LLM fine-tuning, RAG pipeline development, and vector search (Databricks Vector Search / Mosaic AI).
  • Implement MLOps practices: model versioning, A/B testing, model serving via Databricks Model Serving endpoints.

Cloud Integration & DevOps

  • Integrate Databricks with cloud-native services: Azure Data Lake Storage (ADLS).
  • Build and maintain CI/CD pipelines for Databricks notebooks and jobs using Azure DevOps, GitHub Actions, or GitLab CI.
  • Implement Databricks Asset Bundles (DABs) or Terraform for infrastructure-as-code (IaC) deployment of Databricks resources.
  • Manage data ingestion using Auto Loader, COPY INTO, and partner integrations (Fivetran, dbt, Airbyte).
  • Monitor pipeline health, cluster utilization, and costs using Databricks system tables and cloud cost management tools.

Governance, Security & Optimization

  • Implement row-level security, column masking, and dynamic data views using Unity Catalog policies.
  • Ensure data quality enforcement using Delta Live Tables expectations and Great Expectations integrations.
  • Conduct performance tuning — query plan analysis, caching strategies, Photon engine enablement.
  • Maintain data cataloging, metadata management, and data lineage tracking within Unity Catalog.
  • Document architecture decisions, runbooks, and operational guides for Databricks workloads.

Required Qualifications

Education

  • Bachelor's or Master's degree in Computer Science, Information Technology, Data Engineering, or related field.

Experience

  • 8+ years of total experience in data engineering or software engineering.
  • 3+ years of dedicated hands-on experience with the Databricks platform in production environments.
  • Strong background in big data engineering, cloud data platforms, and distributed computing.

Databricks Platform

  • Deep expertise in Databricks Workspaces, Clusters, Jobs, Workflows, and Repos.
  • Proficiency with Unity Catalog — metastore setup, catalog/schema/table management, access controls, and data lineage.
  • Hands-on experience with Delta Live Tables (DLT) — pipeline development, expectations, and monitoring.
  • Strong command of Delta Lake internals — transaction log, ACID guarantees, file layout, and optimization techniques.
  • Experience with Databricks SQL Warehouses, SQL Analytics, and dashboard creation.
  • Knowledge of Databricks Photon engine, serverless compute, and cost optimization strategies.

PySpark & SQL

  • 4+ years of PySpark development — DataFrames, Datasets, Spark SQL, RDD operations.
  • Expert-level SQL — window functions, lateral joins, CTEs, recursive queries, and analytical functions.
  • Experience with Spark performance tuning — AQE, query plans (EXPLAIN), partitioning, and caching.
  • Proficiency with Python for pipeline development, utilities, and automation.

Cloud Platforms

  • Hands-on experience with at least one: Azure (ADLS Gen2, ADF, Azure Databricks), AWS (S3, EMR, Glue, AWS Databricks), or GCP (GCS, BigQuery, Dataproc).
  • Experience with cloud networking for Databricks: VNet/VPC injection, private endpoints, and firewall configurations.
  • Familiarity with IAM roles, managed identities, and service principal authentication for Databricks.

MLflow & ML Engineering (Nice to Have)

  • Working knowledge of MLflow — experiment tracking, model registry, and deployment.
  • Experience supporting ML pipelines on Databricks for training, evaluation, and serving.
  • Exposure to Databricks Feature Store and Mosaic AI / GenAI capabilities.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151810671

Similar Jobs

Pune, India

Skills:

Document Databases Working knowledge of MongoDBETL Tools Proficiency in Ab Initio GDE EME Co Operating SystemUnix Linux ScriptingSQL PL SQL QueryingTesting QA Exposure to SIT E2E and OAT testing cyclesMessaging Systems Experience with IBM MQ Apache KafkaAgile Delivery Experience with JIRAScheduling Tools Familiarity with Tivoli Workload Scheduler TWSSecurity Compliance Awareness of access control vulnerability assessments

Pune, India

Skills:

.NETdata warehouses Spark SQLPower BiAzure Log AnalyticsPowershell ScriptingHiveAzure Data FactoryData lakesMicrosoft Azure Data platformAzure SQL Data WarehouseAzure Storage ServicesAzure Application InsightsData BricksStream Analyticsdata martsEvent HubsAzure SQL DBAzure Analysis Services

Pune, India

Skills:

JavaUnixApache FlinkData ModelingSchema DesignPysparkData CleansingApache SparkData Warehouse ConceptsShell ScriptingSqlELTApache AirflowLinuxApache KafkaRestful ApisPythonEtlData TransformationData Quality Validation

Pune, India

Skills:

snowflake KafkaELTDevopsSparkIncident ManagementDatabricksRestful ApisAzureEtlAWSAirflowMonitoring

Pune, India

Skills:

Data GovernanceData ModelingSqlPythondata orchestration workflowsdata quality frameworksdistributed data processing frameworksETL pipelines

Beware of Scammers

We don’t charge money for job offers