Search by job, company or skills

Lead Data Engineer

6-8 Years
  • Posted 4 hours ago
  • Be among the first 10 applicants

Job Description

Role Overview

We are looking for a Lead Data Engineer with strong experience building scalable batch, near-real-time, and streaming data platforms on Microsoft Azure.

The role requires hands-on expertise in Python, PySpark, advanced SQL, Azure data services, Medallion Architecture, and deploying Apache Spark workloads on Kubernetes or Azure Kubernetes Service (AKS).

.Key Responsibilities

  • Design and build batch, near-real-time, and streaming data pipelines on Azure.
  • Develop Bronze, Silver, and Gold data layers using Medallion Architecture.
  • Build and deploy containerized PySpark workloads on Kubernetes or AKS.
  • Configure Spark drivers, executors, CPU, memory, scaling, dependencies, and storage access.
  • Integrate data from REST APIs, SFTP, databases, files, enterprise systems, and Azure Event Hubs.
  • Develop complex transformation, cleansing, enrichment, reconciliation, and validation workflows.
  • Implement incremental loads, CDC, watermarking, deduplication, schema evolution, retries, and recovery.
  • Optimize Spark jobs, partitioning, shuffles, joins, file sizes, and query performance.
  • Implement monitoring, logging, alerting, audit controls, and data-quality checks.
  • Build reusable Python, PySpark, and SQL components.
  • Create CI/CD pipelines for Spark applications, Docker images, and Kubernetes deployments.
  • Review technical designs and support data engineers with implementation standards.

Required Skills

  • 6+ years of hands-on data engineering experience.
  • Strong experience with Microsoft Azure data platforms.
  • Advanced Python, PySpark, and SQL skills.
  • Strong hands-on experience with: Apache Spark, Kubernetes and AKS, Docker, Azure Data Lake Storage Gen2, Azure Event Hubs, Azure DevOps and Git
  • Experience deploying and operating Spark applications on Kubernetes.
  • Strong understanding of Spark drivers, executors, resource allocation, partitioning, caching, broadcast joins, shuffle optimization, and skew handling.
  • Experience with batch, streaming, ETL, ELT, and event-driven processing patterns.
  • Experience implementing Medallion Architecture.
  • Experience with REST APIs, SFTP, JSON, CSV, Parquet, Delta Lake, and relational databases.
  • Experience with CDC, incremental processing, schema enforcement, and schema evolution.
  • Strong understanding of data modelling, schema design, partitioning, and storage optimization.
  • Experience with Kubernetes Jobs, Cron Jobs, Config Maps, Secrets, resource limits, node pools, and autoscaling.
  • Experience implementing pipeline observability, data validation, monitoring, alerting, and error recovery.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 153637963

Similar Jobs

Pune, India

Skills:

Api ManagementApache SparkAzure DatabricksAzure Data FactoryTerraformDelta Lake ArchitectureAzure Data Lake Storage Gen2Service PrincipalsAzure Key VaultUnity CatalogDatabricks Asset BundlesrbacCI CD pipelinesEvent GridAzure Data Services

Pune, India

Skills:

CsvPysparkJsonELTGitDockerMongoDBRest ApisAdvanced SqlKubernetesPythonEtlOLAP engines

Pune, India

Skills:

Data LineageReactTypescriptJavascriptPythonPipeline BuilderAIP toolsetPalantir FoundryOSDK

Pune, India

Skills:

Apache SparkKafkaApache AirflowPandasElasticsearchMongoDBpgvectorPolarsVector databasesQdrantNoSQL databasesData models and storage architecturesDistributed data systemsData processing performance and optimisationData pipelinesMilvus

Pune, India

Skills:

Spark SQLS3RDSAws ServicesPysparkEmrJenkinsGitDevops ToolsPythonBig Data conceptsAWS architectureAWS cloud microservicescloud implementation projectsAuroraSnowflake SQLGlueData Vault 2.0

Beware of Scammers

We don’t charge money for job offers