

Search by job, company or skills

Lead Databricks Data Engineer
Type: Full-time
Location- Remote
Experience- 6+ years
About the role
We are hiring a Lead Databricks Data Engineer to own the data engineering workstream of a large-scale predictive analytics programme. You will architect and lead the build of production-grade data pipelines on Databricks, integrating complex enterprise source systems (SAP and Oracle) into a governed lakehouse that feeds ML models, scoring engines, and executive dashboards. You will lead a small engineering team and set the standards the whole programme builds on.
What you'll do
1-Architect end-to-end ingestion and transformation pipelines on Databricks (Delta Lake, Workflows, Unity Catalog) from SAP and Oracle source systems
2-Design the integration approach for SAP data extraction (PM, MM, S/4HANA) in read-only patterns, working within enterprise security and data governance constraints
3-Implement data quality validation, anomaly detection, and lineage tracking as first-class pipeline components with contractual quality thresholds
4-Lead and review the work of data engineers; own pipeline reliability through SIT, UAT, go-live, and hypercare
5-Establish CI/CD-based deployment practices for data assets in coordination with the client platform team
Contribute to solution architecture, estimation, and technical documentation
What you'll need
1-6+ years in data engineering, with production Databricks delivery (Delta Lake, Spark, Workflows)
2-Demonstrated experience integrating SAP systems (ideally PM/MM or S/4HANA) into cloud data platforms
3-Strong pipeline orchestration, data modelling, and data quality engineering skills
4-Experience leading small engineering teams in fixed-deadline delivery
5-Clear technical communication with architects, security teams, and business stakeholders
Nice to have
1-Databricks Certified Data Engineer Associate or Professional (certification support provided)
2-Experience in asset-intensive or regulated industries
3-Familiarity with Oracle P6 / EAM or finance data sources
4-Prior delivery experience in the GCC region
Interested candidates share their resume at [Confidential Information]
Job ID: 152471191
Skills:
Github, Machine Learning, Metadata Management, Pyspark, BI, ELT, Terraform, Advanced Analytics, Azure DevOps, Apache Spark, Sql, Jenkins, Data Quality, Data Governance, Databricks, Azure, Kubernetes, Etl, lineage, Privacy, infrastructure automation, Databricks Workflows, Generative AI, Unity Catalog, DataOps, CI CD pipelines, Delta Lake, compliance frameworks
Skills:
Pyspark, Apache Spark, Sql, Git, Terraform, Databricks, Azure DevOps, MLflow, Delta Live Tables, Unity Catalog, GitHub Actions, Delta Lake, Databricks Workflows
Skills:
amazon dynamodb , snowflake , Cassandra, Amazon S3, PostgreSQL, AWS Glue, Docker, Neo4j, Python, Aws Lambda, Oracle Sql Server, Apache Spark, Sql, Git, Amazon Redshift, Sqs, Apache Kafka, Databricks, MongoDB, Amazon Athena, AWS Step Functions, GraphDB, Amazon SNS, Amazon Neptune, Amazon ECS Fargate, AWS Cloud Technologies, CI CD pipelines, Microsoft Fabric
Skills:
Apache Spark, Emr, Redshift, Sql, Apache Airflow, Docker, Terraform, Kubernetes, Python, Aws S3, Data observability, BI platforms, Flink, Lake Formation, Glue, Athena, Data pipelines, Data modeling techniques
Skills:
Azure Databricks, Apache Spark RDD, Apache Spark SQL, Delta Lake, Azure Cosmos DB NoSQL