

Search by job, company or skills

Senior Data Engineer - Databricks | PySpark | dbt | Lakehouse Engineering
⚠️ This is NOT a typical Data Engineering position.
We are looking for a hands-on Senior Data Engineer with strong, production-grade Databricks experience - someone who has actually built and optimized Lakehouse/Medallion architectures, PySpark & dbt pipelines, PostgreSQL data layers, and CI/CD workflows, rather than someone with only surface-level exposure to these technologies.
⚠️ If Databricks is not a core part of your recent hands-on experience, this role may not be the right fit.
Job Title: Senior Data Engineer
Specialization: Data Engineering | Databricks | Lakehouse | Data Platforms | AI-Ready Data
Experience: 5+ Years
Location: BNG | HYD | PUN | MUM | GGN
Work Model: Hybrid
What We're Looking For
Mandatory-Strong hands-on experience required:
• Databricks – with practical experience building production data platforms and Lakehouse architectures.
• PySpark – strong production experience working with non-trivial data volumes.
• Python – advanced programming and data engineering experience.
• dbt – hands-on experience with staging, transformations, incremental models, testing, and documentation.
• CI/CD – Azure DevOps – experience deploying and managing data assets through CI/CD pipelines.
• PostgreSQL – database design, schema management, query optimization, and performance tuning.
• Strong understanding of Medallion Architecture and Delta Lake.
• Advanced SQL and experience working with large-scale datasets.
Good to Have
• MLflow and experience supporting ML/data science workflows.
• LangChain or similar AI/Agent frameworks.
• Exposure to AI Agents / AI-ready data platforms.
• Experience with SQL Server.
• Experience with DuckDB.
• Terraform / Infrastructure-as-Code exposure.
• Docker and containerized data workloads.
• Experience with APIs and structured/semi-structured formats such as JSON, CSV, Parquet, and Apache Arrow.
• Exposure to marketing analytics, marketing measurement, or enterprise consulting environments.
Why This Role Is Different
This isn't simply a role where you'll build pipelines and move data.
You'll be responsible for building the data foundation behind modeling, optimization, analytics, customer-facing applications, and AI Agents - with an emphasis on architecture, performance, data quality, governance, reproducibility, and production reliability.
Primary Skills (Mandatory):
Databricks | PySpark | Python | dbt | Azure DevOps / CI-CD | PostgreSQL | Advanced SQL | Delta Lake | Medallion Architecture
Good to Have:
MLflow | LangChain | AI Agents | SQL Server | DuckDB | Terraform | Docker
Job ID: 153645177
Skills:
amazon emr , Pyspark, Apache Hadoop, S3, RDS, AWS Glue, Dynamodb, Sql, Lambda, Kinesis, Ec2, Amazon Redshift, Spark, Python, Step Functions, Iceberg, Glue, Athena
Skills:
Pyspark, Dimensional Modeling, Jira, Sql, Git, Confluence, Gcp, Azure, Python, Azure DevOps, AWS, Etl, cdc, Medallion Architecture, SAP HANA to Snowflake migration, Snowflake Architecture, dbt, ELT Development, SCD Type 1
Skills:
Metadata Management, Databricks, Python, Data lineage techniques, Master data management, Data modeling methodologies, Big Data storage architecture
Skills:
Hadoop, Pyspark, Kafka, Presto, Spark, Python, AI tools, Pinot, pgvector, Codex, vector databases, Claude, Pinecone, Flink, cloud data warehouses, Hudi, Milvus, Weaviate
Skills:
snowflake , Python, Sql, Gen AI, data engineering design patterns, Microsoft DevOps, Matillion, Streamsets