Key Skills & Technologies
Core Tech Stack
- Data Engineering Tools: DBT, Databricks, PySpark, Apache Airflow, Prefect
- Databases & Querying: SQL, Delta Lake, Apache Iceberg
- Cloud & Infrastructure: AWS (EC2, S3, EMR, Kinesis), Terraform, AWS CDK, Pulumi
- Additional Tools (Preferred): Sigma, Kinesis, EMR
Key Competencies
- Designing and maintaining scalable data models and transformation layers
- Building and operating robust data pipelines (batch & streaming)
- Optimizing Spark/DBT workloads for cost, performance, and reliability
- Orchestrating end-to-end workflows with SLA-driven execution
- Collaborating with Data Science teams on feature engineering and ML pipeline productionization
- Infrastructure-as-Code (IaC) practices and cloud-native deployments
- Data consistency, latency, throughput, and fault-tolerance trade-offs in distributed systems
Soft Skills & Attributes
- Self-starter mindset with bias for action and ownership
- Strong communication (written & verbal) in English
- Collaborative, curious, and passionate about solving novel problems
- Interest or experience in supply chain data challenges (preferred)
Bonus (Nice-to-Have)
- Experience with LLMs: building, evaluating, or integrating AI into data pipelines or internal tools
Shift Timing: Evening Shift (6:30 PM onwards)
Work Location: Onsite – Hyderabad
Key Responsibilities
- Design and maintain DBT models that produce trusted datasets, features, and metrics for Data Science, ML, analysis, and reporting.
- Build and operate scalable data pipelines in Databricks using PySpark and Delta/Iceberg tables to transform raw operational events into analysis-ready data.
- Develop deep business domain knowledge (especially in operations) to ensure data models reflect real-world workflows.
- Orchestrate end-to-end data workflows in Airflow (and Prefect where applicable), ensuring SLA adherence for daily models, dashboards, and operational decisions.
- Participate in peer code reviews and champion data quality, testing, and documentation standards.
- Partner with data scientists to scope, design, and productionize feature pipelines and supporting data infrastructure.
- Optimize DBT and Spark workloads for cost, performance, and reliability as data volume scales.
- Adapt quickly to new technologies — curiosity and learning agility are essential.
Requirements
- 7+ years of experience building, testing, and deploying data engineering systems
- Experience with at least one distributed data system, with the ability to reason about consistency, latency, throughput, and fault tolerance
- Strong SQL proficiency and hands-on experience with at least one of: PySpark, DBT, or Airflow in production environments
- Familiarity or experience with Infrastructure-as-Code (e.g., Terraform, AWS CDK, Pulumi)
- Interest or understanding of supply chain data challenges (a plus)
- Self-driven, collaborative, and passionate about delivering high-impact solutions
Skills: dbt,airflow,databricks,sql,aws,pyspark