B.Tech/BS degree in Computer Science, Computer Engineering, Statistics, or other Engineering disciplines.
6+ years of experience in data engineering, building scalable and reliable data pipelines in production environments.
Strong experience with any cloud data platforms such as Azure/ AWS/ GCP.
Hands-on expertise with distributed data processing frameworks such as Apache Spark for large-scale batch and/or streaming data processing.
Expert-level SQL skills with the ability to write optimized queries, perform complex transformations, and tune queries for large-scale analytical workloads.
Proficiency in Python for data engineering tasks including pipeline development, automation, data transformations, and integration with APIs or cloud services.
Strong knowledge of data modeling techniques, including dimensional modeling, star/snowflake schemas, and designing data models optimized for analytics and reporting.
Experience designing scalable ETL/ELT architectures, ensuring high data quality, reliability, and performance for large-scale data platforms.
Strong experience or knowledge of Databricks will be a plus.
Responsibilities
Design, implementation, and improvement of processes & automation of Data infrastructure.
Tuning of Data pipelines for reliability & performance.
Building tools and scripts to develop, monitor, and troubleshoot ETLs.
Perform scalability, latency, and availability tests on a regular basis.
Perform code reviews and QA data imported by various processes.
Investigate, analyze, correct and document reported data defects.
Create and maintain technical specification documentation.