Lead Data Engineer
Location : Chennai - onsite
Experience : 6+ years
Mandatory skills : Python, Pyspark, SQL, DSA, Cloud (AWS/GCP - 1st priority, Azure)
Job Description
Eucloid is looking for a Lead Data Engineer to join our Data Platform team supporting various business applications.
The ideal candidate will support development of data infrastructure for our clients by participating in activities which may include starting from up-stream and down-stream technology selection to designing and building of different components.
Candidate will also involve in projects like integrating data from various sources, managing big data pipelines that are easily accessible with optimized performance of overall ecosystem.
The ideal candidate is an experienced data wrangler who will support our software developers, database architects and data analysts on business initiatives.
You must be self-directed and comfortable supporting the data needs of cross-functional teams, systems, and technical solutions.
Key Skills
- B.Tech/BS degree in Computer Science, Computer Engineering, Statistics, or other Engineering disciplines
- 6+ years of experience in data engineering, building scalable and reliable data pipelines in production environments.
- Strong experience with any cloud data platforms such as Azure/ AWS/ GCP.
- Hands-on expertise with distributed data processing frameworks such as Apache Spark for large-scale batch and/or streaming data processing.
- Expert-level SQL skills with the ability to write optimized queries, perform complex transformations, and tune queries for large-scale analytical workloads.
- Proficiency in Python for data engineering tasks including pipeline development, automation, data transformations, and integration with APIs or cloud services.
- Strong knowledge of data modeling techniques, including dimensional modeling, star/snowflake schemas, and designing data models optimized for analytics and reporting.
- Experience designing scalable ETL/ELT architectures, ensuring high data quality, reliability, and performance for large-scale data platforms.
- Strong experience or knowledge of Databricks will be a plus.
Responsibilities
- Design, implementation, and improvement of processes & automation of Data infrastructure
- Tuning of Data pipelines for reliability & performance
- Building tools and scripts to develop, monitor, and troubleshoot ETLs
- Perform scalability, latency, and availability tests on a regular basis.
- Perform code reviews and QA data imported by various processes.
- Investigate, analyze, correct and document reported data defects.
- Create and maintain technical specification documentation.
(ref:hirist.tech)