I
Hadoop / PySpark
I
- Posted 8 hours ago
- Be among the first 10 applicants
Job Description
echnology->Big Data - Data Processing->PySpark,Technology->Big Data - Hadoop->Hadoop Administration->Hadoop
- Design, develop, and support scalable data processing workflows using Hadoop and PySpark for batch and large-volume processing.
- Build and maintain data pipelines that ingest, transform, and validate data from multiple sources into curated datasets.
- Optimize Spark jobs for performance (partitioning, caching, shuffle tuning) and improve overall pipeline efficiency and reliability.
- Perform data quality checks, reconciliation, and root-cause analysis for pipeline failures or data anomalies.
- Collaborate with cross-functional teams to understand requirements, translate them into technical solutions, and deliver within timelines.
- Create clear technical documentation for workflows, data mappings, and operational runbooks.
- Participate in code reviews, follow engineering best practices, and contribute to continuous improvement of standards and tooling.
- 2–5 years of experience in big data engineering or data processing roles.
- Bachelor's/Master's degree in Engineering, Computer Science, or equivalent (BTech/BE/MCA/MSc/MTech or related).
- Hands-on experience with Hadoop ecosystem concepts (HDFS, distributed processing fundamentals).
- Practical experience developing data transformations using PySpark.
- Strong problem-solving skills with the ability to debug data and job execution issues in distributed environments. Preferred Qualifications:
- Experience building end-to-end ETL/ELT pipelines and managing dependencies across multiple data workflows.
- Working knowledge of Spark optimization techniques and handling skewed/large datasets efficiently.
- Familiarity with data modeling concepts and designing curated datasets for analytics and reporting use cases.
- Exposure to production support practices such as monitoring, incident triage, and improving pipeline resiliency.
- Strong communication skills to collaborate effectively with stakeholders and explain technical trade-offs clearly. Good to have skills: Hive, HBase, Kafka, Airflow, Scala
