Search Jobs

Search by job, company or skills

Hadoop / PySpark

Hadoop / PySpark

Infosys Limited
  • Posted 13 hours ago
  • Be among the first 10 applicants

Job Description

Responsibilities :

Lead end-to-end delivery of big data solutions using Hadoop and PySpark, from requirements to production rollout.

Design and implement scalable batch/ETL pipelines for large datasets with strong focus on reliability and performance.

Drive technical architecture discussions, define standards, and ensure best practices for distributed data processing.

Optimize Spark jobs through partitioning strategies, caching, shuffle tuning, and efficient file formats where applicable.

Collaborate with stakeholders to translate business needs into technical specifications, delivery plans, and milestones.

Establish data quality checks, validation frameworks, and operational monitoring for production pipelines.

Perform root-cause analysis for pipeline failures and performance bottlenecks implement preventive fixes.

Mentor engineers through code reviews, design reviews, and knowledge sharing to uplift team capability.

Ensure documentation, runbooks, and handover artifacts are created and maintained for support readiness.

Additional Responsibilities:

Bachelor's or Master's degree (or equivalent) in Engineering/Technology/Computer Applications/Science (BE/BTech/MTech/MCA/MSc or equivalent).

9-11 years of overall experience in data engineering and large-scale data processing environments.

Strong hands-on experience with Hadoop ecosystem components and distributed data processing concepts.

Strong hands-on experience building data pipelines using PySpark.

Proven ability to lead technical delivery, guide teams, and manage stakeholder expectations in consulting engagements. Preferred Qualifications:

Experience designing and implementing Spark-based data processing patterns (batch and incremental loads).

Strong understanding of data modeling and storage patterns for big data platforms (partitioning, compaction, schema evolution).

Experience with workflow orchestration and scheduling for data pipelines and dependency management.

Demonstrated expertise in production hardening: monitoring, alerting, SLAs, and incident management for data jobs.

Strong consulting mindset with ability to present solutions, document designs clearly, and influence technical decisions across teams.

Technical and Professional Requirements:

Primary skills:Technology- Big Data - Data Processing- PySpark,Technology- Big Data - Hadoop- Hadoop

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

Big Data - Hadoop

Big Data - Data Processing

About Company

Similar Jobs

Bengaluru, India
Skills:
Big Data - Hadoop, Big Data - Data Processing, Hadoop Administration, Hadoop, Pyspark, Technology
2-5 yrs
Bengaluru, India
Skills:
Hadoop, Pyspark, Scala, Kafka, Data Modeling, HBase, ELT, Hive, Etl, Airflow, HDFS, Spark optimization