Search Jobs

Search by job, company or skills

Hadoop / PySpark

Hadoop / PySpark

Infosys
  • Posted 9 hours ago
  • Be among the first 10 applicants

Job Description

  • Primary skills:Technology->Big Data - Data Processing->PySpark,Technology->Big Data - Hadoop->Hadoop
  • Lead end-to-end delivery of big data solutions using Hadoop and PySpark, from requirements to production rollout.
  • Design and implement scalable batch/ETL pipelines for large datasets with strong focus on reliability and performance.
  • Drive technical architecture discussions, define standards, and ensure best practices for distributed data processing.
  • Optimize Spark jobs through partitioning strategies, caching, shuffle tuning, and efficient file formats where applicable.
  • Collaborate with stakeholders to translate business needs into technical specifications, delivery plans, and milestones.
  • Establish data quality checks, validation frameworks, and operational monitoring for production pipelines.
  • Perform root-cause analysis for pipeline failures and performance bottlenecks; implement preventive fixes.
  • Mentor engineers through code reviews, design reviews, and knowledge sharing to uplift team capability.
  • Ensure documentation, runbooks, and handover artifacts are created and maintained for support readiness.
  • Bachelor's or Master's degree (or equivalent) in Engineering/Technology/Computer Applications/Science (BE/BTech/MTech/MCA/MSc or equivalent).
  • 9–11 years of overall experience in data engineering and large-scale data processing environments.
  • Strong hands-on experience with Hadoop ecosystem components and distributed data processing concepts.
  • Strong hands-on experience building data pipelines using PySpark.
  • Proven ability to lead technical delivery, guide teams, and manage stakeholder expectations in consulting engagements. Preferred Qualifications:
  • Experience designing and implementing Spark-based data processing patterns (batch and incremental loads).
  • Strong understanding of data modeling and storage patterns for big data platforms (partitioning, compaction, schema evolution).
  • Experience with workflow orchestration and scheduling for data pipelines and dependency management.
  • Demonstrated expertise in production hardening: monitoring, alerting, SLAs, and incident management for data jobs.
  • Strong consulting mindset with ability to present solutions, document designs clearly, and influence technical decisions across teams.

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

data quality checks

workflow orchestration

root-cause analysis

ETL pipelines

About Company