Search by job, company or skills

Senior Data Engineer - Pyspark & AWS

5-7 Years
Early Applicant
  • Posted 9 hours ago
  • Be among the first 10 applicants

Job Description

Senior Data Engineer - Pyspark & AWS

Experience: 5-7 Years

Location: Pune

Work Mode: Work from Office (5 days working)

Job Summary:

We are seeking an experienced and leadership-minded Senior Data Engineer - Pyspark & AWS to drive the reliability, scaling, and operation of our enterprise data infrastructure. In this role, you will balance core development with production support and operational excellence, ensuring high availability through proactive monitoring and debugging of cloud data streams.

As a senior member of the data team, you will take full ownership of the data ecosystem, troubleshoot complex live incidents, and actively mentor and upskill junior engineers.

Key Responsibilities:

Data Engineering & Pipeline Design:

Orchestration: Program complex workflow dependencies and multi-stage DAGs within Apache Airflow.

ETL Architecture: Construct serverless data preparation workflows using AWS Glue and PySpark.

Compute Processing: Deploy targeted, cost-optimized scripts via AWS Lambda for event-driven backend tasks.

Storage Management: Structure secure Amazon S3 data lakes with clear landing, raw, and gold curation tiers.

Leadership & Mentorship:

Technical Ownership: Accept end-to-end accountability for platform health, security compliance, and architecture updates.

Team Mentorship: Guide junior developers through systematic code reviews, architectural advice, and pair programming sessions.

Process Standards: Establish baseline engineering documentation, testing frameworks, and on-call rotation protocols.

Technical Skills:

Primary Skills (Required)

AWS Production Environment: Background supporting enterprise-scale cloud infrastructures under tight data freshness targets.

Observability & Triage: Hands-on triage using AWS CloudWatch, CloudTrail, X-Ray, or third-party log collectors.

Serverless Ingestion: Advanced setup of AWS Glue (Crawlers, Catalogs, Jobs) and event-driven AWS Lambda components.

Data Lakes & Orchestration: Deep mastery of Amazon S3 object storage policies and workflow automation with Apache Airflow.

Apache Spark: Processing large-scale datasets using distributed computing engines via PySpark or Spark SQL.

Python Programming: Advanced scripting for data automation, API bindings, custom transformations, and object mutation.

Database Foundations: Ability to query and locate performance degradation points inside Amazon Redshift, Athena, or RDS systems.

Production Support & Operations:

Production Support: Provide Tier 3 technical support for live cloud datasets, maintaining strict platform uptime SLAs.

System Monitoring: Build dashboards using AWS CloudWatch to track resource consumption and data flow metrics.

Issue Debugging: Inspect execution logs, trace data failures, and resolve runtime bottlenecks across active microservices.

Incident Management: Lead root-cause analysis (RCA) activities for data dropouts, pipeline breaks, or structural data quality shifts.

Qualifications & Experience:

Experience:5+ years of dedicated data engineering or platform support experience, with strong Pyspark programming skills and at least 3 years directly focused on AWS ecosystems.

Education: Bachelor's or Master's degree in Computer Science, Software Engineering, or an equivalent technical discipline.

Soft Skills: Highly analytical communicator capable of translating business roadblocks into immediate technical fixes.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151570053