We're looking for a Senior Data Engineer to design, build, and scale distributed data pipelines and platforms. You'll own critical data infrastructure end-to-end — from ingestion to processing to storage — enabling reliable, high-quality data for analytics and downstream applications across the organization.
Responsibilities
- Design, build, and maintain large-scale batch and streaming data pipelines using Apache Spark and Kafka
- Develop and optimize distributed data processing jobs in Scala and Python
- Build and manage data ingestion pipelines using Kafka Connect and Hadoop/HDFS
- Design and maintain data warehousing solutions using Hive
- Develop, tune, and manage workflows on Databricks
- Write efficient, production-grade code in Scala, Java, Python, and where needed, C++
- Design and optimize relational data models and queries in MySQL
- Architect and deploy data solutions across cloud platforms (Azure, AWS, GCP)
- Ensure data quality, reliability, and performance through monitoring, testing, and validation frameworks
- Collaborate with data scientists, analysts, and backend engineers to understand data requirements and deliver scalable solutions
- Troubleshoot and resolve performance bottlenecks in large-scale distributed systems
- Contribute to architectural decisions around data platform design and technology selection
- Mentor junior data engineers and promote best practices in data engineering
Requirements
- 6–10 years of experience in data engineering or distributed systems development
- Strong programming proficiency in Scala and Java; working knowledge of Python; C++ a plus
- Hands-on expertise with Apache Spark (Scala-based) for large-scale data processing
- Strong experience with Apache Kafka and Kafka Connect for real-time data streaming and integration
- Solid experience with Hadoop/HDFS ecosystem and Hive for data warehousing
- Hands-on experience with Databricks for data engineering and pipeline orchestration
- Strong SQL skills and experience with MySQL for relational data management
- Experience deploying and managing data platforms on at least one major cloud provider (Azure, AWS, or GCP); multi-cloud experience preferred
- Proficiency with standard development tooling (IntelliJ, Eclipse, PyCharm, VS Code, VIM)
- Strong understanding of distributed systems concepts, data partitioning, and performance optimization
- Experience with CI/CD practices and version control (Git) for data pipeline deployment
- Excellent problem-solving skills and ability to work with large, complex datasets
- Strong communication skills and ability to collaborate across engineering, analytics, and business teams
Nice to Have
- Experience with data orchestration tools (Airflow, Databricks Workflows, etc.)
- Familiarity with data governance, lineage, and cataloging tools
- Exposure to machine learning pipelines or MLOps
- Experience with infrastructure-as-code (Terraform, CloudFormation)