S
Cloud Data Engineer
S
Cloud Data Engineer
swits digital private limited- Posted 2 hours ago
- Be among the first 10 applicants
Job Description
Job Title: Cloud Data Engineer - Location - Bangalore - (Hybrid - 3Days a Week)
Looking for Immediate Joiners / Serving Notice Peroid.
Experience
Candidates should have 5 7 years of hands-on experience in data engineering or big data platforms, with strong practical expertise in Hadoop, Hive, HDFS, and PySpark. Experience in Cloudera, Kubernetes, containerized Spark, CI/CD, and platform performance optimization will be an added advantage.
Looking for Immediate Joiners / Serving Notice Peroid.
Experience
- 5 7 years of relevant experience in Data Engineering, Big Data, or Cloud Data Engineering.
- Strong hands-on experience with Hadoop, Hive, HDFS, and PySpark is mandatory.
- Develop, maintain, and optimize scalable data engineering solutions using Python and Apache Spark/PySpark.
- Design and build reusable, scalable data ingestion and data transformation frameworks for large-scale data processing.
- Work extensively with the Hadoop ecosystem, including Hadoop, Hive, HDFS, and Impala, preferably in Cloudera-based environments.
- Demonstrate strong understanding of HDFS architecture and internals, including:
- Data storage architecture
- Partition management
- Compaction strategies
- Resource utilization
- Cluster performance optimization
- Troubleshoot complex platform, storage, compute, and cluster-level issues, perform detailed root cause analysis, and implement performance improvements.
- Work with a variety of database technologies and object storage platforms, ensuring optimized connectivity, data access, and processing patterns.
- Apply strong knowledge of distributed data processing, Spark optimization, and data platform engineering best practices.
- Develop and support applications using Kubernetes, including:
- Containerized Spark workloads
- Pod management
- Application scaling
- Troubleshooting
- Operational support
- Build, package, and deploy Spark workloads using JFrog-managed container images and automated CI/CD pipelines.
- Leverage AI Agents, Generative AI, and LLM-powered development tools to improve engineering productivity and accelerate software delivery.
- Collaborate with engineering and platform teams to develop reliable, scalable, and high-performance data solutions.
- Hadoop
- Hive
- HDFS
- PySpark / Apache Spark
- Python
- Distributed data processing
- Spark optimization
- Cloudera ecosystem
- Impala
- Kubernetes
- Containerized Spark workloads
- JFrog
- CI/CD pipelines
- Object storage platforms
- Scala development
- Generative AI / AI Agents / LLM-powered development tools
Candidates should have 5 7 years of hands-on experience in data engineering or big data platforms, with strong practical expertise in Hadoop, Hive, HDFS, and PySpark. Experience in Cloudera, Kubernetes, containerized Spark, CI/CD, and platform performance optimization will be an added advantage.
More Info
Key Skills
Object storage platforms
Scala development
Generative AI
Distributed data processing
AI Agents
Containerized Spark workloads
Spark optimization
LLM-powered development tools
CI CD pipelines



