Search by job, company or skills

5-7 Years
Not Disclosed
  • Posted 2 hours ago
  • Be among the first 10 applicants

Job Description

Job Title: Cloud Data Engineer - Location - Bangalore - (Hybrid - 3Days a Week)

Looking for Immediate Joiners / Serving Notice Peroid.

Experience

  • 5 7 years of relevant experience in Data Engineering, Big Data, or Cloud Data Engineering.
  • Strong hands-on experience with Hadoop, Hive, HDFS, and PySpark is mandatory.

Key Responsibilities & Technical Skills

  • Develop, maintain, and optimize scalable data engineering solutions using Python and Apache Spark/PySpark.
  • Design and build reusable, scalable data ingestion and data transformation frameworks for large-scale data processing.
  • Work extensively with the Hadoop ecosystem, including Hadoop, Hive, HDFS, and Impala, preferably in Cloudera-based environments.
  • Demonstrate strong understanding of HDFS architecture and internals, including:
    • Data storage architecture
    • Partition management
    • Compaction strategies
    • Resource utilization
    • Cluster performance optimization
  • Troubleshoot complex platform, storage, compute, and cluster-level issues, perform detailed root cause analysis, and implement performance improvements.
  • Work with a variety of database technologies and object storage platforms, ensuring optimized connectivity, data access, and processing patterns.
  • Apply strong knowledge of distributed data processing, Spark optimization, and data platform engineering best practices.
  • Develop and support applications using Kubernetes, including:
    • Containerized Spark workloads
    • Pod management
    • Application scaling
    • Troubleshooting
    • Operational support
  • Build, package, and deploy Spark workloads using JFrog-managed container images and automated CI/CD pipelines.
  • Leverage AI Agents, Generative AI, and LLM-powered development tools to improve engineering productivity and accelerate software delivery.
  • Collaborate with engineering and platform teams to develop reliable, scalable, and high-performance data solutions.
Mandatory Skills

  • Hadoop
  • Hive
  • HDFS
  • PySpark / Apache Spark
  • Python
  • Distributed data processing
  • Spark optimization

Good to Have / Added Advantage

  • Cloudera ecosystem
  • Impala
  • Kubernetes
  • Containerized Spark workloads
  • JFrog
  • CI/CD pipelines
  • Object storage platforms
  • Scala development
  • Generative AI / AI Agents / LLM-powered development tools

Preferred Profile

Candidates should have 5 7 years of hands-on experience in data engineering or big data platforms, with strong practical expertise in Hadoop, Hive, HDFS, and PySpark. Experience in Cloudera, Kubernetes, containerized Spark, CI/CD, and platform performance optimization will be an added advantage.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

Object storage platforms

Scala development

Generative AI

Distributed data processing

AI Agents

Containerized Spark workloads

Spark optimization

LLM-powered development tools

CI CD pipelines

Similar Jobs

8-10 yrs
Bengaluru, India
Skills:
snowflake , Github, Data Modeling, Cloudformation, Sql, Apache Airflow, Git, Terraform, Data Warehousing, Python, CI CD Pipelines, AWS IAM Security, AWS Cloud Services, dbt, Monitoring
4-7 yrs
Bengaluru
Skills:
AWS, Sql, Databricks, AWS Glue, Python, Azure Data Factory, Azure, Gcp, Pyspark, Spark, GCP Dataflow
7-9 yrs
Bengaluru, India
Skills:
bigtable , Solr, Cassandra, Prometheus, Bash, HBase, Cloud Infrastructure, Datadog, Redis, Devops, Gcp, Terraform, Elasticsearch, Kubernetes, Python, AWS, SRE, Go, AI LLM tools, OpenSearch, OpenTofu, Search Data Platforms
3-6 yrs
Bengaluru, India
Skills:
Pl Sql, MongoDB, Python, Kubernetes, AWS, CI CD frameworks, AI best practices
5-7 yrs
Bengaluru, India
Skills:
data warehouses , BigQuery, Machine Learning, System Design, Data Governance, Automation, Sql, Python, Data Pipelines, Data Analysis