Search by job, company or skills

Data Engineer

Data Engineer

Randomtrees
4-5 Years
Not Disclosed
Quick Apply
  • Posted a month ago
  • Over 200 applicants have applied

Job Description

Role & Responsibilities:

  • Design, develop, and maintain ELT/ETL pipelines on GCP using Dataflow/Beam, Dataproc/Spark, and Airflow/Composer.
  • Model and optimize datasets in BigQuery using partitioning, clustering, materialized views, and UDFs.
  • Build streaming and near-real-time data ingestion using Pub/Sub, Dataflow, and CDC where applicable.
  • Implement data quality checks, validation frameworks, and SLAs, monitoring pipelines via Cloud Monitoring/Logging.
  • Optimize performance and cost across GCS, Dataproc autoscaling, and BigQuery slot usage.
  • Contribute to coding, CI/CD standards, observability, documentation, and perform code reviews.
  • Partner with Analytics, BI, and ML teams to productize datasets and ensure strong data contracts.
  • Support production operations, including on-call rotations for critical pipelines.

Preferred Candidate Profile:

  • Hands-on expertise in GCP data stack: BigQuery, Dataflow (Apache Beam), Dataproc, Cloud Storage, Pub/Sub, Cloud Composer (Airflow).
  • Strong experience with Spark (PySpark or Scala) for batch processing.
  • Solid Airflow DAG design knowledge (idempotent tasks, backfills, retries, SLAs).
  • Advanced SQL and data modeling (star/snowflake schemas, slowly changing dimensions, partition strategies).
  • Proficiency in Python (preferred) or Scala/Java.
  • Experience with Git and CI/CD tools (Cloud Build, GitHub Actions, GitLab CI).
  • Familiarity with GCP security and governance (IAM, service accounts, secrets management, VPC-SC).
  • Strong debugging, problem-solving, and communication skills.

Good-to-Have:

  • Snowflake (migration, performance tuning, tasks/streams).
  • Terraform for infrastructure as code on GCP.
  • Kafka or other streaming tools; Cloud Run/Functions for glue services.
  • BI exposure (Looker, Tableau, Power BI).
  • GCP Professional Data Engineer certification.

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

About Company

RandomTrees is a global provider of Data and AI services. Led by distinguished Industry experts with a passion to innovate with data. We help reinvent your enterprise with cutting edge Generative AI backed by our niche expertise to handle any type of data. We bring a unique set of data & and AI accelerators that help jumpstart & accelerate AI journey in the enterprise. Our differentiators are backed by the leading enterprise cloud partners and clients span across life sciences, Automotive and other manufacturing Industries Our portfolio comprises of: (1) Cloud data engineering services & Accelerators (2) Gen AI Products, Accelerators & services.

Similar Jobs

5-7 yrs
Hyderabad, India
Skills:
Software configuration management, Data Modeling, Data Architecture, Advanced Sql, Python, Data pipeline design
3-5 yrs
Hyderabad, India
Skills:
graph databases , object storage , S3, Golang, C#, AWS Glue, Data Modeling, Nosql, Lambda, Kinesis, Ruby, Oracle, Python, Java, Rust, C++, Emr, Redshift, Non-relational databases, Data stores, IAM roles and permissions, ETL pipelines, FireHose, AWS technologies, Key-value stores, Column-family databases
7-9 yrs
Hyderabad, India
Skills:
Apache Spark, Kafka, Apache Airflow, Pandas, Elasticsearch, MongoDB, pgvector, Polars, Qdrant, Vector databases, NoSQL databases, Data models and storage architectures, Distributed data systems, Data processing performance and optimisation, Data pipelines, Milvus
5-15 yrs
Bengaluru, Chennai, Hyderabad
Skills:
Gcp, Sql Server My Sql, Data Engineer
6-8 yrs
Hyderabad, India
Skills:
Azure Functions, Adf, Pyspark, Databricks, Azure, Sql, Synapse, Azure Purview, Delta Lake, Unity Catalog