Job Description
We are looking for 5+ years of experience in data engineering and data science roles across fintech, e-commerce, and consumer platforms for Bangalore Location.
Work from Office opportunity - Client Interaction Involved.
Key Responsibilities
- Architect and deploy scalable data pipelines to ensure seamless data ingestion and transformation across diverse enterprise environments.
- Implement and manage modern lakehouse architectures to provide a unified, performant interface for both batch and real-time analytics.
- Develop and optimize streaming applications using Apache Flink and Kafka to support low-latency data processing requirements.
- Automate data transformation workflows using dbt to ensure high-quality, consistent, and well documented data models for downstream consumption.
- Manage the lifecycle of data stored in Apache Iceberg tables to improve query performance and data reliability for analytical workloads.
- Collaborate with cloud infrastructure teams to provision and maintain secure, cost-effective environments that support high-volume data processing.
Skills Required Languages
- Proficient in Python and SQL.
Big Data & Data Engineering
- Expert in building large-scale data pipelines and lakehouse architectures using PySpark, DBT, Apache Kafka / Apache Flink, and Apache Iceberg.
Cloud Infrastructure
- Strong experience with AWS, including cost optimization (e.g., reducing EMR costs) and building cloud-agnostic MLOps platforms.
Data Orchestration & Workflow
- Hands-on expertise with Apache Airflow for automating and scaling production pipelines.
Data Warehousing & Storage
- Experience with Snowflake (Snow pipes, DBT models), Presto, and Ceph RADOS.
MLOps & Data Science
- Skilled in deploying NLP/ML models (such as sentence transformers and collaborative filtering), building end-to-end MLOps solutions, and feature engineering.
Analytics & Visualization
- Proficient in creating insights through Tableau and Python-based Dash.
Data Governance
- Experience evaluating and implementing tools like Immuta, Collibra, Atlan, and CastorDoc for cataloging and security.
(ref:hirist.tech)