Search by job, company or skills

Data Engineer

Data Engineer

imerit technology
4-8 Years
Not Disclosed
Early Applicant
  • Posted 2 months ago
  • Be among the first 40 applicants

Job Description

iMerit is a leading AI data solutions company specializing in transforming unstructured data into structured intelligence for advanced machine learning and analytics applications. Our clients span autonomous mobility, medical AI, agriculture, and more—powering next-generation AI systems with high-quality data services.

About the Role

We are seeking a skilled Data Engineer to help scale and enhance our internal data observability and analytics platform. This platform integrates with data annotation tools and ML pipelines to provide visibility, insights, and automation across large-scale data operations. You will design and optimize robust data pipelines, build integrations with internal platforms (e.g., AngoHub, 3DPCT) and customer platforms, and support real-time metrics, dashboards, and workflows critical to customer delivery and operational excellence.

Key Responsibilities

Design and build scalable batch and real-time data pipelines across structured and unstructured sources.

Integrate analytics and observability services with upstream annotation tools and downstream ML validation systems to enable full-cycle traceability.

Collaborate with product, platform, and analytics teams to define event models, metrics, and data contracts.

Develop ETL/ELT workflows using tools like AWS Glue, PySpark, or Airflow; ensure data quality, lineage, and reconciliation.

Implement observability pipelines and alerts for mission-critical metrics (e.g., annotation throughput, quality KPIs, latency).

Build data models and queries to power dashboards and insights via tools like Athena, QuickSight, or Redash.

Contribute to infrastructure-as-code and CI/CD practices for deployment across cloud environments (preferably AWS).

Document architecture, data flow, and support runbooks; continuously improve platform performance and resilience.

Integrate with customer data platforms and pipelines, including bespoke data frameworks.

Minimum Qualifications

4–8 years of experience in data engineering or backend development in data-intensive environments.

Proficient in Python and SQL; familiarity with PySpark or other distributed processing frameworks.

Strong experience with cloud-native data tools and services (S3, Lambda, Glue, Kinesis, Firehose, RDS).

Familiarity with frameworks like Apache Hadoop, Apache Spark, and related tools for handling large datasets.

Experience with data lake and warehouse patterns (e.g., Delta Lake, Redshift, Snowflake).

Solid understanding of data modeling, schema design, and versioned datasets.

Data Governance and Security: Understanding and implementing data governanc policies and security measures.

Proven experience in building resilient, production-grade pipelines and troubleshooting live systems.

Working knowledge of messaging frameworks like Kafka, Firehose etc

Working knowledge of API frameworks, robust and performant API design

Good working knowledge of Database fundamentals, relational databases and SQL

Preferred Qualifications

Experience with observability/monitoring systems (e.g., Prometheus, Grafana, OpenTelemetry) is a plus.

Familiarity with data governance, RBAC, PII redaction, or compliance in analytics platforms.

Exposure to annotation/ML workflow tools or ML model validation platforms.

Comfort working in Agile, distributed teams using tools like Git, JIRA, and Slack.

Why Join Us

You'll work at the intersection of AI, data infrastructure, and impact—contributing to platforms that ensure AI is explainable, auditable, and ethical at scale. Join a team building the next generation of intelligent data operations.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

About Company

Similar Jobs

6-10 yrs
Bengaluru, India
Skills:
Agile Methodologies, Pyspark, Sql, Big Data Technologies, Microservices, Git, Gcp, Apache Kafka, Spark, Rest Apis, Azure, Python, AWS, Kafka Connect, CI CD tools, ETL ELT pipelines, Kafka Streams
5-8 yrs
Bengaluru, India
Skills:
BigQuery, Pyspark, Dataproc, Data Modeling, Sql, ELT, Git, Python, Etl, GCP data services, Cloud Dataflow, Pub Sub, Cloud Composer, CI CD tools, performance optimization, partitioning, GCS
6-8 yrs
Bengaluru, India
Skills:
Azure Data Factory (ADF), T-sql, Data Warehouse Concepts, Pyspark, SQL Server, Sql, Git, Python, Star Schema, Azure DevOps, Generative AI, AI-Assisted Development, Claude Claude Code, Fact Dimension modeling, Azure SQL Database, Stored Procedures, CI/CD, Microsoft Fabric
7-11 yrs
Bengaluru, India
Skills:
PostgreSQL, Sql, ELT, Python, Etl, canonical data modeling, orchestration tools, data contracts, AI coding assistants, lakehouse concepts, data-quality validation, cloud data platforms, config-driven architecture
8-10 yrs
Bengaluru, India
Skills:
snowflake , Github, Data Modeling, Cloudformation, Sql, Apache Airflow, Git, Terraform, Data Warehousing, Python, CI CD Pipelines, AWS IAM Security, AWS Cloud Services, dbt, Monitoring