Search by job, company or skills

Senior Data Engineer

Early Applicant
  • Posted 2 days ago
  • Be among the first 10 applicants

Job Description

Company Overview:

  • Mid-Sized Pioneering IT and Engineering Services Company
  • Domains: Hi-Tech, Automotive, Manufacturing, Telecom, Medical and Life Sciences, Pharmaceutical
  • Successfully service Fortune 500 Companies
  • Customer Geographies: North America, Europe, Japan, Korea, China

Job Description – Senior Data Engineer

ABOUT THE ROLE:

We are looking for an accomplished Senior Data Engineer for the design, development, & deployment of cutting-edge data solutions. In this leadership role, you will drive technical strategy, mentor a team of data engineers, and collaborate with cross-functional stakeholders to deliver impactful, production-ready data systems. You will serve as technical authority on ingestion, processing and transformation, data modeling, translating complex business problems into scalable solutions.

KEY RESPONSIBILITIES:

  • Design & implement governed ingestion pipelines consuming on-site & off-site events into canonical event schema
  • Build and maintain Kafka or Kinesis-based streaming pipelines with schema validation, data quality checks, and alerting on source failures
  • Integrate batch ingestion from databases and alongside the streaming path, managing orchestration via Airflow or equivalent
  • Enforce privacy flags at ingestion time quarantining or anonymising events without valid consent before they reach downstream layers
  • Maintain a raw event store with partitions by date and source, serving as the audit source for the pipeline
  • Implement deterministic identity across fragmented systems
  • Build and maintain the cookie-to-login stitching mechanism
  • Design identity logic spanning on-site sessions, CRM records etc.
  • Monitor identity resolution quality — match rate, false positive rate, unresolved session ratio — and iterate on matching logic to improve coverage over time
  • Compute windowed aggregations over the event store
  • Build and maintain governed, versioned feature definitions in a feature store ensuring normalization, encoding, and embedding lookup logic is consistent between training and serving pipelines
  • Collaborate with data scientists to translate model feature requirements into production-grade pipeline implementations with no training-serving skew
  • Implement automated data contract validation between pipeline layers, schema compatibility checks, completeness assertions, and anomaly detection
  • Define and enforce SLAs on pipeline freshness, completeness, and accuracy with monitoring dashboards and escalation paths for breaches
  • Ensure PII is classified, tagged, and handled according to GDPR and CCPA requirements at every stage of the pipeline like ingestion, storage, feature compute, and serving
  • Contribute and maintain data lineage document, full traceability from raw event to feature to model prediction
  • Work closely with the Data Architect to implement data contracts and schema agreements between pipeline layers, and flag design risks early
  • Partner with ML Engineers to ensure feature pipelines meet model training and online inference requirements
  • Participate in code reviews, contribute to engineering standards, and mentor junior engineers where applicable

REQUIRED QUALIFICATIONS:

  • 6-8 years of hands-on data engineering experience, with at least 2 years in a senior or lead capacity on production systems
  • Demonstrable experience building and operating event-driven, streaming data pipelines at scale in a production environment
  • Prior experience on a personalization, recommendation, or user behavioral analytics platform is strongly preferred
  • Experience working within a regulated data environment like GDPR, CCPA, or equivalent.
  • Batch and structured streaming, including windowed aggregations, stateful processing, and performance tuning
  • Event streaming platforms including producer/consumer design, partition management, and exactly-once semantics
  • Data Lakehouse engineering like Delta Lake, Apache Iceberg, or equivalent, including ACID transactions, schema evolution, and time travel
  • Pipeline orchestration using Apache Airflow, AWS Glue, Azure Data Factory, or Databricks Workflows including DAG design, dependency management, and failure handling
  • SQL and Python at production standard quality with clean, tested, version-controlled code
  • Cloud data platform exposure on at least one of: AWS (S3, Glue, Kinesis, Redshift), Azure (ADLS, Data Factory, Synapse), or Databricks
  • Feature store design and operation: Databricks Feature Store, Feast, Tecton, or equivalent
  • Data quality frameworks: Great Expectations, dbt tests, or equivalent for automated pipeline validation
  • Deterministic and probabilistic matching, entity deduplication, graph-based stitching
  • CI/CD for data pipelines with automated testing, deployment, and monitoring using GitHub Actions, Azure DevOps, or equivalent

PREFERRED QUALIFICATIONS:

  • Experience with graph data modelling and graph processing frameworks — Neo4j, Amazon Neptune, GraphX, or GraphFrames
  • Familiarity with vector embedding pipelines - batch encoding of text at scale, embedding storage, and integration with ANN search infrastructure
  • Working knowledge of NLP pipeline engineering — tokenisation, embedding generation, and chunking for unstructured text at scale
  • Experience with real-time feature serving and integration like Redis, DynamoDB, or equivalent cache stores
  • Familiarity with MLflow or equivalent model registry — understanding of how feature pipelines connect to model training and deployment workflows
  • IaC experience like Terraform, Bicep, or AWS CDK for reproducible data infrastructure deployment
  • Experience with data mesh or federated data architecture patterns in multi-team environments
  • Strong written and verbal communication — able to explain complex pipeline design decisions to non-engineering stakeholders clearly
  • Able to make design decisions where requirements are incomplete and flag risks proactively
  • Collaborative working style with cross-functional teams spanning data engineering, ML, platform, and product
  • Understanding that interfaces between pipeline layers are as important as the implementations within them
  • Takes responsibility for pipeline reliability, DQ, and SLA adherence end-to-end, not just the code written

WHAT WE OFFER:

  • Awesome Culture: Creative Synergies has a flat organization and an agile culture of positivity, entrepreneurial spirit, customer centricity, celebrating technical excellence, teamwork, and meritocracy
  • Opportunity to work with Customers who are technology Leaders (including Global Fortune 500 Customers) & work on Real-World Problems that matter and are often mission-critical
  • Leadership role with significant influence over AI strategy and team direction.
  • Access to state-of-the-art GPU infrastructure and cutting-edge AI tools.
  • Competitive compensation package with performance-based incentives.
  • Flexible working arrangements with hybrid options.
  • Continuous learning budget for conferences, courses, and certifications.

More Info

Job Type:
Industry:
Employment Type:

Job ID: 151584465

Similar Jobs

Bengaluru, India

Skills:

Spark SQLGitAdfPysparkSqlDelta LakeAutoloaderUnity Catalog

Bengaluru, India

Skills:

snowflake JavaUnixCPower BiOBIEEInformatica EtlPythonAWSAirflowOASCI CD pipelinesOracle PL-SQL

Bengaluru, India

Skills:

ElkPostgreSQLPysparkPrometheusKafkaGrafanaELTNosqlCeleryDockerNeo4jTerraformFlaskOraclePythonPytestCloudformationSqlJenkinsDjangoApache BeamFastAPIKubernetesEtlJanusGraphunittestGitHub ActionsOpenTelemetryBitbucket PipelinesasyncioPulsar

Bengaluru

Skills:

BigQueryAwsAzurePython

Bengaluru, India

Skills:

Azure Data FactoryMulesoftApi IntegrationSqlELTEtlData MappingAzure SQL ServerTroubleshooting