Company Overview:
- Mid-Sized Pioneering IT and Engineering Services Company
- Domains: Hi-Tech, Automotive, Manufacturing, Telecom, Medical and Life Sciences, Pharmaceutical
- Successfully service Fortune 500 Companies
- Customer Geographies: North America, Europe, Japan, Korea, China
Job Description – Senior Data Engineer
ABOUT THE ROLE:
We are looking for an accomplished Senior Data Engineer for the design, development, & deployment of cutting-edge data solutions. In this leadership role, you will drive technical strategy, mentor a team of data engineers, and collaborate with cross-functional stakeholders to deliver impactful, production-ready data systems. You will serve as technical authority on ingestion, processing and transformation, data modeling, translating complex business problems into scalable solutions.
KEY RESPONSIBILITIES:
- Design & implement governed ingestion pipelines consuming on-site & off-site events into canonical event schema
- Build and maintain Kafka or Kinesis-based streaming pipelines with schema validation, data quality checks, and alerting on source failures
- Integrate batch ingestion from databases and alongside the streaming path, managing orchestration via Airflow or equivalent
- Enforce privacy flags at ingestion time quarantining or anonymising events without valid consent before they reach downstream layers
- Maintain a raw event store with partitions by date and source, serving as the audit source for the pipeline
- Implement deterministic identity across fragmented systems
- Build and maintain the cookie-to-login stitching mechanism
- Design identity logic spanning on-site sessions, CRM records etc.
- Monitor identity resolution quality — match rate, false positive rate, unresolved session ratio — and iterate on matching logic to improve coverage over time
- Compute windowed aggregations over the event store
- Build and maintain governed, versioned feature definitions in a feature store ensuring normalization, encoding, and embedding lookup logic is consistent between training and serving pipelines
- Collaborate with data scientists to translate model feature requirements into production-grade pipeline implementations with no training-serving skew
- Implement automated data contract validation between pipeline layers, schema compatibility checks, completeness assertions, and anomaly detection
- Define and enforce SLAs on pipeline freshness, completeness, and accuracy with monitoring dashboards and escalation paths for breaches
- Ensure PII is classified, tagged, and handled according to GDPR and CCPA requirements at every stage of the pipeline like ingestion, storage, feature compute, and serving
- Contribute and maintain data lineage document, full traceability from raw event to feature to model prediction
- Work closely with the Data Architect to implement data contracts and schema agreements between pipeline layers, and flag design risks early
- Partner with ML Engineers to ensure feature pipelines meet model training and online inference requirements
- Participate in code reviews, contribute to engineering standards, and mentor junior engineers where applicable
REQUIRED QUALIFICATIONS:
- 6-8 years of hands-on data engineering experience, with at least 2 years in a senior or lead capacity on production systems
- Demonstrable experience building and operating event-driven, streaming data pipelines at scale in a production environment
- Prior experience on a personalization, recommendation, or user behavioral analytics platform is strongly preferred
- Experience working within a regulated data environment like GDPR, CCPA, or equivalent.
- Batch and structured streaming, including windowed aggregations, stateful processing, and performance tuning
- Event streaming platforms including producer/consumer design, partition management, and exactly-once semantics
- Data Lakehouse engineering like Delta Lake, Apache Iceberg, or equivalent, including ACID transactions, schema evolution, and time travel
- Pipeline orchestration using Apache Airflow, AWS Glue, Azure Data Factory, or Databricks Workflows including DAG design, dependency management, and failure handling
- SQL and Python at production standard quality with clean, tested, version-controlled code
- Cloud data platform exposure on at least one of: AWS (S3, Glue, Kinesis, Redshift), Azure (ADLS, Data Factory, Synapse), or Databricks
- Feature store design and operation: Databricks Feature Store, Feast, Tecton, or equivalent
- Data quality frameworks: Great Expectations, dbt tests, or equivalent for automated pipeline validation
- Deterministic and probabilistic matching, entity deduplication, graph-based stitching
- CI/CD for data pipelines with automated testing, deployment, and monitoring using GitHub Actions, Azure DevOps, or equivalent
PREFERRED QUALIFICATIONS:
- Experience with graph data modelling and graph processing frameworks — Neo4j, Amazon Neptune, GraphX, or GraphFrames
- Familiarity with vector embedding pipelines - batch encoding of text at scale, embedding storage, and integration with ANN search infrastructure
- Working knowledge of NLP pipeline engineering — tokenisation, embedding generation, and chunking for unstructured text at scale
- Experience with real-time feature serving and integration like Redis, DynamoDB, or equivalent cache stores
- Familiarity with MLflow or equivalent model registry — understanding of how feature pipelines connect to model training and deployment workflows
- IaC experience like Terraform, Bicep, or AWS CDK for reproducible data infrastructure deployment
- Experience with data mesh or federated data architecture patterns in multi-team environments
- Strong written and verbal communication — able to explain complex pipeline design decisions to non-engineering stakeholders clearly
- Able to make design decisions where requirements are incomplete and flag risks proactively
- Collaborative working style with cross-functional teams spanning data engineering, ML, platform, and product
- Understanding that interfaces between pipeline layers are as important as the implementations within them
- Takes responsibility for pipeline reliability, DQ, and SLA adherence end-to-end, not just the code written
WHAT WE OFFER:
- Awesome Culture: Creative Synergies has a flat organization and an agile culture of positivity, entrepreneurial spirit, customer centricity, celebrating technical excellence, teamwork, and meritocracy
- Opportunity to work with Customers who are technology Leaders (including Global Fortune 500 Customers) & work on Real-World Problems that matter and are often mission-critical
- Leadership role with significant influence over AI strategy and team direction.
- Access to state-of-the-art GPU infrastructure and cutting-edge AI tools.
- Competitive compensation package with performance-based incentives.
- Flexible working arrangements with hybrid options.
- Continuous learning budget for conferences, courses, and certifications.