Search Jobs

Search by job, company or skills

Data Scientist

Data Scientist

techaivv technologies
Early Applicant
  • Posted 6 days ago
  • Be among the first 10 applicants

Job Description

  • Own canonical data modeling: define stable, source-agnostic entities and relationships that represent the business domain (e.g., payroll constructs such as Gross Pay, Net Pay, Regular Pay, Bonus, Termination) independent of any single vendor's data shape.
  • Design config-driven architecture: ensure new source systems, clients, or field variations are onboarded through configuration and mapping rules rather than one-off code changes or schema forks.
  • Enforce correctness and reconciliation discipline: define and validate that canonical data reconciles against source system totals and known ground truth, and build in checks that surface drift or silent data-quality regressions early.
  • Own production engineering and abstraction discipline: review and approve how abstraction layers are actually implemented in code, ensuring the conceptual model and the production implementation do not diverge over time.
  • Partner with the Solution Architect and Data Integration Architect: this role owns the correctness and stability of the canonical model itself, while those roles own end-to-end platform coherence and pipeline/ingestion mechanics respectively.
  • Lead data modeling reviews and design walkthroughs for new entities, new client onboarding, or new source-system integrations, pressure-testing proposed models against edge cases and real production data.
  • Define data governance practices for the canonical layer: versioning of the model, change management for schema evolution, lineage, and documentation standards.
  • Validate architecture and modeling decisions directly against real client data (not sampled or synthetic data) before sign-off, and flag contradictions between assumed and actual data behavior.
  • Create architecture artifacts (canonical model documentation, ADRs, entity-relationship diagrams, mapping specifications) and maintain them as the model evolves.
  • Mentor data engineers on modeling discipline and abstraction practices; review code and configuration for adherence to the canonical model rather than workarounds.
  • Support onboarding of new foundation clients by assessing how their source systems map to the existing canonical model and where the model needs to be extended versus where a client-specific exception is being incorrectly introduced.

Required Qualifications

  • 7+ years (Senior Consultant) or 11+ years (Manager) total experience in data architecture or data modeling roles, including demonstrated ownership of a canonical or common data model spanning multiple source systems or clients.
  • Strong hands-on experience with dimensional and canonical data modeling, including designing entities that stay stable while underlying source systems vary or change.
  • Demonstrated experience designing config-driven (not hardcoded) architecture for onboarding new data sources, clients, or variations without redesigning the core model.
  • Strong track record of building or enforcing reconciliation and data-quality validation practices against ground-truth data at production scale.
  • Proficiency in SQL and at least one programming language (Python preferred) sufficient to review and validate production data pipeline code against the intended model.
  • Experience with relational databases (PostgreSQL preferred) for canonical model implementation at scale.
  • Ability to distinguish role-level architecture responsibility from years of experience alone - evaluating what a candidate actually owned versus what they were exposed to or supervised.
  • Hands-on proficiency with AI coding assistants (Claude Code, GitHub Copilot, Cursor, Windsurf, or equivalent) for architecture work, design documentation, and engineering workflows. Familiarity with agentic engineering patterns is expected.
  • Strong communication and stakeholder management skills; ability to explain modeling trade-offs to both engineers and business/delivery stakeholders.

Preferred Qualifications

  • Experience building a canonical data model specifically for payroll, HR, finance, or other transaction-critical domains with strict correctness requirements.
  • Experience onboarding multiple vendor systems (e.g., ADP, Workday, SAP, or equivalents in other domains) into a single normalized model.
  • Familiarity with orchestration tools (Prefect, Airflow, Dagster, or equivalent) sufficient to understand how the canonical model is populated and refreshed in production, even if not owning the orchestration layer directly.
  • Experience with cloud data platforms (Azure preferred) hosting the canonical model and its supporting infrastructure.
  • Experience defining data contracts or schemas consumed by downstream AI/ML models, understanding how model-readiness requirements should shape canonical model design.
  • Experience in regulated industries and implementing audit-ready data governance and change-control practices.
  • Cross-functional fluency with data integration, cloud, and AI architecture patterns - working familiarity with ETL/ELT and lakehouse concepts, cloud-native deployment, and LLM data consumption patterns. Not expected to own these, but must collaborate effectively with Data Integration, Cloud Solution, and AI Solutions architects in joint reviews.

Key Competencies

  • Rigorous, first-principles approach to canonical modeling - able to separate what is a stable business concept from what is an artifact of a particular source system.
  • Discipline around config-driven extensibility; instinctively pushes back on hardcoded, one-off solutions.
  • Correctness and reconciliation mindset - treats data validation against ground truth as non-negotiable, not an afterthought.
  • Production engineering ownership - reviews how abstractions are actually implemented, not just how they are designed on paper.
  • Strong architectural judgment and trade-off management; comfortable pushing back and flagging contradictions rather than deferring to assumptions.
  • Mentorship and technical leadership; comfortable with hands-on review of code, configuration, and data when needed.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

canonical data modeling

orchestration tools

data-quality validation

cloud data platforms

AI coding assistants

config-driven architecture

Similar Jobs

12-14 yrs
Hyderabad, India
Skills:
causal analysis , Hadoop, Cosmos, Sql, Data Science, Statistical Inference, Spark, Python, Generative AI, Forecasting, Analytics, R, Applied Statistics, experimentation, Responsible AI, Metric Design
5-7 yrs
Hyderabad, India
Skills:
Sql, Tensorflow, Git, Pytorch, Docker, Kubernetes, Python, prompt evaluation methodologies, scikit-learn, LLM prompt engineering, CI CD workflows, MLflow, context engineering
8-10 yrs
Hyderabad, India
Skills:
Machine Learning, Predictive Modeling, Python, decision systems, generative AI
8-10 yrs
Hyderabad, India
Skills:
Hadoop, Arima, Spark, Python, Exponential Smoothing, R, data visualization tools
5-10 yrs
Hyderabad, India
Skills:
Pyspark, Machine Learning, Cicd, Databricks, Azure, Python, Time Series Forecasting, ADF pipelines, Demand Forecasting