Search by job, company or skills

Functional AI Tester

Functional AI Tester

Michelin
5-7 Years
Not Disclosed
  • Posted 18 hours ago
  • Be among the first 10 applicants

Job Description

We are seeking a Quality Assurance (QA) Engineer focused on testing Generative AI (GenAI) applications with a strong emphasis on Python-based test automation, GenAI evaluation, and ETL/data quality validation. You will design and execute end-to-end test strategies

that ensure our AI solutions are accurate, reliable, safe, and compliant.

About the Role

You will be involved in QA for GenAI features including Retrieval-Augmented Generation (RAG), conversational AI and Agentic

evaluations.

The role centers on:

• Systematic GenAI evaluation (qualitative and quantitative metrics)

• ETL and data quality testing for the data flows that feed AI systems

• Python-driven automated testing

This position is hands-on and collaborative, partnering with AI engineers, data engineers, and product teams to define measurable acceptance criteria and ship high-quality AI features.

Key Responsibilities

• Test strategy and planning

o Define risk-based test strategies and detailed test plans for GenAI features.

o Establish clear acceptance criteria with stakeholders for functional, safety, and data quality aspects.

• Python test automation

o Build and maintain automated test suites using Python (e.g., PyTest, requests).

o Implement reusable utilities for prompt/response validation, dataset management, and result scoring.

o Create regression baselines and golden test sets to detect quality drift.

• GenAI evaluation

o Develop evaluation harnesses covering factuality, coherence, helpfulness, safety, bias, and toxicity etc.

o Design prompt suites, scenario-based tests, and golden datasets for reproducible measurements.

o Implement guardrail tests including prompt-injection resilience, unsafe content detection, and PII redaction

checks.

o Track quality metrics over time.

• RAG and semantic retrieval testing

o Verify alignment between retrieved sources and generated answers.

o Verify adversarial tests.

o Measure retrieval relevance, precision/recall, grounding quality, and hallucination reduction.

• API and application testing

o Test REST endpoints supporting GenAI features (request/response contracts, error handling, timeouts).

• ETL and data quality validation

o Test ingestion and transformation logic; validate schema, constraints, and field-level rules.

o Implement data profiling, reconciliation between sources and targets, and lineage checks.

o Verify data privacy controls, masking, and retention policies across pipelines.

• Non-functional testing

o Performance and load testing focused on latency, throughput, concurrency, and rate limits for LLM calls.

o Cost-aware testing (token usage, caching effectiveness) and timeout/retry behavior validation.

o Reliability and resilience checks including error recovery and fallback behavior.

• Share results and insights; recommend remediation and preventive actions.

Required Qualifications

Experience

o 5+ years in software QA, including test strategy, automation, and defect management.

o 2+ years testing AI/ML or GenAI features, with hands-on evaluation design.

o 4+ years testing ETL/data pipelines and data quality.

Technical skills

o Python: Strong proficiency building automated tests and tooling (PyTest, requests, pydantic or similar).

o API testing: REST contract testing, schema validation, negative testing.

o GenAI evaluation: crafting prompt suites, golden datasets, rubric-based scoring, and automated evaluation

pipelines.

o RAG testing: retrieval relevance, grounding validation, chunking/indexing verification, and embedding checks.

o ETL/data quality: schema and constraint validation, reconciliation, lineage awareness, data profiling.

• Quality and governance

o Understanding of LLM limitations and methods to detect/reduce hallucinations.

o Safety and compliance testing including PII handling and prompt-injection resilience.

o Strong analytical and debugging skills across services and data flows.

Soft skills

o Excellent written and verbal communication; ability to translate quality goals into measurable criteria.

o Collaboration with AI engineers, data engineers, and product stakeholders.

o Organized, detail-oriented, and outcomes-focused.

Nice to Have

• Experience with evaluation frameworks or tooling for LLMs and RAG quality measurement.

• Experience creating synthetic datasets to stress specific behaviors

More Info

Job Type:
Industry:
Employment Type:

Key Skills

pydantic

GenAI evaluation

RAG testing

REST contract testing

requests

schema validation

ETL data quality

About Company