Search by job, company or skills

AI Product Experience Harness

15-21 Years
  • Posted 17 hours ago
  • Be among the first 10 applicants

Job Description

Role- AI Product Experience Harness

Location- Noida & Hyderabad

Experience- 15-21 Years

Role Summary:

Build and run the automated harness that validates the delivered product against the experience specification and the AI compliance/guardrail specification, so validation is a pipeline, not a headcount. This role executes automated checks against criteria that the Experience lead (functional/visual/accessibility) and the Compliance lead (safety/red-team/regulatory) define — it does not independently set the bar for what pass means; it makes pass mechanically checkable and evidence-backed.

Key Responsibilities:

Functional & experience validation:

  • Build automated end-to-end test suites (Playwright or Cypress) covering critical user flows
  • Build visual-regression testing into the pipeline (screenshot diffing against approved baselines)
  • Build automated accessibility test suites (WCAG 2.1/2.2 AA conformance checks)
  • Convert every experience-spec acceptance criterion into an executable check; flag criteria that cannot be automated and require manual review — this is a real skill, not a formality: a spec clause like feels responsive needs to become a measurable threshold (e.g., time-to-first-token, time-to-interactive) before it's testable

AI/agentic-specific validation:

  • Build automated checks for agentic UI states defined by the Experience spec streaming/incremental rendering correctness, progress-indicator accuracy during multi-step agent tasks, correct rendering of failure/retry/escalation states
  • Build and run automated red-team regression suites (using tools defined by/coordinated with the Compliance role) so that known jailbreak/injection/PII-leakage test cases are re-run on every release, not just at initial audit
  • Automate LLM output evaluation checks — factual consistency, refusal-behavior conformance, format/schema conformance — as repeatable pipeline gates rather than one-off manual review

Operations:

  • Run the harness against CTO org deliveries; produce pass/fail evidence per release, feeding the audit/evidence pack the Compliance role maintains
  • Maintain test data, environments, and fixtures, including synthetic/adversarial prompt datasets for AI-specific test suites
  • Publish a live quality dashboard for the product org, showing pass rates, flaky-test rates, and trend lines release over release
  • Integrate all of the above into CI/CD so validation runs automatically pre-merge and pre-release, not on request

Must-Have Requirements:

  • 12+ years full-stack engineering / SDET with real automation depth in Playwright or Cypress — not just test-writing, but framework/harness architecture
  • Has built a test harness from scratch for at least one product — be ready to describe the architecture decisions, not just the tests inside it
  • CI/CD pipeline integration experience (GitHub Actions, Jenkins, GitLab CI, or equivalent) — tests must run automatically, gating merges/releases
  • Can read a spec and identify untestable clauses — work sample: given a sample spec, identify which acceptance criteria can't be automated as written and propose a measurable alternative
  • Has built or integrated visual-regression testing (Percy, Chromatic, Applitools, BackstopJS, or Playwright's native visual comparison
  • Practical experience automating evaluation of LLM/agentic outputs as part of a test pipeline — using named tools such as Ragas, DeepEval, Promptfoo, or a custom eval harness — be ready to describe one specific automated check and what it caught
  • Working familiarity with red-team test automation for LLM applications (Garak, PyRIT, or equivalent), sufficient to integrate Compliance-authored test cases into the CI pipeline as repeatable regression checks — this role executes and automates red-team suites; it does not independently design the adversarial test strategy (that's the Compliance role's ownership)

Preferred:

  • Accessibility testing depth beyond automated checks (manual screen-reader testing, WCAG AAA familiarity)
  • Performance testing (Lighthouse, k6, WebPageTest) integrated into the same pipeline
  • Direct prior collaboration with an AI-safety/compliance function, having built the automation layer for someone else's test design

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152946775

Beware of Scammers

We don’t charge money for job offers