Search Jobs

Search by job, company or skills

Senior Software QA Engineer Agentic AI Systems

Senior Software QA Engineer Agentic AI Systems

Readyly
3-5 Years
  • Posted an hour ago
  • Be among the first 10 applicants

Job Description

Remote | Full-time | ₹9L – ₹15L/year

About the Role

We're looking for a Senior Software QA Engineer to be the quality backbone of our agentic AI systems. This role is less about writing test scripts and more about deep session review — analyzing how our AI agents behave in the real world, spotting patterns in failures, and turning those insights into engineering action. You'll also play an ongoing role in strengthening our AI Eval framework and Smartflows, helping both evolve alongside the systems they're built to test.

What You'll Do
  • Review AI Sessions: Regularly review agent/LLM session transcripts and logs to evaluate output quality, reasoning correctness, and tool-use behavior.
  • Pattern-Level Analysis: Analyze session data across volume to identify recurring failure patterns — bad tool selection, context drift, hallucination clusters, agent loops, latency spikes, prompt regressions — rather than one-off bugs.
  • Create Engineering Tickets: Translate findings into clear, well-documented tickets with reproduction context, impact assessment, and priority, so engineering can act on them efficiently.
  • QA Verification Loop: Once engineering ships a fix, verify it against the original issue and re-check for regressions before sign-off and deployment.
  • Improve AI Eval & Smartflows: Continuously refine our AI Eval framework and Smartflows based on session findings — updating rubrics, scoring criteria, and flow logic so they keep pace with real-world agent behavior.
  • Process Improvement: Identify systemic gaps in the QA and engineering workflow and propose concrete improvements — better logging/observability, review cadences, ticket templates, evaluation rubrics, or automation opportunities.
  • Build Review Frameworks: Help define what good looks like for agent behavior — success criteria, scoring rubrics, or checklists that make session review consistent across the team.
  • Cross-functional Collaboration: Work closely with engineers, Customer Success, and Internal QA to surface real-world issues, align on priorities, and close the loop between observed problems and shipped fixes.
  • Track Quality Trends: Maintain visibility into quality metrics over time (failure rates, ticket turnaround, regression rates) and report on trends to leadership.
What You Bring
  • 3+ years of QA experience, ideally with some exposure to AI/LLM-powered products, or strong analytical QA experience with a fast ability to ramp up on agentic systems.
  • Strong pattern-recognition skills — comfortable reviewing large volumes of unstructured session data and spotting what matters.
  • Experience writing clear, actionable engineering tickets (bug reports, reproduction steps, priority/severity judgment).
  • Understanding of agentic AI concepts — planning loops, memory, tool use, multi-agent coordination — enough to judge whether an agent behaved correctly.
  • Familiarity with LLM APIs (Anthropic Claude, OpenAI, Gemini, or open-source models) and how prompt/tool design affects output quality.
  • Process-oriented mindset — genuinely enjoys finding inefficiencies and proposing fixes, not just executing a checklist.
  • Comfort working with (and improving) internal eval and workflow tooling, not just executing tests against a fixed spec.
  • Strong cross-functional communication skills — able to work fluidly with engineers, Customer Success, and Internal QA to get a full picture of issues and drive them to resolution.
  • Working knowledge of React.js, Node.js, and/or Python is a plus for understanding the systems under review, though this is not primarily a coding role.
  • Bachelor's degree in Computer Science or related field, or equivalent experience.
Bonus Points
  • Direct experience QA-ing agentic workflows, RAG pipelines, or MCP-based tool integrations.
  • Familiarity with AI observability/eval tools (LangSmith, Arize, Helicone).
  • Experience building or improving QA processes or internal eval frameworks from scratch.
  • Exposure to AWS Lambda or other AWS services used in AI deployments.
  • AWS Certification.

Skills: AI Session Review · Pattern Analysis · AI Eval & Smartflows Improvement · QA Process Design · Bug Triage & Ticketing · Agentic AI · LLM Evaluation · Cross-functional Collaboration

Annual Salary: ₹9.0 lakhs – ₹15.0 lakhs

More Info

Key Skills

Internal eval frameworks

AI Session Review

QA Process Design

LLM Evaluation

Agentic AI

Cross-functional Collaboration

Bug Triage Ticketing

AWS Certification

AI Eval Smartflows Improvement

About Company