QA Engineer
Job Description
EXL (NASDAQ: EXLS) is a leading data analytics and digital operations and solutions company. We partner with clients using a data and AI-led approach to reinvent business models, drive better business outcomes and unlock growth with speed. EXL harnesses the power of data, analytics, AI, and deep industry knowledge to transform operations for the world's leading corporations in industries including insurance, healthcare, banking and financial services, media and retail, among others. EXL was founded in 1999 with the core values of innovation, collaboration, excellence, integrity and respect. We are headquartered in New York and have more than 54,000 employees spanning six continents. For more information, visit www.exlservice.com.
EXL never requires or asks for fees/payments or credit card or bank details during any phase of the recruitment or hiring process and has not authorized any agencies or partners to collect any fee or payment from prospective candidates. EXL will only extend a job offer after a candidate has gone through a formal interview process with members of EXL's Human Resources team, as well as our hiring managers.
About the Role
We are looking for an experienced Agentic AI Testing Lead to build and lead an AI Quality Engineering team responsible for validating LLM-based, multi-agent, RAG, and autonomous AI workflows. The role involves defining AI testing strategies, evaluation frameworks, automation pipelines, quality KPIs, and governance practices to ensure reliable, safe, accurate, and scalable AI solutions. About the Role We are looking for an experienced Agentic AI Testing Lead to build and lead an AI Quality Engineering team responsible for validating LLM-based, multi-agent, RAG, and autonomous AI workflows. The role involves defining AI testing strategies, evaluation frameworks, automation pipelines, quality KPIs, and governance practices to ensure reliable, safe, accurate, and scalable AI solutions.
Key Responsibilities
1. Agentic AI Testing, Evaluation & Automation • Define and execute testing strategies for LLM-based, multi-agent, RAG, and Agentic AI systems. • Validate autonomous agent behavior, reasoning, memory, tool usage, API/database integrations, and end-to-end workflows. • Evaluate AI outputs for accuracy, relevance, groundedness, consistency, completeness, toxicity, bias, hallucination risk, and guardrail compliance. • Define AI quality KPIs such as hallucination rate, groundedness score, agent success rate, task completion rate, response relevancy, latency, cost efficiency, and user satisfaction. • Build automated evaluation pipelines, quality scoring mechanisms, dashboards, and CI/CD-integrated quality gates. • Develop reusable test harnesses, simulators, and benchmarking frameworks to compare models, prompts, and agent configurations.
2. Team Leadership & Capability Building • Build and lead a team of Agentic AI Quality Engineers. • Define team structure, testing standards, best practices, and governance models. • Mentor QA engineers in AI testing methodologies, evaluation techniques, and automation frameworks. • Drive innovation and adoption of emerging AI testing tools and technologies. • Collaborate with Product, Engineering, Data Science, and AI Research teams to improve overall AI quality.
3. Reporting & Stakeholder Management • Provide quality assessments and recommendations to leadership and stakeholders. • Present testing outcomes, risk assessments, KPI trends, and model evaluation reports. • Drive quality governance for Agentic AI initiatives across the organization. • Ensure traceability of testing activities, evaluation criteria, and quality benchmarks.
Required Skills & Experience Technical Skills
• 4–8 years of experience in Software Testing, Quality Engineering, or Test Automation.
• Minimum 2+ years of hands-on experience in GenAI, LLM Testing, Agentic AI Testing, or AI Quality Engineering.
• Strong understanding of LLMs, AI agents, RAG, prompt validation, tool calling, agent memory, MCP, and multi-agent orchestration.
• Experience defining AI quality metrics, evaluation methodologies, benchmarking frameworks, and model comparison approaches.
• Hands-on automation experience with Python, Playwright, Pytest, API automation, test framework development, and CI/CD quality gates.
• Experience with AI evaluation frameworks such as DeepEval, Ragas, LangSmith, OpenAI Evals, or equivalent tools.
• Exposure to cloud platforms such as Azure, AWS, or GCP. Soft Skills • Strong communication and stakeholder management skills.
• Analytical mindset with strong problem-solving ability.
• Self-driven, outcome-oriented, and capable of leading multiple initiatives in a fast evolving AI ecosystem.
Preferred Qualifications
• Experience testing enterprise Agentic AI platforms and autonomous AI systems.
Preferred Qualifications
EXL never requires or asks for fees/payments or credit card or bank details during any phase of the recruitment or hiring process and has not authorized any agencies or partners to collect any fee or payment from prospective candidates. EXL will only extend a job offer after a candidate has gone through a formal interview process with members of EXL's Human Resources team, as well as our hiring managers.
About the Role
We are looking for an experienced Agentic AI Testing Lead to build and lead an AI Quality Engineering team responsible for validating LLM-based, multi-agent, RAG, and autonomous AI workflows. The role involves defining AI testing strategies, evaluation frameworks, automation pipelines, quality KPIs, and governance practices to ensure reliable, safe, accurate, and scalable AI solutions. About the Role We are looking for an experienced Agentic AI Testing Lead to build and lead an AI Quality Engineering team responsible for validating LLM-based, multi-agent, RAG, and autonomous AI workflows. The role involves defining AI testing strategies, evaluation frameworks, automation pipelines, quality KPIs, and governance practices to ensure reliable, safe, accurate, and scalable AI solutions.
Key Responsibilities
1. Agentic AI Testing, Evaluation & Automation • Define and execute testing strategies for LLM-based, multi-agent, RAG, and Agentic AI systems. • Validate autonomous agent behavior, reasoning, memory, tool usage, API/database integrations, and end-to-end workflows. • Evaluate AI outputs for accuracy, relevance, groundedness, consistency, completeness, toxicity, bias, hallucination risk, and guardrail compliance. • Define AI quality KPIs such as hallucination rate, groundedness score, agent success rate, task completion rate, response relevancy, latency, cost efficiency, and user satisfaction. • Build automated evaluation pipelines, quality scoring mechanisms, dashboards, and CI/CD-integrated quality gates. • Develop reusable test harnesses, simulators, and benchmarking frameworks to compare models, prompts, and agent configurations.
2. Team Leadership & Capability Building • Build and lead a team of Agentic AI Quality Engineers. • Define team structure, testing standards, best practices, and governance models. • Mentor QA engineers in AI testing methodologies, evaluation techniques, and automation frameworks. • Drive innovation and adoption of emerging AI testing tools and technologies. • Collaborate with Product, Engineering, Data Science, and AI Research teams to improve overall AI quality.
3. Reporting & Stakeholder Management • Provide quality assessments and recommendations to leadership and stakeholders. • Present testing outcomes, risk assessments, KPI trends, and model evaluation reports. • Drive quality governance for Agentic AI initiatives across the organization. • Ensure traceability of testing activities, evaluation criteria, and quality benchmarks.
Required Skills & Experience Technical Skills
• 4–8 years of experience in Software Testing, Quality Engineering, or Test Automation.
• Minimum 2+ years of hands-on experience in GenAI, LLM Testing, Agentic AI Testing, or AI Quality Engineering.
• Strong understanding of LLMs, AI agents, RAG, prompt validation, tool calling, agent memory, MCP, and multi-agent orchestration.
• Experience defining AI quality metrics, evaluation methodologies, benchmarking frameworks, and model comparison approaches.
• Hands-on automation experience with Python, Playwright, Pytest, API automation, test framework development, and CI/CD quality gates.
• Experience with AI evaluation frameworks such as DeepEval, Ragas, LangSmith, OpenAI Evals, or equivalent tools.
• Exposure to cloud platforms such as Azure, AWS, or GCP. Soft Skills • Strong communication and stakeholder management skills.
• Analytical mindset with strong problem-solving ability.
• Self-driven, outcome-oriented, and capable of leading multiple initiatives in a fast evolving AI ecosystem.
Preferred Qualifications
• Experience testing enterprise Agentic AI platforms and autonomous AI systems.
Preferred Qualifications
- Bachelor's degree/MCA in Computer Science, Information Technology, or a related field.
- Proven track record of delivering high-quality automation testing solutions.
- Strong analytical and problem-solving skills.
- Ability to work independently and as part of a team.
- Ability to multitask and flexibility with timing.
More Info
Key Skills
DeepEval
multi-agent orchestration
tool calling
CI CD quality gates
Agentic AI Testing
LangSmith
prompt validation
OpenAI Evals
Ragas
benchmarking frameworks
Playwright
LLMs
model comparison approaches
AI Quality Engineering
RAG
GenAI LLM Testing
AI agents
agent memory
AI quality metrics




