Haystack
← Back to Jobs
Technology
VI

QA Lead (Exp-10+ Years) with Agentic AI- Full Time Role

Visionary Innovative Technology SolutionsSan Diego, CA🇺🇸United StatesPosted Sep 17, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
San Diego, CA, United States
Posted
Yesterday
SQLSeleniumAgileCypressGenerative AIJavaJavaScriptLLMPlaywrightPostmanPythonRESTTypeScript

Job Description

We are seeking an experienced QA Lead with strong Agentic AI / Generative AI testing experience to lead quality engineering initiatives for enterprise applications and AI-powered solutions.

The ideal candidate will have strong experience in test strategy, QA leadership, automation, API/UI testing, CI/CD, and Agile methodologies, along with hands-on knowledge of Agentic AI, LLMs, RAG, AI agents, prompt engineering, tool/function calling, and AI evaluation.

The candidate will be responsible for defining comprehensive testing strategies for traditional software as well as AI-driven applications, ensuring functional correctness, reliability, security, performance, accuracy, and responsible behavior of AI agents.

Mandatory Skills

  • 10+ years of experience in Software Quality Assurance / Quality Engineering
  • 3+ years of experience leading QA teams or QA automation initiatives
  • Strong experience in Test Strategy, Test Planning, Test Execution, and Defect Management
  • Strong hands-on experience with UI and API automation
  • Experience with Selenium, Playwright, Cypress, or equivalent
  • Strong programming experience in Java, Python, JavaScript, or TypeScript
  • Strong experience with REST APIs, Postman, REST Assured
  • Experience with SQL and database validation
  • Strong knowledge of CI/CD and DevOps
  • Hands-on experience testing Generative AI / LLM applications
  • Strong understanding of Agentic AI and AI Agent workflows
  • Experience testing RAG-based applications
  • Experience with prompt testing and LLM evaluation
  • Experience validating AI responses for accuracy, relevance, groundedness, consistency, and hallucinations

Agentic AI Testing

  • Design QA strategies for AI agents and autonomous agent workflows.
  • Test agent planning, reasoning, decision-making, and execution workflows.
  • Validate AI agents interacting with external tools, APIs, databases, and enterprise systems.
  • Test tool/function calling and verify correct tool selection and parameters.
  • Validate multi-step and multi-agent workflows.
  • Test agent behavior across different user prompts and scenarios.
  • Validate agent state, context retention, and conversation history.
  • Test failure handling, retries, fallbacks, and recovery mechanisms.
  • Validate agent guardrails and restricted actions.
  • Test unauthorized or unexpected agent behavior.
  • Verify that agents produce appropriate responses when required information is unavailable.
  • Develop test scenarios for human-in-the-loop and autonomous workflows.
  • Validate deterministic and non-deterministic AI behavior.

LLM / Generative AI Testing

  • Test applications powered by LLMs and Generative AI.
  • Validate prompt-response behavior across positive, negative, edge, and adversarial scenarios.
  • Evaluate LLM responses for:
    • Accuracy
    • Relevance
    • Completeness
    • Consistency
    • Groundedness
    • Hallucinations
    • Toxicity
    • Bias
    • Safety
  • Perform prompt regression testing.
  • Create reusable prompt test suites and evaluation datasets.
  • Validate structured JSON and schema-based LLM outputs.
  • Test model responses across different models and model versions.
  • Perform LLM regression testing following model, prompt, or knowledge-base changes.
  • Validate token usage, latency, response quality, and error handling.

RAG Testing

  • Test Retrieval-Augmented Generation (RAG) pipelines.
  • Validate document ingestion and processing.
  • Test chunking, embeddings, indexing, and retrieval.
  • Validate retrieved context against expected source documents.
  • Test retrieval relevance and ranking.
  • Identify incorrect or missing context retrieval.
  • Validate grounded responses based on retrieved information.
  • Test hallucination and unsupported-answer scenarios.
  • Validate citations and source references.
  • Test RAG applications using vector databases.

Similar jobs