Haystack
← Back to Jobs
Remote
Technology
VI

QA Automation Engineer (Exp-10+ Years) with AI Exp-Full time-Remote

Visionary Innovative Technology SolutionsUnited States🇺🇸United StatesPosted 9 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
Yesterday
SQLAWSScrumSeleniumAgileAzureBDDCucumberCypressGenerative AIGitGitHub ActionsGitLab CIHugging FaceJavaJavaScriptJenkinsJiraLLMPlaywrightPostmanPythonRESTTypeScript

Job Description

We are looking for a highly experienced QA Automation Engineer with 10+ years of experience in software quality assurance, test automation, and AI/GenAI testing. The ideal candidate will have strong expertise in automation frameworks, API/UI testing, BDD, CI/CD, and modern AI/LLM application testing.

The candidate will be responsible for designing scalable automation solutions and validating AI-powered applications, LLM-based features, RAG pipelines, AI agents, and GenAI outputs for accuracy, reliability, performance, security, and functional correctness.

Mandatory Skills

  • 10+ years of experience in QA Automation / Software Testing
  • Strong hands-on experience with Cypress / Playwright / Selenium
  • Strong programming experience in JavaScript / TypeScript / Java / Python
  • Strong experience with BDD, Gherkin, and Cucumber
  • UI, API, integration, regression, and end-to-end automation
  • Strong experience with REST APIs and JSON
  • API automation using Postman / REST Assured
  • Experience with AI/GenAI/LLM application testing
  • Understanding of LLMs, RAG, Prompt Engineering, and AI Agents
  • Experience validating AI-generated responses and outputs
  • Strong SQL/database testing experience
  • Git and CI/CD experience
  • Jenkins / GitHub Actions / Azure DevOps / GitLab CI
  • Strong Agile/Scrum experience

Key Responsibilities

  • Design, develop, and maintain robust UI and API automation frameworks.
  • Create automated test scenarios using Cypress, Playwright, Selenium, or equivalent tools.
  • Develop BDD test cases using Gherkin and Cucumber.
  • Perform functional, regression, integration, system, and end-to-end testing.
  • Design automation strategies for complex enterprise applications.
  • Test and validate AI/GenAI-powered applications and features.
  • Develop test scenarios for LLM-based applications, RAG systems, and AI agents.
  • Validate AI responses for accuracy, relevance, consistency, completeness, and groundedness.
  • Test prompts, prompt variations, system instructions, and AI workflows.
  • Identify and validate LLM hallucinations, incorrect responses, bias, toxicity, and unexpected outputs.
  • Validate AI model responses against expected business rules and reference data.
  • Automate repetitive AI/LLM validation scenarios.
  • Test AI applications across different models, prompts, contexts, and datasets.
  • Validate RAG retrieval quality, context relevance, citations, and response grounding.
  • Test AI agent workflows, tool calling, function calling, and multi-step execution.
  • Integrate automated tests into CI/CD pipelines.
  • Analyze automation failures, application defects, logs, and test results.
  • Collaborate with developers, product owners, data scientists, and AI/ML engineers.
  • Participate in code reviews and contribute to automation framework improvements.
  • Track defects using JIRA or similar defect-management tools.

AI / GenAI Testing

Hands-on or strong working knowledge of:

  • Generative AI / GenAI
  • Large Language Models (LLMs)
  • RAG – Retrieval-Augmented Generation
  • Prompt Engineering
  • Prompt Testing
  • AI Agents / Agentic AI
  • AI chatbot testing
  • LLM response validation
  • Hallucination detection
  • Response accuracy and relevance testing
  • Context/groundedness validation
  • Embeddings and vector databases
  • Semantic similarity testing
  • AI model evaluation
  • Toxicity and bias testing
  • Guardrails validation
  • Function/tool calling validation
  • Multi-turn conversational testing
  • AI output consistency testing
  • Model comparison and regression testing

AI Tools / Platforms

Experience with one or more:

  • OpenAI / ChatGPT
  • Azure OpenAI
  • AWS Bedrock
  • Google Gemini
  • Anthropic Claude
  • GitHub Copilot
  • LangChain
  • LangGraph
  • Hugging Face
  • AI evaluation frameworks
  • Vector databases such as Pinecone, Azure AI Search, or similar

Automation & Programming

  • Cypress
  • Playwright
  • Selenium WebDriver
  • JavaScript / TypeScript

Similar jobs