Haystack
← Back to Jobs
Remote
Other

AI Automation Tester / Lead (100% Remote) Only Citizen & (Heathcare or Insurance Exp)

Sapience, IncUnited States🇺🇸United StatesPosted 22 Jul 2026

Quick Overview

Work Type
Remote
Level
Mid Senior

Job Description

We are seeking a talented AI Quality & Automation Engineer to join our AI Engineering team and help ensure the quality, reliability, and safety of next-generation Generative AI, Conversational AI, RAG (Retrieval-Augmented Generation), and AI Agent solutions. The ideal candidate will have a strong background in QA automation, hands-on experience with Playwright/Selenium, API testing, and a solid understanding of LLM evaluation, AI quality benchmarking, prompt engineering, and RAG validation.

This role involves building automated test frameworks, validating AI responses, ensuring production readiness through AI safety and observability, and collaborating with engineering, DevOps, and QA teams to deliver high-quality AI-powered applications.

Duration: 6-12+ Months

Key Responsibilities

  • Test and validate Generative AI, Conversational AI, and AI Agent applications.
  • Perform LLM evaluation, benchmarking, and response quality validation.
  • Validate RAG systems for retrieval accuracy, relevance, grounding, and hallucination detection.
  • Develop and maintain automation frameworks using Playwright or Selenium.
  • Execute API, regression, integration, and end-to-end testing.
  • Perform AI safety testing, red teaming, bias validation, and Human-in-the-Loop (HITL) testing.
  • Support AI observability, monitoring, and quality metrics.
  • Collaborate with Engineering, QA, DevOps, and Product teams throughout the software development lifecycle.
  • Integrate automated tests into CI/CD pipelines and improve overall test automation processes.

Required Skills

  • Strong experience in QA Automation and software testing.
  • Hands-on experience with Playwright and/or Selenium.
  • Experience with API testing using Postman and REST APIs.
  • Knowledge of Generative AI, LLM evaluation, and RAG validation.
  • Experience with prompt engineering and conversational AI testing.
  • Understanding of AI safety, observability, bias testing, and HITL workflows.
  • Proficiency in Python, SQL, Git, and CI/CD concepts.
  • Strong analytical, troubleshooting, and communication skills.

Preferred Skills

  • Experience with OpenAI, Azure OpenAI, ChatGPT, Claude, or similar LLM platforms.
  • Knowledge of LangChain, LangGraph, CrewAI, and Model Context Protocol (MCP).
  • Experience working with vector databases such as Pinecone, ChromaDB, or Weaviate.
  • Familiarity with LLM evaluation frameworks such as RAGAS and DeepEval.
  • Experience with Docker, Kubernetes, GitHub Actions, and modern DevOps practices.
  • Knowledge of Python or JavaScript for automation development.

Skills

Docker
SQL
Selenium
Azure
Generative AI
Git
GitHub Actions
JavaScript
Kubernetes
LLM
Playwright
Postman
Python
REST

Similar jobs