Quick Overview
Job Description
About the Role
Join an early-stage AI infrastructure team as its first QA engineer, owning quality and test automation across a GPU fleet management platform. Your work will help ensure that conversational and web interfaces, backend services, and Kubernetes operations behave reliably and safely in production.
What You'll Do
Test conversational workflows end to end, including command accuracy and actions performed on live clusters.
Build evaluations for AI agent behavior, checking response quality, safe actions, and handling of ambiguous or invalid requests.
Create end-to-end tests for the web interface and verify consistent behavior across product interfaces.
Write Go integration tests for backend workflows such as installation, upgrades, and failover.
Provision and tear down Kubernetes test clusters to validate controllers, operators, and custom resources across versions.
Verify alert delivery across product interfaces and chat, and ensure alerts fire appropriately.
Gate releases in CI and turn bugs and incidents into regression tests.
What We're Looking For
At least 3 years of experience in QA, SDET, or test automation, including ownership of product quality and test suites.
Strong Go skills for integration testing or backend test automation, plus experience building end-to-end web UI tests.
Hands-on experience with Kubernetes, React, RPC systems such as gRPC, CI/CD, and Linux fundamentals.
Familiarity with monitoring and alerting tools such as Prometheus, Alertmanager, or Grafana.
Experience testing conversational or LLM-powered products, building agent evaluations, or working with GPU infrastructure is valuable.
Exposure to multicluster tooling, chaos engineering, Kubernetes testing tools, and UI automation frameworks is a plus.
Compensation & Benefits
Annual salary: $180,000 to $200,000 USD. Visa sponsorship is not available.
Location
On-site in San Francisco, California.