Haystack
← Back to Jobs
Remote
Contract
Technology
TU

Senior Software Engineer, AI Training / Evaluation

TuringUnited States🇺🇸United StatesPosted Oct 2, 2026

Quick Overview

Seniority
Mid Senior
Employment type
Contract
Work mode
Remote
Location
United States
Posted
1 hour ago
Start date
ASAP
PythonJavaScript/TypeScriptJavaC++GoC#RubyPHPRustGit

Job Description

About the Role

We are looking for senior software engineers with 7+ years of industry experience to support the training and evaluation of large language models. In this role, you will review, analyze, and improve AI-generated code, ensuring that it is correct, secure, maintainable, scalable, and aligned with production engineering standards.

This is a code quality and engineering evaluation role—not a traditional software testing or manual QA position. You will apply your software engineering judgment to assess implementation quality, identify root causes, and recommend or implement robust solutions.

Why Join Us?

Turing is one of the world’s fastest-growing AI companies, accelerating the development and deployment of advanced AI systems. You will contribute to improving AI coding capabilities by applying your real-world experience with production systems, code reviews, debugging, and software architecture.

What Does Day-to-Day Look Like?

  • Review and evaluate AI-generated code across different programming languages and software engineering scenarios.
  • Assess code for correctness, reliability, security, scalability, readability, and maintainability.
  • Identify bugs, logical errors, incomplete implementations, edge-case failures, and architectural weaknesses.
  • Analyze unfamiliar codebases and understand the impact of proposed changes.
  • Review bug fixes, feature implementations, refactoring, API integrations, configuration changes, and database operations.
  • Compare alternative implementations and determine which solution best satisfies the requirements.
  • Rewrite or improve code to create high-quality reference solutions.
  • Provide clear technical explanations for identified issues and recommended improvements.
  • Develop evaluation criteria, technical annotations, and rubrics for code-quality assessment.
  • Collaborate with researchers and engineering teams to design coding benchmarks and improve LLM evaluation methodologies.

Required Qualifications

  • 7+ years of professional software engineering experience.
  • Strong proficiency in at least one programming language, such as Python, JavaScript/TypeScript, Java, C++, Go, C#, Ruby, PHP, or Rust.
  • Significant experience building, maintaining, debugging, and reviewing production-grade software.
  • Strong understanding of software design principles, clean code, modular architecture, abstraction, error handling, and maintainability.
  • Proven ability to identify functional, performance, security, and design issues in complex codebases.
  • Strong debugging, root-cause analysis, and problem-solving skills.
  • Experience with code reviews and collaborative software development workflows.
  • Familiarity with Git and modern engineering practices.
  • Strong written English and the ability to communicate technical feedback clearly.
  • Good understanding of data structures, algorithms, APIs, databases, and application architecture.

Evaluation Process (approximately 60 mins) :

  • AI interview (20 mins approx)
  • Delivery interview (45 - 60 min)

Similar jobs