Haystack
← Back to Jobs
Technology
XO

GenAI Solutions Architect / Sr AI Engineer

Xoriant CorporationNew York, NY🇺🇸United StatesPosted Sep 28, 2026

Why This Role Stands Out

This role offers an exciting opportunity to architect cutting-edge GenAI solutions, leveraging your expertise in LLMs and RAG to drive innovation. You'll thrive here if you possess strong software engineering skills and a passion for building scalable AI systems, so consider applying to join Xoriant Corporation's dynamic team.

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
New York, NY, United States
Posted
2 days ago
LLMPython

Job Description

GenAI Solutions Architect / Sr AI Engineer

NYC Location

In-person interview - Must have

Role Overview

The ideal candidate combines strong software engineering skills with deep understanding of LLMs, RAG, agents, context management, evaluation, and scalability.


Key Responsibilities

Solution Design & Architecture

·                  Design end-to-end AI, RAG, and agentic solutions for enterprise use cases.

·                  Evaluate architectural trade-offs and select appropriate patterns, models, and platforms.

·                  Create and defend Architecture Decision Records (ADRs) and technical designs.

·                  Identify risks, failure modes, scalability concerns, and optimisation opportunities.

AI Engineering & Development

·                  Build production-grade AI applications using LLMs, agents, workflows, and retrieval systems.

·                  Develop and integrate tools, APIs, vector databases, and knowledge systems.

·                  Implement memory, context management, guardrails, evaluation, and observability capabilities.

·                  Leverage AI-assisted coding tools (Claude Code, Cursor, GitHub Copilot, etc.) while maintaining engineering ownership of the solution.

Production Readiness

·                  Improve consistency, reliability, and performance of AI systems.

·                  Troubleshoot issues such as hallucinations, context bloat, latency, cost overruns, and output variability.

·                  Design monitoring, testing, evaluation, and governance frameworks for production systems.

·                  Optimize inference, retrieval, caching, and overall system performance.

Collaboration

·                  Work with product, architecture, data, and platform teams to define and deliver solutions.

·                  Translate business requirements into scalable technical architectures.

·                  Contribute to engineering standards, best practices, and reusable AI assets.


Required Skills & Experience

Core AI & LLM Engineering

·                  Hands-on experience building GenAI, RAG, and agentic applications.

·                  Strong understanding of LLM architectures, prompting, model selection, and evaluation.

·                  Experience with multi-agent systems, tool calling, MCP, workflow orchestration, or similar patterns.

·                  Understanding of fine-tuning, embeddings, vector search, and retrieval architectures.

Architecture & System Thinking

·                  Ability to justify technology choices and architectural decisions.

·                  Experience designing solutions for enterprise-scale workloads and large data sets.

·                  Strong understanding of scalability, reliability, cost, performance, and maintainability trade-offs.

·                  Familiarity with Architecture Decision Records (ADR) and solution documentation.

Context & Memory Management

·                  Understanding of:

o        Context management strategies

o        Context compression and summarization

o        Short-term and long-term memory patterns

o        Retrieval optimisation

o        Token and prompt efficiency

Engineering & Coding

·                  Strong programming skills in Python and modern software engineering practices.

·                  Experience with version control, testing, CI/CD, code reviews, and SDLC processes.

·                  Ability to read, review, optimise, and troubleshoot AI-generated code.

Optimization & Production Operations

·                  Understanding of:

o        KV Cache

o        Prompt caching

o        Response caching

o        Guardrails

o        Evaluation frameworks

o        Monitoring and observability

o        Performance optimisation techniques

Similar jobs