Haystack
← Back to Jobs
Technology

AI Engineer – Generative & Agentic Workflows

Galaxy i Technologies, Inc.Sunnyvale, CA🇺🇸United StatesPosted 7 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Job: AI Engineer – Generative
& Agentic Workflows

Location: Sunnyvale,
CA/Austin, TX (Hybrid)


VISA: USC


Generative
AI Core Development
  • Model Optimization & Fine-Tuning: Fine-tune, distill, and optimize open-source and
proprietary foundation models (e.g., Llama, Claude, GPT-4, Mistral) using
techniques like PEFT/LoRA and RLHF/DPO.
  • Retrieval-Augmented Generation (RAG): Build scalable, low-latency Hybrid RAG architectures
integrating vector databases, graph databases, and semantic routing for
complex domain contexts.
  • Multimodal Pipelines:
Design and deploy multimodal pipelines (text, vision, audio, structured
data) to extract insights and generate rich artifacts.

Agentic
AI & Autonomous Systems
  • Multi-Agent Architectures: Design and deploy multi-agent orchestration frameworks
(e.g., LangGraph, AutoGen, CrewAI, Semantic Kernel) with dedicated roles,
shared memory, and cross-agent negotiation strategies.
  • Planning & Reasoning Protocols: Implement advanced reasoning paradigms—such as
Chain/Tree/Graph-of-Thought, ReAct loops, self-reflection, and
reflection-based error correction.
  • Tool Augmentation & Function Calling: Connect agents to external APIs, databases, software
environments, and browser tools, enabling reliable function calling,
schema validation, and tool execution.
  • Human-in-the-Loop (HITL) Workflows: Build human-in-the-loop safety checkpoints, approval
triggers, and oversight interfaces into autonomous agent loops.

Platform
Performance, Guardrails & MLOps
  • Agentic Observability & Evaluation: Set up rigorous evaluation frameworks (LLM-as-a-judge,
trajectory tracing, latency profiling) using tools like LangSmith,
Phoenix, or Arize.
  • Guardrails & Alignment: Implement strict safety, hallucination mitigation,
context-window optimization, and prompt injection defenses using guardrail
frameworks (e.g., NeMo Guardrails, Guardrails AI).
  • Production Deployment: Scale agent workflows on cloud infrastructure
(AWS/Google Cloud Platform/Azure) with async task queues, durable execution state, and
low-latency API integration.

Key Skills & Technologies:

Languages
Python (Expert), TypeScript / Node.js (Plus)

Gen AI Frameworks
PyTorch, Hugging Face Transformers, vLLM, Ollama,
LangChain, LlamaIndex

Agentic Frameworks
LangGraph, AutoGen, CrewAI, Semantic Kernel, Temporal

Databases & Search
Qdrant, Pinecone, Milvus, Weaviate, Pgvector, Neo4j

Infrastructure & MLOps
Docker, Kubernetes, Ray, FastAPI, LangSmith, MLflow,
AWS/Google Cloud Platform/Azure

Experience
& Qualifications
  • Experience: 6+ years of professional software engineering experience, with 2+ years
dedicated to building and deploying Gen AI and/or LLM applications in
production.
  • Proven Track Record:
Experience building stateful LLM applications, custom RAG systems, or
autonomous agentic workflows deployed to real users.
  • Strong Algorithmic Foundation: Solid understanding of transformer architectures,
attention mechanisms, vector embeddings, and non-deterministic state
machine design.
  • Problem-Solving Mindset: Comfort dealing with model non-determinism, edge cases
in tool calling, and designing robust fallback mechanisms.

Skills

Docker
FastAPI
Neo4j
Node.js
AWS
MLOps
MLflow
Azure
GPT
Google Cloud
Hugging Face
Kubernetes
LLM
Phoenix
PyTorch
Python
React
TypeScript

Similar jobs