Haystack
← Back to Jobs
Technology

Cloud infrastructure (ideally AWS) with AI

KeylentUnited States🇺🇸United StatesPosted 17 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Cloud infrastructure (ideally AWS) with AI

Requirement 
Proficiency building and operating on cloud infrastructure (ideally AWS): containerized services (ECS/EKS), serverless (Lambda), data services (S3, DynamoDB, Redshift), orchestration (Step Functions), model serving (SageMaker), and infra-as-code (Terraform/CloudFormation). 
Significant software development experience in one or more languages (Python, C/C++, Go, Java); strong hands-on large-scale Python experience preferred. 
Three or more years of designing, architecting, testing, and launching production ML systems (deployment/serving, evaluation, monitoring, data pipelines, fine-tuning workflows). 
Practical LLM experience: API integration, prompt engineering, fine-tuning/adaptation, RAG, and tool-using agents (vector retrieval, function calling, secure tool execution). 
Understanding of commercial and open-source LLMs and their capabilities (e.g., OpenAI, Gemini, Llama, Qwen, Claude). 

Focus Area 
Description 
Build agentic AI systems 
Design and implement tool-calling agents combining retrieval, structured reasoning, and secure action execution (function calling, change orchestration, policy enforcement) following the MCP protocol; engineer guardrails for safety, compliance, and least-privilege access. 
Productionize LLMs 
Build evaluation frameworks for open-source and foundational LLMs; implement retrieval pipelines, prompt synthesis, response validation, and self-correction loops for production operations. 
Integrate with runtime ecosystems 
Connect agents to observability, incident management, and deployment systems for automated diagnostics, runbook execution, remediation, and post-incident summarization with full traceability. 
Collaborate with users 
Partner with production and application teams to translate pain points into agentic AI roadmaps; define objective functions linked to reliability, risk reduction, and cost. 
Enhance governance 
Build validator models, adversarial prompts, and policy checks; enforce deterministic fallbacks, circuit breakers, and rollback strategies; instrument continuous evaluations. 
Improve scale & performance 
Optimize cost and latency via prompt engineering, context management, caching, model routing, and distillation; leverage batching, streaming, and parallel tool-calls to meet SLOs.

Skills

DynamoDB
AWS
CloudFormation
C++
Java
LLM
Python
Redshift
Terraform

Similar jobs