Cloud infrastructure (ideally AWS) with AI
Quick Overview
Job Description
Cloud infrastructure (ideally AWS) with AI
Requirement
Proficiency building and operating on cloud infrastructure (ideally AWS): containerized services (ECS/EKS), serverless (Lambda), data services (S3, DynamoDB, Redshift), orchestration (Step Functions), model serving (SageMaker), and infra-as-code (Terraform/CloudFormation).
Significant software development experience in one or more languages (Python, C/C++, Go, Java); strong hands-on large-scale Python experience preferred.
Three or more years of designing, architecting, testing, and launching production ML systems (deployment/serving, evaluation, monitoring, data pipelines, fine-tuning workflows).
Practical LLM experience: API integration, prompt engineering, fine-tuning/adaptation, RAG, and tool-using agents (vector retrieval, function calling, secure tool execution).
Understanding of commercial and open-source LLMs and their capabilities (e.g., OpenAI, Gemini, Llama, Qwen, Claude).
Focus Area
Description
Build agentic AI systems
Design and implement tool-calling agents combining retrieval, structured reasoning, and secure action execution (function calling, change orchestration, policy enforcement) following the MCP protocol; engineer guardrails for safety, compliance, and least-privilege access.
Productionize LLMs
Build evaluation frameworks for open-source and foundational LLMs; implement retrieval pipelines, prompt synthesis, response validation, and self-correction loops for production operations.
Integrate with runtime ecosystems
Connect agents to observability, incident management, and deployment systems for automated diagnostics, runbook execution, remediation, and post-incident summarization with full traceability.
Collaborate with users
Partner with production and application teams to translate pain points into agentic AI roadmaps; define objective functions linked to reliability, risk reduction, and cost.
Enhance governance
Build validator models, adversarial prompts, and policy checks; enforce deterministic fallbacks, circuit breakers, and rollback strategies; instrument continuous evaluations.
Improve scale & performance
Optimize cost and latency via prompt engineering, context management, caching, model routing, and distillation; leverage batching, streaming, and parallel tool-calls to meet SLOs.
Skills
Similar jobs
Insider Threat Analyst
Motion Recruitment Partners, LLC · Washington, United States
1 minute agoFirmware Engineer 5
Randstad Digital · Redmond, United States
1 minute ago$70 - $80/hrSr. Developer - UI
Motion Recruitment Partners, LLC · Atlanta, United States
1 minute agoSoftware Engineer / Typescript, Java, Robotics / Stoneham, MA
Motion Recruitment Partners, LLC · Boston, United States
1 minute agoSalesforce Development Technical Manager -PERM FTE
Robert Half · Cedar Rapids, United States
1 minute ago$160k - $185k/yrSoftware Engineer -Delphi & Firebird
Capgemini America, Inc. · United States
1 minute ago$53.6k - $122.4k/yr