Quick Overview
Job Description
We are seeking a highly capable Lead AI Engineer to design, build, and scale enterprise-grade agentic AI systems that operate reliably in production environments.
This role requires strong hands-on expertise in agent orchestration, distributed systems, LLM application engineering, retrieval architectures, and production AI delivery. The ideal candidate is not only technically strong, but also able to lead implementation decisions, guide engineers, and translate evolving business problems into robust AI systems.
You will work closely with product teams, data scientists, platform engineers, and business stakeholders to architect intelligent agent workflows, memory-aware reasoning systems, tool-driven execution pipelines, and cloud-ready AI platforms.
Experience: 8+ years (Full-time)Location- Remote Key Responsibilities
Technical Leadership
Lead design and implementation of scalable agentic AI systems for enterprise use cases
Own architecture decisions across orchestration, memory, tool execution, retrieval, and system reliability
Translate ambiguous business requirements into modular technical solutions with clear execution plans
Guide engineering teams on design patterns, implementation standards, and production readiness
Drive technical reviews, design discussions, and solution trade-off decisions
Agentic AI Engineering
Build multi-agent systems involving planning, reasoning, execution, coordination, and controlled autonomy
Implement agent workflows with state handling, retries, guardrails, memory persistence, and fault recovery
Design agent communication patterns including sequential, hierarchical, and collaborative orchestration
Build tool-first execution models integrating APIs, databases, enterprise systems, and external services
Implement short-term and long-term memory patterns across agent sessions
LLM Systems & Prompt Engineering
Design production-grade LLM pipelines including prompt orchestration, function calling, structured outputs, and tool invocation
Apply advanced prompt engineering techniques including:
Zero-shot and few-shot prompting
Chain-of-thought reasoning
Reflection and iterative prompting
Prompt optimization for reliability and cost control
Build robust guardrails for hallucination reduction, response validation, and deterministic behavior
Production Engineering
Build fault-tolerant AI systems with clear success/failure handling
Ensure observability through logging, tracing, and execution monitoring
Implement evaluation pipelines for prompts, agents, and retrieval quality
Drive performance tuning for latency, cost, and scalability
Ensure enterprise compliance including security, access control, and data governance
Team Collaboration & Delivery
Partner with Data Scientists to integrate ML models into agent workflows
Work with platform teams to productionize AI systems across cloud environments
Contribute actively within Agile delivery cycles
Mentor engineers and raise technical standards across the team
Required Experience & Expertise
Professional Experience8+ years of industry experience in AI/ML and Intelligent systems development
Proven experience delivering AI or ML solutions in large-scale or enterprise environments
Strong understanding of Agentic AI architectures, including both neural-based and symbolic agents
Hands-on experience building multi-agent systems, including:
Agent collaboration and coordination
Reinforcement learning or feedback-driven optimization
Dynamic or flexible workflows
State, caching, and memory management
Experience with one or more agentic AI frameworks, such as:
LangGraph / LangChain
CrewAI
Semantic Kernel
AutoGen or equivalent frameworks
Programming & ML
Strong proficiency in Python for building scalable, production-grade systems
Experienced or foundational knowledge in machine learning frameworks such as TensorFlow, PyTorch, Scikit-learn, or AutoML tools
Solid understanding of model lifecycle management, including training, evaluation, and deployment
Prompt Engineering & LLMs
Practical experience with prompt engineering techniques, including:
Zero-shot and few-shot prompting
Chain-of-thought and structured reasoning
Prompt iteration and optimization
Experience building LLM-based applications, including tool use and function calling
IR / RAG & Knowledge Systems
Experience designing and implementing Information Retrieval (IR) and RAG systems
Hands-on work with vector databases, embeddings, and optionally knowledge graphs
Familiarity with hybrid search approaches (vector + lexical + metadata-based retrieval)
Model Evaluation
Experience evaluating AI systems using quantitative and qualitative metrics
Familiarity with A/B testing, benchmarking, and performance analysis of LLMs and prompts
Technical Skills
Programming Languages: Python (required)
Agentic AI: LangGraph, LangChain, CrewAI, Semantic Kernel, AutoGen, OpenAI Agent SDK, or similar
Generative AI: LLMs, RAG architectures, NLP pipelines
Cloud Platforms: Experience with at least one major cloud provider (e.g., Google Cloud Platform, Azure, AWS); ability to design cloud-agnostic architectures
Version Control: Git / GitHub
Development Practices: Model testing, validation, CI/CD awareness
Collaboration: Experience working in Agile / Scrum teams
What Success Looks Like
The candidate designs and mentors the team to implement enterprise-grade agentic AI solutions with minimal supervision and zero hand-holding.
Translates ambiguous or loosely defined business requirements into well-architected, scalable, and production-ready agentic systems.
Delivers solutions that adhere to enterprise IT, security, and compliance standards, including data governance and access controls.
Builds fault-tolerant, resilient agentic software, with clear handling of both success and failure scenarios.
Implements comprehensive testing strategies, covering positive paths, edge cases, and failure modes, incorporating explicit business validation inputs.
Proactively collaborates with cross-functional team members to promote shared learning, technical excellence, and best practices.
Acts as a reliable team contributor during high-pressure situations, including production incidents or critical system failures, supporting root-cause analysis and rapid recovery.
Demonstrates ownership, accountability, and a production-first mindset throughout the lifecycle of agentic AI solutions.
Similar jobs
- IN
ServiceNow AI Developer - InDev with Security Clearance
NewInDev
Ashburn, VA🇺🇸On-site20 hours agoAgileCSSHTML+2Technology - PC
Forward Deployed AI Engineer
NewParkar Consulting Group, LLC
Westchester, IL🇺🇸On-site20 hours agoAzurePower BITechnology - 3B
Senior AI Engineer
New3S Business Corporation Inc.
San Jose, CA🇺🇸On-site20 hours agoLLMTechnology - CS
AI Engineer
NewCynet Systems
Redmond, WA🇺🇸Hybrid20 hours agoAWSMLOpsNLP+10Technology - RT
Artificial Intelligence AI LLM Engineer
NewRequest Technology, LLC
Chicago, IL🇺🇸HybridYesterdayDockerSQLAWS+6 - AI
AI Native Engineer
NewARK Infotech Spectrum
New York, NY🇺🇸Hybrid20 hours agoMicroservicesAWSAzure+4