← Back to Jobs
Technology
Sr. Application Engineer/Lead/Architect
Intake IT SolutionsSanta Clara, CA🇺🇸United StatesPosted 6 Aug 2026
Quick Overview
Work Type
Hybrid
Level
Mid Senior
Job Description
Job Title: AI – Sr. Application Engineer/Lead/Architect
Location: Bay area, CA
Duration: FTE
- Will work on the intelligence layer for multiple programs — owns all model quality, RAG accuracy, prompt engineering, and AI safety across applications
- Socratic tutor persona, adaptive learning recommendation engine, multi-modal AI (text and voice), RAG evaluation framework, and feedback loop into retrieval
- 6-LLM call chain orchestration (NeMoGuardrails → intent classification → query rewriting → RAG → synthesis), , and compatibility check logic
- Production-grade AI quality from launch — this is not a research or prototyping role; accuracy thresholds, latency requirements, and safety guardrails must pass InfoSec adversarial testing before Release 1
Required Skills
Experience
- Total IT 10+ Years
- 4–7 years of software engineering with at least 2 years focused on LLM application development in production — not research, not demos, not internal tools with 10 users
- Has shipped an LLM-powered feature or product to production where real users depend on the accuracy and the engineer owns the quality metrics
- Has owned an AI safety or guardrails implementation for a customer-facing product — not just added an off-the-shelf filter; designed and tested the safety layer
- Has built RAG evaluation pipelines and used them to make go/no-go release decisions — accuracy gating is part of the workflow.
- Has profiled and optimized a multi-step LLM call chain for latency
LLM Application Development
- LLM prompt engineering—system prompts, few-shot examples, chain-of-thought, instruction following · Expert · Must-have
- Multi-step LLM chain orchestration—LangChain, LlamaIndex, or custom orchestration · Expert · Must-have
- Multi-turn conversation design—context window management, conversation summarization, session memory · Advanced · Must-have
- Streaming LLM response handling—token-by-token streaming, partial response rendering · Advanced · Must-have
- Model selection and benchmarking—matching model size to task; balancing latency, cost, and accuracy · Advanced · Must-have
RAG Pipeline Design & Quality
- RAG pipeline design—chunking strategy, embedding model selection, retrieval configuration · Expert · Must-have
- Vector similarity search tuning—index parameters, similarity thresholds, retrieval depth · Advanced · Must-have
- Reranking—cross-encoder rerankers, relevance scoring · Advanced · Must-have
- RAG evaluation frameworks—RAGAS, TruLens, or equivalent; automated eval pipelines · Advanced · Must-have
- Hybrid search — combining dense vector retrieval with BM25 or keyword search · Proficient · Nice to have
AI Safety & Guardrails
- Prompt injection detection and mitigation · Advanced · Must-have
- Jailbreak testing and red-teaming LLM systems · Advanced · Must-have
- Content safety classifier integration · Advanced · Must-have
- Hallucination detection and mitigation strategies · Advanced · Must-have
- Topical control—enforcing scope boundaries on LLM responses · Advanced · Must-have
Evaluation & Production Quality
- Automated evaluation pipeline design—test set curation, metric selection, regression detection · Advanced · Must-have
- A/B evaluation methodology for prompt and model changes · Proficient · Must-have
- Latency profiling for LLM call chains—identifying bottlenecks across multi-step pipelines · Proficient · Must-have
- Feedback loop design—user signal collection, signal-to-retrieval-weight integration · Proficient · Must-have
- Production model monitoring—accuracy drift detection, quality degradation alerting · Proficient · Must-have
Development
- Python—ML/AI application development, async programming · Expert · Must-have
- API design for AI services—streaming endpoints, error handling, timeout management · Advanced · Must-have
- Embedding model operations—model selection, batch embedding, index updates · Advanced · Must-have
Nice to Have
- Adaptive learning systems or personalization engine experience
- Knowledge graph integration with RAG
- Multi-agent orchestration patterns
- ServiceNow API integration
- Prior experience building AI products on NVIDIA infrastructure
Skills
LLM
Python
Similar jobs
Software Engineer, Fleet Infrastructure with Security Clearance
Anduril Industries · Boston, United States
1 minute ago$166k - $220k/yrFlight Test Instrumentation Engineer with Security Clearance
Anduril Industries · Costa Mesa, United States
20 minutes ago$146k - $194k/yrFlight Software Engineer, Embedded C/C++, Air Dominance & Strike with Security Clearance
Anduril Industries · Costa Mesa, United States
20 minutes ago$129k - $220k/yrLead Software Engineer, Front End
Capital One · McLean, United States
21 minutes ago$197.3k - $225.1k/yrLead Software Engineer, Back End (Java, JavaScript, Python, Go, Spring, Kafka, CI/CD, AWS, Claude Code)
Capital One · McLean, United States
31 minutes ago$197.3k - $225.1k/yrSoftware Engineer, AI Satellites (Starmind)
SpaceX · Bastrop, United States
31 minutes ago