Haystack
← Back to Jobs
Remote
Technology
PA

AI Engineer - Agentic AI (Healthcare FDE) (No C2C - Fully Remote!!)

Palni IncUnited States🇺🇸United StatesPosted 3 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
19 hours ago
FastAPILLMPostgreSQLPython

Job Description

Role: Healthcare AI/ML Engineer – Agentic AI, Agentic Code

Location: USA (Remote)

Duration: 6+ months

Visa: Permanent Residents only!!

 Job Description:

 What we need you to bring

  • Production agentic systems, not demos. You have shipped multi-step agents that ran against real traffic and real failure modes. LangGraph, LangChain, DSPy, Temporal-backed orchestration, or something you built yourself. What we look for is whether you can explain why the graph is shaped the way it is and what happens when step four returns garbage.
  • Strong production Python. FastAPI, async patterns, PostgreSQL, pgvector or ChromaDB, clean typed interfaces, real test discipline. Comfort with at least one strongly-typed language is a plus. Code quality matters more than framework familiarity.
  • Retrieval engineering depth. You have built and debugged RAG in production over messy, structured, and semi-structured source material. You know why naive chunking destroys citation traceability, you have measured retrieval quality independently of answer quality, and you have opinions about reranking that came from data.
  • Evaluation as a first-class discipline. You have built eval sets rather than eyeballing outputs. Offline and online evaluation, human agreement measurement, regression gating, and honest calibration of where automated judges break down. If you have run evals against expert-labeled ground truth, lead with that.
  • Prompt and context engineering at production scale. Structured output, tool and function-call design, context budgeting, failure containment. Treated as engineering with tests and versions, not as copywriting.
  • Healthcare regulatory literacy. CPT, ICD-10, and HCPCS fluency. NCD and LCD structure. Medical necessity criteria and how InterQual or MCG-equivalent criteria sets are actually applied. CMS-0057-F at minimum. PHI handling discipline that is instinct rather than a checklist. This is a hard requirement for this seat, not a preference.
  • Claude as coding partner, fluently. Spec-first prompting, agent-driven refactors, code review by Claude as muscle memory, Claude-assisted eval and test generation, ADR drafting. We expect Claude visible in your day-to-day, not as a toy.
  • Customer-facing maturity. You can sit with a chief medical officer, a UM nurse, and a payer's head of engineering in the same room and lead a productive conversation about how an agent will fit their workflow without losing any of them.
  • 5+ years of production engineering experience, with at least 2 building LLM or ML systems that ran in production. Healthcare or regulated-systems depth strongly preferred. Exceptional candidates slightly under the bar with unusually strong production

Roles & Responsibilities:

  • Agent graph design on Aether One™. Multi-agent decomposition of clinical decision workflows. State machines, not conversation loops. Tool design, control flow, retry and fallback semantics, partial-failure behavior, and typed contracts between agents. LangGraph is our common shape; the reasoning behind the graph matters more than the framework.
  • Retrieval over clinical and regulatory knowledge. RAG across NCDs, LCDs, payer medical policy, formulary criteria, and coding references (CPT, ICD-10, HCPCS, LOINC, RxNorm, SNOMED). Chunking and indexing strategies that preserve citation granularity, because a decision has to point at the exact criterion it turned on. Hybrid retrieval, reranking, and retrieval evaluated separately from generation.
  • Per-criterion citation chains. Every determination traces to the specific policy language that produced it. This is an architectural requirement, not a feature. If the chain breaks, the decision is not defensible, and if it is not defensible we do not ship it.
  • Evaluation harness and quality gates. Gold-standard case sets built with clinician adjudication. Regression suites wired into CI. Agreement metrics against human reviewers, drift monitoring on live traffic, and a clear-eyed view of where LLM-as-judge is useful and where it quietly lies. Eval pass rate is a number we publish, not a number we claim.
  • Deterministic guardrails and the no-auto-deny invariant. Zero auto-denials is enforced in architecture, not in a prompt and not in a config flag. You build the routing, the human-in-the-loop gates on adverse outcomes, and the assertions that fail the build when an invariant is violated.
  • Model routing and the containment boundary. Frontier reasoning is rented and swap-capable; the knowledge overlay is owned and stays customer-side. You work across that boundary: Claude and other frontier models where they fit, open-weight local substrates where sovereign or air-gapped deployment demands it, and the routing logic that decides which handles what.
  • Read paths into the Healthcare Brain. Knowledge packs, MCP servers, and Claude Skills are how agents and customer teams consume governed context instead of prompts and hope. You build and maintain those surfaces for the workflows you own.
  • Customer-facing agent engineering. You will sit with utilization-management nurses, pharmacy operations teams, medical directors, and the customer's own engineers. You will watch your agent handle their real cases, take the feedback in the room, and tune. You stay embedded through the first 90 days of production.
  • Patent-grade specification work. Genzeon Platforms files patents on the architectures we build. ADRs you author may become claim language. We will train you on this, but you have to be willing to write with that level of care.

 

Similar jobs