Quick Overview
Job Description
NO H1S OR 3RD PARTIES.
We are seeking an experienced GenAI engineer with deep expertise in LLM evaluation, production-grade AI systems, and financial services. The successful candidate will design and improve reliable AI applications that answer complex financial questions using structured and unstructured data.
This is primarily an LLM engineering role—not a trading or quantitative research position. However, candidates must understand front-office financial workflows and the standards of accuracy, reliability, and responsiveness expected by investment professionals.
Ideal candidates may come from financial-data and AI platforms, banks, asset managers, fintech companies, or specialist firms such as Rogo, RavenPack, LSEG, FactSet, Bloomberg, S&P Global, or similar organizations.
Key responsibilities
Design comprehensive evaluation frameworks for LLM-powered financial applications, covering factual accuracy, relevance, completeness, grounding, consistency, and appropriate refusal behavior.
Build representative evaluation datasets and financial “golden sets,” including difficult, ambiguous, adversarial, and business-critical test cases.
Establish automated evaluation pipelines combining deterministic tests, model-based evaluation, human review, and expert financial validation.
Identify and reduce hallucinations, unsupported conclusions, citation errors, retrieval failures, and other sources of unreliable output.
Develop production-grade applications that answer questions using structured and unstructured financial information, including research, filings, transcripts, news, market data, and proprietary documents.
Design and optimize retrieval, RAG, search, ranking, reranking, context construction, and tool-use workflows.
Build or integrate MCP servers and other interfaces that allow models to query enterprise data, research systems, analytical tools, and external services securely.
Evaluate the complete application rather than the model in isolation, including retrieval quality, tool selection, tool execution, context quality, answer generation, citations, and end-to-end task success.
Implement monitoring and observability for production AI systems, including quality degradation, data drift, latency, cost, failure modes, and user feedback.
Apply fine-tuning, supervised learning, preference optimization, synthetic-data generation, and related techniques where they provide measurable improvements.
Develop bespoke smaller language models through model distillation—for example, transferring capabilities from a large teacher model to a smaller model optimized for a particular financial domain or workflow.
Make informed decisions about when to use prompting, retrieval, tools, fine-tuning, distillation, or a combination of techniques.
Optimize models and applications for accuracy, latency, scalability, inference cost, security, and operational reliability.
Work closely with financial-domain experts, product teams, data engineers, and application engineers to translate front-office requirements into measurable AI capabilities.
Define production-readiness criteria, release gates, regression tests, and quality thresholds for new models and application changes.
Required experience
Significant hands-on experience building and operating LLM or GenAI applications in production.
Deep knowledge of LLM evaluation methods, benchmarks, test-set design, error analysis, regression testing, and continuous quality measurement.
Demonstrated ability to improve the factual accuracy and reliability of answers generated from enterprise or domain-specific information.
Strong experience working with unstructured information and document-heavy workflows.
Practical expertise in retrieval-augmented generation, semantic and hybrid search, embeddings, reranking, grounding, citations, and context management.
Experience building tool-using or agentic systems, ideally including MCP servers or comparable model-to-data and model-to-tool interfaces.
Hands-on experience with model fine-tuning, distillation, or the development of domain-specific small and large language models.
Strong understanding of the trade-offs among model quality, model size, latency, throughput, infrastructure requirements, and inference cost.
Experience implementing production monitoring, evaluation pipelines, quality controls, and safeguards for probabilistic AI systems.
Meaningful financial-services experience, with a working understanding of front-office users, terminology, data, and workflows.
Familiarity with how investment professionals consume research, interrogate financial information, evaluate evidence, and make time-sensitive decisions.
Strong software-engineering skills and the ability to build robust, testable, maintainable production systems.
Financial-domain expectations
Candidates should understand at least several of the following:
Equity, fixed-income, credit, commodities, foreign-exchange, or multi-asset workflows.
Financial research, company analysis, market intelligence, news analytics, and investment decision support.
Company filings, earnings transcripts, financial statements, estimates, corporate actions, and market data.
The distinction between facts, estimates, opinions, forecasts, and model-generated conclusions.
Data entitlements, auditability, source attribution, information security, and regulatory or compliance considerations.
The importance of timeliness, point-in-time correctness, provenance, and reproducibility in financial applications.
Direct experience as a trader or quant is not required. The essential requirement is sufficient domain fluency to understand front-office use cases and recognize when an AI-generated answer is incomplete, misleading, unsupported, or financially implausible.
Core competencies
LLM evaluation: Can define what “good” means, measure it rigorously, and create repeatable systems for improving quality.
Production AI engineering: Can take an LLM application from prototype to a reliable, observable, scalable production service.
Financial-domain fluency: Understands front-office users, financial information, relevant workflows, and the consequences of inaccurate output.
Grounded answer generation: Knows how to produce accurate, traceable responses supported by authoritative sources.
Retrieval and data integration: Can connect models effectively to structured data, unstructured content, search systems, tools, and enterprise platforms.
Model optimization: Understands fine-tuning, distillation, model selection, and the construction of specialized models for defined tasks.
Analytical problem-solving: Can diagnose whether failures originate in the underlying data, retrieval, tool use, context construction, prompting, model behavior, or application logic.
Risk and quality mindset: Anticipates edge cases and designs controls appropriate for high-stakes financial applications.
Cross-functional collaboration: Communicates effectively with financial experts, engineers, researchers, product leaders, and senior stakeholders.
Outcome orientation: Focuses on measurable user and business outcomes rather than model demonstrations alone.
Preferred experience
Experience at a financial-data provider, investment-research platform, financial AI company, bank, asset manager, hedge fund, or capital-markets technology business.
Experience supporting research analysts, portfolio managers, traders, salespeople, or other front-office professionals.
Experience building domain-specific models or AI applications for financial research and decision support.
Experience with SaaS products, multi-tenant platforms, enterprise deployments, or customer-facing AI applications.
Familiarity with model governance, access controls, data privacy, audit trails, and regulated production environments.
Experience evaluating both proprietary and open-weight models and selecting the appropriate model for a given task.
What success looks like
The successful candidate will build AI systems that financial professionals can use with confidence. Their work will result in measurable improvements in answer accuracy, grounding, coverage, latency, cost, and reliability. They will also establish a disciplined evaluation process that makes quality visible, catches regressions before release, and supports the development of specialized models for high-value financial workflows.
Similar jobs
- AN
Senior Detection and Response Engineer
NewAnyscale
San Francisco🇺🇸8 hours agoAWSMachine LearningAzure+2Engineering - AN
Senior Mission Engineer
NewAntares
Los Angeles🇺🇸8 hours agoAgileBusiness DevelopmentCompliance+4Engineering - AN
Mission Engineer II
NewAntares
Los Angeles🇺🇸8 hours agoAgileBusiness DevelopmentCompliance+4Engineering - AM
Design Engineer, Growth
NewAmbrook
New York🇺🇸Remote2 hours agoFirestoreGCPNext.js+13Engineering - AL
Forward-Deployed Engineer
NewAlexai
San Francisco🇺🇸11 hours agoApplicant Tracking SystemsOnboardingEngineering - DT
Weapon System Mechanical Engineer with Security Clearance
NewDecision Technologies Inc
Arlington, VA🇺🇸$90k - $120k/yrOn-siteYesterdayEngineering