Haystack
← Back to Jobs
Remote
Full time
Technology
FI

Software Engineer, Agents (Internal Audit)

FieldguideSan Francisco🇺🇸United StatesPosted Sep 17, 2026

Quick Overview

Seniority
Mid Senior
Employment type
Full Time
Work mode
Remote
Location
San Francisco, United States
Posted
7 hours ago
GraphQLLLMPhoenixPythonReactTypeScript

Job Description

About Us

Fieldguide is establishing a new state of trust for global commerce and capital markets by automating and streamlining the work of assurance and audit practitioners — specifically in cybersecurity, privacy, and financial audits. We build software for the people who enable trust between businesses.

We're based in San Francisco, CA, and we're backed by Goldman Sachs Alternatives, Bessemer Venture Partners, 8VC, Floodgate, Y Combinator, and more. Over 50 of the top 100 accounting and consulting firms trust Fieldguide to power mission-critical work.

About the Role

You'll join a genuine 0→1 team on the ground floor of one of the company's biggest new bets. This seat is specifically product-focused: you'll own agent quality, ship agents that do real audit work, and work alongside practitioners.

Depending on your experience and what you're looking to own, you may join building within a major agent area, owning one end-to-end, or setting technical direction for agentic audit work across the team. We're hiring across all levels and will calibrate during interviews based on scope and demonstrated experience.

What You'll Do

  • Make agent judgment repeatable: run error analysis on real testing data and turn findings into concrete fixes

  • Tradeoffs such as quality/latency/cost across a long multi-phase run

  • Build structured-output pipelines that turn model output into real audit artifacts

  • Take ambiguous problem statements and turn them into a plan, a shipped feature, and a clear read on what was cut and why

  • Work directly with an embedded subject matter expert and with design-partner firms, turning their feedback into agent changes within days

  • Expand agent coverage into new controls and new areas of internal audit

Who You Are (All Levels)

  • Product-minded and full-stack: you've shipped LLM-backed features to production against real users, and you measure yourself on whether they got used

  • You're fluent in evals and error analysis, and you apply them in service of shipping something practitioners trust

  • You have real opinions on model selection, prompting, and orchestration tradeoffs, and you can defend them with evidence rather than vibes

  • Energized by 0→1 work: you'd rather define the problem than inherit a spec, and you don't stall on ambiguity

  • Strong instincts for human-in-the-loop design

  • A genuine team player across the organization, not just within engineering: you'll work daily with PM, design, domain experts, and customer-facing teams, and you treat that as the best part of the job

  • Ship fast without leaving a mess: your code is reviewable, tested where it counts, and instrumented

  • Able to internalize a hard domain fast. You don't need to know SOX today, but you'll understand it well enough to make the right product calls

Higher-Level Responsibilities

At the Senior level, you may:

  • Own a major agent area end-to-end, from how the agent reasons about a class of controls through to the artifact a reviewer signs

  • Set the evals and error-analysis practice for the team's agent work, and decide what evidence justifies shipping a change or rolling it back

  • Collaborate with PMs and designers to shape roadmaps and define architectural tradeoffs, including where the agent acts and where the auditor decides

  • Own the harder model and orchestration judgment calls across a long multi-phase run

  • Mentor other engineers and raise the bar on 0→1 execution and applied eval rigor

At the Staff level, you may:

  • Drive agent initiatives that reach beyond Internal Audit and influence how agents are built across Fieldguide

  • Set and champion engineering standards for agent reliability, reproducibility, and defensibility

  • Partner with engineering and product leadership to define long-term technical strategy for agentic audit work

  • Serve as a trusted advisor to leaders across Engineering, Product, and Design

  • Represent Fieldguide externally through writing, speaking, and open-source contributions

Experience

Must-have:

  • Shipped LLM-backed product features to production against real users

  • Applied AI skillset: evals, error analysis, and model-selection decisions you owned and can explain

  • Comfortable full-stack, with enough backend depth to work in agent orchestration

  • Autonomy working from an ambiguous spec

  • A collaborative mode that works across PM, design, and domain experts

Nice-to-have:

  • Python, TypeScript, React, Postgres, Hasura, GraphQL

  • Temporal or comparable durable-execution / workflow orchestration

  • Hands-on eval experience (Langfuse, Braintrust, LangSmith, Arize Phoenix, or comparable)

  • Structured-output work including schema contracts, generating real artifacts from model output

  • Startup experience, as a founder or as an early engineer

  • A 0→1 track record: things you started where no scaffolding existed

  • Experience working directly with customers, and comfort being in the room when they use what you built

  • Background in internal audit, SOX, accounting, or another regulated domain

  • Document processing, including PDF and Excel manipulation and annotation

Not a fit if:

  • Prompt engineering is your whole skill set

  • Your agent work never carried production traffic

  • You want to own eval methodology or the evaluation harness itself rather than ship product features (better fit on Foundation Agents)

  • You need a fully specified ticket to start

  • You'd rather not be in the room with customers and domain experts

What Should Excite You

  • 0→1 on the biggest bet: You're building the agent and the product from scratch, on the ground floor of where the company is going

  • Repeatable judgment: Making an agent reach the same defensible conclusion twice, in a domain where ground truth requires expert judgment

  • Real audit stakes: Your work directly affects what firms put in front of their clients, and what a reviewer is willing to sign

  • Customer proximity: Design-partner firms and an embedded SOX expert use what you ship within days of it landing

  • Human-in-the-loop design: Deciding where the agent acts and where the auditor decides, on work that genuinely matters

  • High trust, high autonomy: You're given ambiguous problems and trusted to define the plan

Benefits

  • Competitive compensation with equity

  • Comprehensive health and wellness benefits

  • Flexible time off and work schedules

  • Technology reimbursements

  • 401(k) plan

  • Twice-yearly in-person offsites across the U.S.

  • Wellness benefits starting on your first day

Our Values

  • Fearless — Inspire and break down seemingly impossible walls

  • Fast — Launch fast with excellence; iterate to perfection

  • Lovable — Deliver happiness and 11-star experiences

  • Owners — Execute and run the business with ownership

  • Win-win — Create mutual value and earn trust for life

  • Inclusive — Scale the best ideas with inclusive teams

Similar jobs