Quick Overview
Job Description
About Us
Fieldguide is establishing a new state of trust for global commerce and capital markets by automating and streamlining the work of assurance and audit practitioners — specifically in cybersecurity, privacy, and financial audits. We build software for the people who enable trust between businesses.
We're based in San Francisco, CA, and we're backed by Goldman Sachs Alternatives, Bessemer Venture Partners, 8VC, Floodgate, Y Combinator, and more. Over 50 of the top 100 accounting and consulting firms trust Fieldguide to power mission-critical work.
About the Role
You'll join a genuine 0→1 team on the ground floor of one of the company's biggest new bets. This seat is specifically product-focused: you'll own agent quality, ship agents that do real audit work, and work alongside practitioners.
Depending on your experience and what you're looking to own, you may join building within a major agent area, owning one end-to-end, or setting technical direction for agentic audit work across the team. We're hiring across all levels and will calibrate during interviews based on scope and demonstrated experience.
What You'll Do
Make agent judgment repeatable: run error analysis on real testing data and turn findings into concrete fixes
Tradeoffs such as quality/latency/cost across a long multi-phase run
Build structured-output pipelines that turn model output into real audit artifacts
Take ambiguous problem statements and turn them into a plan, a shipped feature, and a clear read on what was cut and why
Work directly with an embedded subject matter expert and with design-partner firms, turning their feedback into agent changes within days
Expand agent coverage into new controls and new areas of internal audit
Who You Are (All Levels)
Product-minded and full-stack: you've shipped LLM-backed features to production against real users, and you measure yourself on whether they got used
You're fluent in evals and error analysis, and you apply them in service of shipping something practitioners trust
You have real opinions on model selection, prompting, and orchestration tradeoffs, and you can defend them with evidence rather than vibes
Energized by 0→1 work: you'd rather define the problem than inherit a spec, and you don't stall on ambiguity
Strong instincts for human-in-the-loop design
A genuine team player across the organization, not just within engineering: you'll work daily with PM, design, domain experts, and customer-facing teams, and you treat that as the best part of the job
Ship fast without leaving a mess: your code is reviewable, tested where it counts, and instrumented
Able to internalize a hard domain fast. You don't need to know SOX today, but you'll understand it well enough to make the right product calls
Higher-Level Responsibilities
At the Senior level, you may:
Own a major agent area end-to-end, from how the agent reasons about a class of controls through to the artifact a reviewer signs
Set the evals and error-analysis practice for the team's agent work, and decide what evidence justifies shipping a change or rolling it back
Collaborate with PMs and designers to shape roadmaps and define architectural tradeoffs, including where the agent acts and where the auditor decides
Own the harder model and orchestration judgment calls across a long multi-phase run
Mentor other engineers and raise the bar on 0→1 execution and applied eval rigor
At the Staff level, you may:
Drive agent initiatives that reach beyond Internal Audit and influence how agents are built across Fieldguide
Set and champion engineering standards for agent reliability, reproducibility, and defensibility
Partner with engineering and product leadership to define long-term technical strategy for agentic audit work
Serve as a trusted advisor to leaders across Engineering, Product, and Design
Represent Fieldguide externally through writing, speaking, and open-source contributions
Experience
Must-have:
Shipped LLM-backed product features to production against real users
Applied AI skillset: evals, error analysis, and model-selection decisions you owned and can explain
Comfortable full-stack, with enough backend depth to work in agent orchestration
Autonomy working from an ambiguous spec
A collaborative mode that works across PM, design, and domain experts
Nice-to-have:
Python, TypeScript, React, Postgres, Hasura, GraphQL
Temporal or comparable durable-execution / workflow orchestration
Hands-on eval experience (Langfuse, Braintrust, LangSmith, Arize Phoenix, or comparable)
Structured-output work including schema contracts, generating real artifacts from model output
Startup experience, as a founder or as an early engineer
A 0→1 track record: things you started where no scaffolding existed
Experience working directly with customers, and comfort being in the room when they use what you built
Background in internal audit, SOX, accounting, or another regulated domain
Document processing, including PDF and Excel manipulation and annotation
Not a fit if:
Prompt engineering is your whole skill set
Your agent work never carried production traffic
You want to own eval methodology or the evaluation harness itself rather than ship product features (better fit on Foundation Agents)
You need a fully specified ticket to start
You'd rather not be in the room with customers and domain experts
What Should Excite You
0→1 on the biggest bet: You're building the agent and the product from scratch, on the ground floor of where the company is going
Repeatable judgment: Making an agent reach the same defensible conclusion twice, in a domain where ground truth requires expert judgment
Real audit stakes: Your work directly affects what firms put in front of their clients, and what a reviewer is willing to sign
Customer proximity: Design-partner firms and an embedded SOX expert use what you ship within days of it landing
Human-in-the-loop design: Deciding where the agent acts and where the auditor decides, on work that genuinely matters
High trust, high autonomy: You're given ambiguous problems and trusted to define the plan
Benefits
Competitive compensation with equity
Comprehensive health and wellness benefits
Flexible time off and work schedules
Technology reimbursements
401(k) plan
Twice-yearly in-person offsites across the U.S.
Wellness benefits starting on your first day
Our Values
Fearless — Inspire and break down seemingly impossible walls
Fast — Launch fast with excellence; iterate to perfection
Lovable — Deliver happiness and 11-star experiences
Owners — Execute and run the business with ownership
Win-win — Create mutual value and earn trust for life
Inclusive — Scale the best ideas with inclusive teams
Similar jobs
- RA
Software Engineer I (Onsite)
NewRaytheon
Tewksbury, Massachusetts🇺🇸Hybrid38 minutes agoSAFeSpringSpring Boot+10Technology - JM
Software Engineer III - Core Java, API, and AI (Commercial & Investment Bank - Payments Technology)
NewJ.P. Morgan
Jersey City, New Jersey🇺🇸On-site38 minutes agoMicroservicesSQLSpring+6Technology - JM
Software Engineer Multiple Positions Available
NewJ.P. Morgan
Chicago, Illinois🇺🇸$133.3k - $155k/yrOn-site38 minutes agoDockerSplunkDatadog+6Technology - JM
Senior Lead Software Engineer - Machine Learning I/O (Consumer & Community Banking)
NewJ.P. Morgan
Columbus, Ohio🇺🇸On-site38 minutes agoAWSMachine LearningAgile+3Technology - VT
Automation Engineer
NewVoyager Technologies, Inc.
Denver🇺🇸Remote10 hours agoDockerShellAWS+20Technology - ZT
Software Engineer III
ZoomInfo Technologies LLC
Waltham🇺🇸6 days agoGCPMicroservicesMySQL+12Technology