Haystack
← Back to Jobs
Full time
Technology
IA

Staff Platform & Reliability Engineer

Interface AiSan Francisco🇺🇸United StatesPosted Sep 18, 2026

Why This Role Stands Out

As a Staff Platform & Reliability Engineer at Interface Ai, you'll have a significant impact on a cutting-edge AI platform used by financial institutions, driving critical reliability and security initiatives. This role is perfect for a mid-senior engineer who thrives on building robust systems and ensuring enterprise-grade uptime in a mission-driven company. You'll be instrumental in shaping the future of AI in financial services, so we encourage you to apply and contribute to this exciting growth phase.

Quick Overview

Seniority
Mid Senior
Employment type
Full Time
Location
San Francisco, United States
Posted
11 hours ago
AWSOWASPSOC 2Service MeshBashComplianceKubernetesLLMMental HealthPythonRESTRisk ManagementTypeScript

Job Description

About interface.ai

interface.ai is the agentic AI platform for financial services — bringing conversational and agentic AI to the credit unions and community banks that serve everyday Americans. We're not a lab, we're not a demo company, and we're not burning runway on hypotheticals. We are in production, generating real revenue, and on a mission that matters: making intelligent financial services available to the millions of people who've never had a private banker. More than 100 banks and credit unions run on interface.ai, reaching over 10 million people. Many products are live today; the biggest bets — an AI-first contact center and an AI-native consumer banking experience — are what comes next. Backed by $30M in Series A funding and already cash-flow positive, we're at the inflection point: a proven product, paying customers, and a profitable business rebuilding itself as an AI-native company to lead a world where agents, not software, do the work.

The Role

You own how the platform ships and how it stays up — the reliability and security backbone behind every conversation on interface.ai.

Our customers are banks and credit unions, and the bar for trust is high: real money, real regulators, and the uptime our customers count on. This role holds the platform to the enterprise-grade, 99.99% reliability we're known for and sets the engineering practices that keep it there as we scale. You'll define the SLOs, own the deploy path, run incidents when they happen, and build the platform so a small, senior team working with AI tooling can operate it with confidence — and you'll set the reliability and security bar the rest of the org builds against.

This is a hands-on, senior-most IC seat. What we mean by senior isn't tenure — it's judgment: you've run production systems people depend on, you think in failure modes and containment, and AI tooling is already part of how you work every day.

What You'll Own

  • Reliability & SLOs — define customer-facing SLIs and SLOs across product surfaces, run an error-budget program that governs release decisions, and hold the platform to its reliability bar.

  • Resilience & disaster recovery — own the DR strategy: regional failover, written RTO/RPO per tier, resilience against third-party dependency failure, and recovery you've actually tested rather than just planned.

  • Deploy & delivery — a GitOps deploy path with progressive delivery, automated analysis, and one-click rollback for every service, plus drift detection and a full change audit trail.

  • Cloud & infrastructure-as-code — the AWS foundation, fully managed as code; the Kubernetes platform and service mesh; capacity, cost, and scale.

  • Incident management — end to end: paging and severity policy, an incident-commander rotation, status-page automation, blameless post-mortems with tracked actions, and SLA reporting to customers.

  • Observability — metrics, logging, and distributed tracing across services; dashboards and alerts as code; burn-rate alerting that pages on what actually matters.

  • AI-native operations — build the automation and guardrails that let a small, senior team operate the platform with AI in the loop: runbooks encoded for safe automation, self-healing for routine work, and golden-path templates that ship new services with SLOs, alerts, secrets, and policy built in.

Security & Compliance

You own the infrastructure's security posture — including the parts unique to running an AI platform in regulated financial services.

  • Cloud & cluster security — IAM least-privilege across accounts, secrets management with a single source of truth, admission control and network policy on every cluster, image signing / SBOM / scanning in CI, and cloud-audit and threat-detection feeds as first-class alert sources.

  • Compliance evidence — own the infrastructure evidence for SOC 2 Type II and for customer and regulator reviews (NCUA/FFIEC examinations, GLBA): change management, access reviews, backup/restore, and DR test records, automated wherever possible.

  • Securing AI systems — protect the platform against AI-specific threats such as prompt injection, data exfiltration through tools, tenant isolation for retrieval stores, and abuse detection on public-facing endpoints — and set the guardrails and credential boundaries for AI-assisted engineering workflows.

What We're Looking For

  • A senior-most, hands-on IC who has run production systems with real uptime commitments — and carried the pager for them.

  • Deep production Kubernetes on AWS, including service mesh, with GitOps-based delivery across many services.

  • Infrastructure-as-code at multi-account scale, including taking over and reshaping a large existing estate.

  • You've built an SLO and error-budget practice that actually changed release decisions, with alerting tuned to burn rate rather than noise.

  • You've delivered multi-region or DR capability with defined RTO/RPO and proven it with real failover tests.

  • Security as daily practice, not a checklist — least-privilege IAM, secrets management, admission and network policy, software supply-chain controls — and you've produced evidence that satisfied auditors (SOC 2, ISO 27001, PCI, or FFIEC-style exams).

  • Extreme AI fluency — you use frontier AI tools (Claude Code, Cursor) daily and have clear opinions about the boundaries agents should operate within.

  • Strong programming in TypeScript and/or Python, plus Bash — and writing clear enough that a regulator could follow your post-mortem.

  • BS/BA in Computer Science required; MS or PhD a strong plus. San Francisco-based and committed to working onsite; on-call participation. H1B transfers welcome.

Bonus Points

  • Real-time voice or telephony systems, or other latency-critical streaming workloads.

  • Operating streaming and analytical data platforms and durable workflow engines at scale.

  • Chaos engineering and running resilience exercises against live environments.

  • Regulated-industry background (fintech, banking, healthcare), including vendor-risk management.

  • Threat modeling for LLM and agentic applications (e.g., OWASP LLM Top 10) and defending tool and agent integrations.

  • Cloud cost engineering, including compute-fleet optimization and LLM spend attribution.

What This Role Is — And Isn't

This is a hands-on, senior-most IC seat — not a management track, and not a hands-off "architect" role. You'll be measured by what you ship, what stays up, and the reliability and security bar you set for the engineers around you. If you want the title without owning the pager, this isn't it.

It's also not a role for someone who needs a mature SRE org and a runbook for everything handed to them. You'll set the standards here, not inherit them, in a space moving faster than most companies can track. If that's the leverage point you've been looking for, we want to talk.

Benefits

🩺 100% paid health, dental & vision care
💰 401(k) & financial wellness perks
🍜 Daily meals on us
🚇 Commuter benefit
💪 Monthly wellness stipend
🤖 Claude Enterprise + frontier AI tools for every employee — build with the best
🏙️ Brand-new 21st-floor SF office at 44 Montgomery — floor-to-ceiling views, worth showing up to
🌴 Discretionary PTO + paid parental leave
🧘 Mental health, wellness & family benefits
🚀 A mission-driven team shaping the future of banking

Why interface.ai

Series A · $30M raised · Cash-flow positive — we're not burning cash hoping the product works. It works. Customers are live, revenue is real, and the mission is one you can explain to your family without a slide deck.

You'll work directly with Bruce Kim (CTO / Co-Founder), who sets the technical bar and goes deep on the hardest problems, and Srinivas Njay (CEO), who is hands-on daily across product and engineering. The team is small enough that your judgment shapes everything, and the market is large enough that what you build will matter for a long time.

interface.ai is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Similar jobs