Quick Overview
Job Description
About the company
Hi, we're Ondo Finance. Our mission is to provide institutional-grade, blockchain-enabled investment products and services. We have both a technology arm that develops decentralized finance technology, and an asset management arm that creates and manages tokenized funds. We are the global leader in tokenized treasuries, tokenized stocks and ETFs, and are building the future of institutional-grade financial services onchain.
Founded by folks from Goldman Sachs Digital Assets Team, weโre backed by some of the best investors in the world including Founders Fund, Coinbase Ventures, Pantera Capital, Tiger Global, and more. We are currently the leaders in the space in terms of AUM and are well capitalized to continue growing the firm. We're fully remote, with team members across the U.S.
About the role
Ondo operates real-time trading systems that run around the clock across traditional and crypto venues. The platform spans low-latency Rust engines, a fleet of Go services for trading, execution, and PnL accounting, and a multi-region Kubernetes footprint on AWS.
We are looking for an SRE with strong systems programming skills to own the reliability, observability, and performance of this platform. This is a hands-on role: you will read and modify Go and Rust code, debug latency regressions down to the feed handler, run incident response during market hours, and build the automation that keeps a 24/7 trading system healthy with a small team.
Target outcomes
- Own production reliability for real-time trading services: trading engines, execution gateways, market data ingestion, and PnL/reconciliation pipelines
- Operate and evolve our multi-region Kubernetes clusters on AWS (EKS), deployed via GitOps (Flux) with SOPS-encrypted secrets
- Build and refine observability: Prometheus metrics and alerting, Datadog logs and dashboards, and the SLOs that catch degradation before it costs money
- Improve deploy safety: progressive rollouts, config-reload behavior, and guardrails that prevent a bad push from touching live trading
Responsibilities
- Debug production incidents end to end: stale market data feeds, exchange rate limits, WebSocket disconnects, order-lifecycle desyncs, and latency regressions in the trading path
- Harden market data ingestion from providers such as Databento and venue-native feeds (REST and WebSocket), including staleness detection, failover, and replay
- Build reconciliation and data-integrity tooling across live gauges, Postgres, and our S3 parquet data lake, so positions, fills, and PnL always agree
- Participate in an on-call rotation covering US equity market hours and 24/7 crypto venues
Requirements
- 5+ years in SRE, production engineering, or infrastructure roles, with meaningful time supporting real-time or latency-sensitive systems
- Strong programming ability in Go or Rust, and willingness to work in both; this role changes application code, not just infrastructure
- Deep, hands-on Kubernetes and AWS experience: you have run stateful, latency-sensitive workloads in production, not just stateless web services
- Strong observability instincts: fluent PromQL, structured-log analysis, and experience designing alerts with high signal and low noise
- Solid Linux internals and networking fundamentals: you can chase a p99 regression through the kernel, the NIC, or the GC
- Sound judgment under pressure and clear written communication during and after incidents
Nice to haves
- Experience operating trading systems, execution infrastructure, or market data infrastructure at a trading firm, exchange, or broker
- Familiarity with market microstructure and order lifecycle (order books, order types, fills and reconciliation)
- Experience with market data providers and protocols (Databento, SIP/prop equity feeds, venue WebSocket APIs)
- Exposure to crypto venues and on-chain trading
- Python for operational tooling and data analysis (pandas, parquet, BigQuery)
- Experience with GitOps workflows, infrastructure as code, and secrets management at scale
Tech stack
Go, Rust, Python | Kubernetes (EKS), Flux, SOPS | AWS (multi-region), S3 parquet lake | Prometheus, Grafana, Datadog | Postgres, CockroachDB, BigQuery | Databento, venue WebSocket/REST feeds
What we offer
- Competitive compensation including but not limited to salary, future token rights, and/or equity (according to your preferences) โ We are well-funded and believe that great talent deserves great compensation.
- Full benefits (medical, vision, and dental) and flexible vacation policy (PTO).
- Remote-first team across many countries โ You will be an early team member helping shape our vision, culture, and design practices.
- A+ colleagues โ Our team includes alumni from: Goldman Sachs, Blackrock, Two Sigma, Bridgewater, SpaceX, AWS, Meta, Google, McKinsey, Coinbase, Circle, Uniswap.
- Best-in-class investors โ We are proud to be backed by leading crypto experts and VCs, including Pantera Capital, Founders Fund and Coinbase Ventures.