Why This Role Stands Out
Embrace a pivotal role in ensuring the reliability and scalability of a cutting-edge commerce iPaaS platform, leveraging AI-assisted development for innovative solutions. This remote opportunity is ideal for experienced SREs passionate about building robust systems and shaping the future of AI-native infrastructure. Join a fast-paced, high-growth company where your expertise will directly impact enterprise retailers and drive operational excellence.
Quick Overview
Job Description
About Satsuma
Satsuma is a commerce iPaaS that builds merchant-specific APIs, MCP Servers, and MCP Apps, enabling retailers to connect their full commerce stack once and deploy branded shopping experiences across every AI channel. We work with enterprise retailers and move fast. Our infra has to match.
The role
We're looking for a Senior SRE to own the reliability, scalability, and operational posture of Satsuma's multi-cloud infrastructure. You'll be the person who keeps things running, builds the systems that prevent fires, and makes on-call not terrible.
This is an infra-first role. But we're an AI-native company, and we expect you to use AI-assisted development (Claude Code) as a core part of your workflow — writing tooling, automating runbooks, building internal utilities.
What you'll do
- Own infrastructure across AWS, GCP, and Azure environments
- Build and maintain CI/CD pipelines, observability stacks, and incident response workflows
- Define and enforce SLOs/SLIs; lead postmortems
- Author and maintain IaC (Terraform preferred)
- Write internal tooling and automation using AI-assisted development workflows
- Partner closely with engineering on reliability reviews and architecture decisions
- 5-8 years in SRE, DevOps, or infrastructure engineering
- Hands-on experience across at least two major cloud providers
- Strong Kubernetes, Terraform, and observability tooling (Datadog, Grafana, or equivalent)
- Comfortable reading and editing code; able to ship scripts and internal tools
- Experience with AI-assisted development (Copilot, Cursor, Claude Code)
- On-call maturity -- you've owned incidents end-to-end and made systems better afterward
- Prior experience at a startup or high-growth SaaS company
- Familiarity with API gateway infrastructure or commerce tech stacks
- Hands-on experience with MCP or agentic AI infrastructure
- Unlimited PTO
- 401(K)
- Healthcare Stipend
- Gym stipend
Similar jobs
- JM
Site Reliability Engineer III
NewJ.P. Morgan
Irvine, California🇺🇸On-site21 minutes agoSpringSpring BootEncryption+8Technology - IP
ServiceNow DevOps Lead Developer
NewiTek People, Inc.
Chicago, IL🇺🇸Hybrid15 hours agoSOAPOAuthSonarQube+4Technology - WS
DevOps Cloud Engineer with Security Clearance
NewWhite Sky Technologies
Annapolis Junction, MD🇺🇸$140k - $235k/yrHybrid15 hours agoDockerRubyAWS+9Technology - TE
DevOps Engineer - Kansas City, KS, Denver, CO, Phoenix, AZ, St. Louis, MO, Nashville, TN.
NewTechniPros, LLC
Kansas City, KS🇺🇸Hybrid15 hours agoDockerAWSSonarQube+10Technology - AS
DevOps Engineer
NewApex Systems
Findlay, OH🇺🇸On-site15 hours agoAzureJenkinsTechnology - OR
Senior Site Reliability Engineer
NewOracle
Nashville, Tennessee🇺🇸$81.1k - $187k/yrHybrid1 hour agoOracleSQLAnsible+5Technology