Quick Overview
Job Description
We are seeking a Staff Site Reliability Engineer to serve as the foundational Technical Lead for our Platform Engineering SRE organization. In this role, you will be the primary architect and visionary for the core technology foundations that underpin the world's leading derivatives marketplace.
As the technical lead for all SRE sub-teams, you will bridge the gap between high-level business goals and deep technical implementation, ensuring our Google Cloud Platform-native stack provides the mission-critical intersection of ultra-low latency and absolute reliability required for high-volume financial ecosystems.
As the Staff SRE, your goal is to evolve our platform from "infrastructure as a service " to "reliability as a product. " You will be responsible for the technical roadmap of our entire SRE domain, mentoring senior engineers and setting the global standard for operational excellence across our Python, Kafka, and Kubernetes shop.
What You Will Do
- Technical Vision & Roadmap: Define the 12–18 month technical strategy for the Platform SRE teams, focusing on the evolution of our global footprint and self-service capabilities.
- Architectural Authority: Act as the final technical authority for major infrastructure changes involving Google Cloud Platform, GKE, and our mission-critical Kafka event bus.
- Incident Command & Systemic Resilience: Lead the response for complex, cross-functional outages and drive a "blameless " culture that prioritizes systemic, code-based fixes over manual intervention.
- Internal Development Platform (IDP): Architect and oversee the building of high-level abstractions in Python to mask underlying complexity, providing a seamless "Golden Path " for our global technology stack.
- Reliability Governance: Standardize SLIs, SLOs, and Error Budgets across all platform teams, ensuring they are technically rigorous and directly tied to market integrity.
- Engineering Mentorship: Level up the entire SRE organization through design reviews, architectural "office hours, " and fostering an environment of continuous technical evolution.
What We're Looking For
- Strategic AI Integration: A mastery of leveraging Generative AI and Agentic workflows (e.g., Gemini) to build self-healing infrastructure and sophisticated automated troubleshooting frameworks.
- Software Engineering Mastery: Expert-level proficiency in Python (and ideally Go) to build production-grade distributed systems and custom Kubernetes operators.
- Cloud-Native Leadership: Deep-seated expertise in Google Cloud Platform (Networking, IAM, GKE) and the ability to scale Kafka clusters for high-throughput, low-latency financial environments.
- Advanced IaC & GitOps: Mastery of Terraform module design and ArgoCD for managing immutable infrastructure at an enterprise scale.
- Distributed Systems Theory: A rigorous understanding of non-linear system behaviors, distributed consensus, and the nuances of high-concurrency architectures.
- Executive Communication: The ability to translate sophisticated technical debt and architectural risks into clear business outcomes for senior leadership.
Experience:
- 10+ years in SRE, Systems Engineering, or Software Engineering roles within high-pressure environments.
- 3+ years in a Staff, Principal, or Tech Lead capacity overseeing multiple teams or complex platform domains.
- Proven Track Record: Experience leading large-scale cloud migrations or re-architecting core messaging/compute platforms in a regulated environment.
- Certifications: Google Cloud Platform Professional Cloud Architect or Kubernetes (CKA/CKAD).
- Full-Stack Exposure: Proficiency in Node.js or modern front-end frameworks.
- Domain Expertise: Experience in Financial Markets or highly regulated, high-concurrency environments.
- Agile Integration: Comfort working within Agile frameworks and highly collaborative software development lifecycles.
Similar jobs
- MS
Cloud Platform Engineer - Hybrid
NewMSYS Inc.
Charleston, WV🇺🇸On-site23 hours agoDockerAzureBash+5Technology - AT
Voice Network Engineer with Security Clearance
NewAbacus Technology
Belleville, IL🇺🇸HybridYesterdayExpressUnityTechnology - AG
Site Reliability Engineer (SRE)
NewASCII Group LLC
Dallas, TX🇺🇸$51/hrHybrid23 hours agoDockerAWSELK+12Technology - AG
Senior DevOps Engineer
NewASCII Group LLC
Weehawken Township, NJ🇺🇸On-site23 hours agoAzureGitLab CIHelm+4Technology - MT
Senior SharePoint & Power Platform Engineer
Milestone Technologies, Inc.
Torrance, CA🇺🇸$65 - $75/hrOn-site1 week agoPowerShellTechnology - TB
Marketo Platform Engineer
NewTrail Blazer Consulting LLC
Malvern, PA🇺🇸Hybrid23 hours agoSOAPAWSRESTTechnology