Senior Reliability Engineer
Quick Overview
Job Description
Job Title: Staff Site Reliability Engineer – Platform Engineering
Job Reference ID: 34344
Position Type: Direct Placement
Industry: Financial Services
Location: Chicago, IL (Hybrid - 2 Days Onsite/Week in Downtown Chicago)
Work Authorization: Open to candidates authorized to work in the US
About the Role
We are seeking a Staff Site Reliability Engineer (Platform Engineering) to serve as the foundational Technical Lead for a premier global Financial Services enterprise. In this role, you will be the primary architect and visionary for core technology foundations that underpin high-volume, ultra-low latency financial marketplaces.
You will bridge the gap between high-level business strategy and deep technical implementation, ensuring our Google Cloud Platform-native stack provides mission-critical reliability, extreme scalability, and operational excellence. Your goal is to evolve the platform from "Infrastructure as a Service" to "Reliability as a Product."
Key Responsibilities
Technical Vision & Strategy: Define and execute the 12–18 month technical roadmap for the Platform SRE ecosystem, building high-level Internal Development Platform (IDP) abstractions in Python.
Architectural Leadership: Serve as the final technical authority for core infrastructure architectures spanning Google Cloud Platform, GKE, and enterprise-grade Kafka messaging clusters.
Incident Command & Resilience: Lead response strategies for complex, cross-functional outages; foster a blameless engineering culture focused on code-driven, automated resiliency.
Reliability Governance: Standardize and enforce SLIs, SLOs, and Error Budgets across all engineering pods to safeguard system integrity.
GenAI & Intelligent Ops: Leverage Generative AI and Agentic workflows (e.g., Gemini) to build self-healing infrastructure and automated root-cause analysis frameworks.
Engineering Mentorship: Elevate the global SRE organization through architectural office hours, design reviews, and engineering best practices.
Required Qualifications
Experience: 10+ years in SRE, Infrastructure, or Software Engineering roles in high-concurrency, high-availability environments.
Leadership: 3+ years in a Staff, Principal, or Tech Lead capacity overseeing complex platform engineering domains.
Cloud Native Mastery: Expertise in Google Cloud Platform (Networking, IAM, GKE) and scaling Kafka event buses for low-latency operations.
Software Engineering: Expert-level proficiency in Python (and ideally Go) for writing production-grade distributed systems and custom Kubernetes operators.
IaC & GitOps: Hands-on mastery of Terraform module design and GitOps patterns via ArgoCD.
Location: Chicago-based or willing to relocate to Chicago (hybrid schedule requiring 2 days/week on-site).
Nice to Have
Prior experience in Financial Markets, High-Frequency Trading (HFT), or heavily regulated financial ecosystems.
Google Cloud Platform Professional Cloud Architect or Certified Kubernetes Administrator (CKA/CKAD).
Full-Stack exposure (Node.js or modern web frameworks).
Skills
Similar jobs
Senior Cyber Security Fusion Analyst
Leidos · Odenton, United States
Just now$131.3k - $237.3k/yrCyber Security Systems Engineer
Leidos · Alexandria, United States
Just now$107.9k - $195.1k/yrSoftware Engineering Intern (On-site) with Security Clearance
RTX · State College, United States
7 minutes ago$37k - $82k/yrTest Systems Engineer II (Onsite Ft. Wayne, IN) with Security Clearance
RTX · Fort Wayne, United States
7 minutes ago$68.9k - $131.1k/yrJava Developer
TEKsystems c/o Allegis Group · Jersey City, United States
7 minutes ago$70 - $80/hrSenior Cloud Security Engineer - Boston/Cloud Assurance/AWS
Motion Recruitment Partners, LLC · Boston, United States
7 minutes ago