Quick Overview
Job Description
Hi,
Please find the job description below. Let me know if there are any questions.
Job Title: Onsite SRE Lead
Location: Pleasanton, CA (Onsite) -NO FLEXIBILITY IT'S 5 DAYS ONSITE
Employment Type: Contract - 12 Months
Retail background + E commerce Platform(Microservices) MANDATORY
Look for AKAMAI OR RELATED CDN like Cloudflare or Amazon CDN+ AWS+ SECURITY(IAM OR PCI)+SPLUNK
| Experience: 8+ Years (3+ Years in Lead / Senior SRE Role) eCommerce Platform Domain Experitise processes millions of transactions annually. As SRE Lead, you own the reliability, performance, and scalability of this mission-critical cloud-native system - from CDN edge to microservice core. |
ROLE OVERVIEW
We are seeking a seasoned Site Reliability Engineering Lead to take end-to-end ownership of our retail eCommerce platform's operational excellence. You will drive SRE culture across engineering teams, champion observability, and lead a cross-functional onsite/offshore team to deliver exceptional uptime, performance, and release velocity. This role blends deep technical expertise with strategic leadership - you'll be equally comfortable reviewing Kubernetes manifests, tuning Akamai rules, and presenting SLO dashboards to senior leadership.
KEY RESPONSIBILITIES
Platform Reliability & Availability
Define, monitor, and enforce SLIs, SLOs, and Error Budgets for all eCommerce services
Lead incident response (ICS model): detection, triage, remediation, post-mortem, and blameless RCA
Drive chaos engineering and game-day exercises to validate fault tolerance
Architect resilient multi-AZ, multi-region AWS deployments for zero-downtime operations
CI/CD & Automation
Own and evolve Jenkins-based CI/CD pipelines for microservices deployments on Kubernetes/EKS
Implement GitOps workflows; enforce trunk-based development and deployment gates
Automate infrastructure provisioning using CloudFormation, Terraform, and AWS CDK
Drive shift-left reliability practices: load testing, chaos gates, and SLO checks in pipeline
Cloud Infrastructure & CDN
Manage Akamai CDN configuration: edge rules, caching policies, TLS, WAF, and traffic shaping
Optimize AWS Auto Scaling Groups, ALBs, and CloudFront for peak eCommerce traffic events
Design and govern caching strategy across CDN, Redis, and Memcached layers
Manage EKS/K8s cluster operations: node pools, Helm releases, resource quotas, and HPA/VPA
Observability & Performance
Build and maintain Splunk APM and Splunk Cloud dashboards, alerts, and runbooks
Implement distributed tracing (OpenTelemetry) across all microservices
Conduct performance profiling and capacity planning aligned to traffic growth projections
Establish DORA metrics tracking: deployment frequency, lead time, MTTR, change failure rate
Security, Compliance & Collaboration
Partner with InfoSec on PCI-DSS compliance, vulnerability management, and secrets governance
Enforce least-privilege IAM, network segmentation, and container security scanning in pipelines
Lead cross-functional reviews with Dev, QA, and InfoSec for new service onboarding
Coordinate onsite and offshore SRE teams; mentor junior engineers and foster SRE culture
| MUST-HAVE SKILLS & TECHNOLOGIES Core Technical Skills |
|
Similar jobs
- MA
AI Platform Engineer
MassMutual
NEW YORK, NY🇺🇸Hybrid4 weeks agoDockerGCPAWS+9Technology - MA
AI Platform Engineer
MassMutual
BOSTON, MA🇺🇸Hybrid4 weeks agoDockerGCPAWS+9Technology - MA
AI Platform Engineer
MassMutual
SPRINGFIELD, MA🇺🇸Hybrid4 weeks agoDockerGCPAWS+9Technology - JT
Platform Engineer
NewJaven Technologies, Inc
Cincinnati, OH🇺🇸Hybrid23 hours agoSQLShellPythonTechnology - TR
Only W2: SRE / Production Reliability Engineer, Woonsocket, RI
NewTror
Woonsocket, RI🇺🇸Hybrid23 hours agoManufacturing - NC
DevOps Engineer — Senior with Security Clearance
NewNeuma Consulting LLC
Washingtn, DC🇺🇸Hybrid23 hours agoDockerMongoDBNode.js+16Technology