Why This Role Stands Out
This hybrid Site Reliability Engineer role offers a fantastic opportunity to deeply influence the performance and reliability of large-scale distributed systems, fostering significant technical growth. You'll thrive here if you possess a strong software engineering mindset combined with systems engineering expertise, ready to tackle complex challenges and collaborate across teams. Apply now to leverage your extensive experience in a role that values in-depth analysis and proactive system design.
Quick Overview
Job Description
Position: Site Reliability Engineer-10+ Year exp required
Location : Sunnyvale CA - Local and F2F
Duration: w2
Job Summary
We are seeking an experienced engineer who can analyze, diagnose, and optimize performance and reliability of large-scale distributed systems. This role requires deep technical understanding across the entire application stack, the ability to read and reason about code, and the capability to provide data-backed answers to both engineering teams and business stakeholders.
This role goes beyond traditional operations or DevOps. The successful candidate will think like a software engineer, act like a systems engineer, and operate with a production-first mindset.
Key Responsibilities
Performance & Reliability Engineering
Analyze and resolve performance issues such as high latency, slow login, throughput degradation, and system instability.
Perform deep, end-to-end investigations across the full stack including:
Load balancers and traffic routing
Web server and application runtime configurations
Middleware and messaging systems
Database performance (queries, indexing, pooling)
Kubernetes clusters (pods, resources, scaling behavior)
Linux OS tuning (CPU, memory, IO, ulimits, networking)
Identify root causes and propose clear, actionable engineering solutions.
Distributed Systems Design
Design, review, and influence high-performance, highly-available distributed architectures.
Evaluate trade-offs related to scalability, latency, fault tolerance, and cost.
Partner with development teams early to prevent reliability and performance issues before production.
Capacity Planning & Scalability
Assess system readiness for growth scenarios such as:
We plan to onboard 10,000 users in 6 months can the system support it?
Perform capacity and scale analysis for:
Application tiers
Databases
Messaging systems
Kubernetes compute and storage
Provide evidence-based recommendations supported by metrics, benchmarks, and production data.
Engineering Collaboration
Work closely with software engineering teams to:
Review performance-critical code paths
Propose improvements at code, configuration, or infrastructure level
Improve system observability (metrics, logs, traces)
Communicate complex technical findings clearly to both engineers and business stakeholders.
Required Technical Skills
Strong understanding of distributed systems and performance engineering
Ability to read, analyze, and troubleshoot Java code
Hands-on experience with:
Kubernetes (resource management, scaling, container behavior)
Linux internals and tuning
PostgreSQL (queries, indexing, performance optimization)
Proven experience building or operating high-availability, high-throughput systems
Strong analytical and problem-solving skills with a data-driven approach
Nice to Have
Experience with Azure cloud services
Messaging systems such as ActiveMQ
Load testing and benchmarking experience
Background in roles such as SRE, Performance Engineering, Platform Engineering
Required Skills & Qualifications
Technical Skills
Hands-on experience with cloud platforms (Azure.
Strong scripting skills (e.g., Python, Bash, PowerShell, or similar).
Experience with deployment pipelines, automation, and monitoring tools.
Solid understanding of cloud infrastructure, networking, and application operations.
LLM & AI Experience
Practical experience working with Large Language Models (LLMs).
Familiarity with applying LLMs to engineering or operational workflows is required.
Professional Attributes
Strong desire to learn and deeply understand complex systems.
Self-starter with the ability to take ownership and drive initiatives independently.
Demonstrates leadership, accountability, and problem-solving mindset.
Strong collaboration and communication skills
Similar jobs
- BT
Cloud-Native Platform Engineer
NewBraintree Technology Solutions
Dearborn, MI🇺🇸On-site22 hours agoSpringSpring BootAPI Gateway+14Technology - QT
Sr.Devops Engineer
NewQuantom Tech LLC
Atlanta, GA🇺🇸Hybrid22 hours agoDockerScrumAgile+7Technology - BT
Palantir Platform Engineer with Security Clearance
NewBespoke Technologies Inc.
Herndon, VA🇺🇸Hybrid22 hours agoETLPythonTechnology - SB
Site Reliability Engineer (SRE) Remote Location
NewSierra Business Solution LLC
United States🇺🇸Remote22 hours agoDockerMicroservicesMySQL+16Technology - DS
Azure DevOps Engineer
NewDecision Six Inc.
Southfield, MI🇺🇸Hybrid22 hours agoDockerMicroservicesSSO+6Technology - TS
Senior Observability Engineer
NewTechgene Solutions LLC
United States🇺🇸Remote22 hours agoEngineering