← Back to Jobs
Technology
Site Reliability Engineer
HPTech Inc.Los Angeles, CA🇺🇸United StatesPosted 7 Aug 2026
Quick Overview
Work Type
On Site
Level
Mid Senior
Job Description
Position: Site Reliability Engineer-10+ Year exp required
Location : Sylmar, Los Angeles , CA (Onsite)
Job Summary
- We are seeking an experienced engineer who can analyze, diagnose, and optimize performance and reliability of large-scale distributed systems. This role requires deep technical understanding across the entire application stack, the ability to read and reason about code, and the capability to provide data-backed answers to both engineering teams and business stakeholders.
- This role goes beyond traditional operations or DevOps. The successful candidate will think like a software engineer, act like a systems engineer, and operate with a production-first mindset.
Key Responsibilities
- Performance & Reliability Engineering
- Analyze and resolve performance issues such as high latency, slow login, throughput degradation, and system instability.
- Perform deep, end-to-end investigations across the full stack including:
- Load balancers and traffic routing
- Web server and application runtime configurations
- Middleware and messaging systems
- Database performance (queries, indexing, pooling)
- Kubernetes clusters (pods, resources, scaling behavior)
- Linux OS tuning (CPU, memory, I/O, ulimits, networking)
- Identify root causes and propose clear, actionable engineering solutions.
- Distributed Systems Design
- Design, review, and influence high-performance, highly-available distributed architectures.
- Evaluate trade-offs related to scalability, latency, fault tolerance, and cost.
- Partner with development teams early to prevent reliability and performance issues before production.
- Capacity Planning & Scalability
- Assess system readiness for growth scenarios such as:
- “We plan to onboard 10,000 users in 6 months — can the system support it?”
- Perform capacity and scale analysis for:
- Application tiers
- Databases
- Messaging systems
- Kubernetes compute and storage
- Provide evidence-based recommendations supported by metrics, benchmarks, and production data.
- Engineering Collaboration
- Work closely with software engineering teams to:
- Review performance-critical code paths
- Propose improvements at code, configuration, or infrastructure level
- Improve system observability (metrics, logs, traces)
- Communicate complex technical findings clearly to both engineers and business stakeholders.
- Required Technical Skills
- Strong understanding of distributed systems and performance engineering
- Ability to read, analyze, and troubleshoot Java code
- Hands-on experience with:
- Kubernetes (resource management, scaling, container behavior)
- Linux internals and tuning
- PostgreSQL (queries, indexing, performance optimization)
- Proven experience building or operating high-availability, high-throughput systems
- Strong analytical and problem-solving skills with a data-driven approach
- Nice to Have
- Experience with Azure cloud services
- Messaging systems such as ActiveMQ
- Load testing and benchmarking experience
- Background in roles such as SRE, Performance Engineering, Platform Engineering
Required Skills & Qualifications
Technical Skills
- Hands-on experience with cloud platforms (Azure.
- Strong scripting skills (e.g., Python, Bash, PowerShell, or similar).
- Experience with deployment pipelines, automation, and monitoring tools.
- Solid understanding of cloud infrastructure, networking, and application operations.
LLM & AI Experience
- Practical experience working with Large Language Models (LLMs).
- Familiarity with applying LLMs to engineering or operational workflows is required.
Professional Attributes
- Strong desire to learn and deeply understand complex systems.
- Self-starter with the ability to take ownership and drive initiatives independently.
- Demonstrates leadership, accountability, and problem-solving mindset.
Strong collaboration and communication skills
Skills
Azure
Bash
Java
Kubernetes
LLM
PostgreSQL
PowerShell
Python
Similar jobs
Senior Data Platform Engineer
NuAxis LLC · Washington, United States
11 minutes ago.NET Application Support / DevOps
BCforward · Charlotte, United States
11 minutes ago$65/hrL2 Operations Support Engineer (Cloud & DevOps)
KONNECTINGTREE INC · United States
12 minutes agoCloud/DevOps Engineer - III
SGS Consulting · San Francisco, United States
12 minutes agoSr. DevSecOps Engineer (Hybrid Remote/On-site) with Security Clearance
TEKsystems c/o Allegis Group · Raleigh, United States
33 minutes agoSRE Engineer
UnivEdge Consulting LLC · United States
34 minutes ago