Quick Overview
Job Description
* Job Summary
* We are seeking an experienced engineer who can analyze, diagnose, and optimize performance and reliability of large-scale distributed systems. This role requires deep technical understanding across the entire application stack, the ability to read and reason about code, and the capability to provide data-backed answers to both engineering teams and business stakeholders.
* This role goes beyond traditional operations or DevOps. The successful candidate will think like a software engineer, act like a systems engineer, and operate with a production-first mindset.
________________________________
* Key Responsibilities
* Performance & Reliability Engineering
* Analyze and resolve performance issues such as high latency, slow login, throughput degradation, and system instability.
* Perform deep, end-to-end investigations across the full stack including:
* Load balancers and traffic routing
* Web server and application runtime configurations
* Middleware and messaging systems
* Database performance (queries, indexing, pooling)
* Kubernetes clusters (pods, resources, scaling behavior)
* Linux OS tuning (CPU, memory, IO, ulimits, networking)
* Identify root causes and propose clear, actionable engineering solutions.
* Distributed Systems Design
* Design, review, and influence high-performance, highly-available distributed architectures.
* Evaluate trade-offs related to scalability, latency, fault tolerance, and cost.
* Partner with development teams early to prevent reliability and performance issues before production.
* Capacity Planning & Scalability
* Assess system readiness for growth scenarios such as:
* We plan to onboard 10,000 users in 6 months - can the system support it?
* Perform capacity and scale analysis for:
* Application tiers
* Databases
* Messaging systems
* Kubernetes compute and storage
* Provide evidence-based recommendations supported by metrics, benchmarks, and production data.
* Engineering Collaboration
* Work closely with software engineering teams to:
* Review performance-critical code paths
* Propose improvements at code, configuration, or infrastructure level
* Improve system observability (metrics, logs, traces)
* Communicate complex technical findings clearly to both engineers and business stakeholders.
* Required Technical Skills
* Strong understanding of distributed systems and performance engineering
* Ability to read, analyze, and troubleshoot Java code
* Hands-on experience with:
* Kubernetes (resource management, scaling, container behavior)
* Linux internals and tuning
* PostgreSQL (queries, indexing, performance optimization)
* Proven experience building or operating high-availability, high-throughput systems
* Strong analytical and problem-solving skills with a data-driven approach
* Nice to Have
* Experience with Azure cloud services
* Messaging systems such as ActiveMQ
* Load testing and benchmarking experience
* Background in roles such as SRE, Performance Engineering, Platform Engineering
Required Skills & Qualifications
Technical Skills
* Hands-on experience with cloud platforms (Azure.
* Strong scripting skills (e.g., Python, Bash, PowerShell, or similar).
* Experience with deployment pipelines, automation, and monitoring tools.
* Solid understanding of cloud infrastructure, networking, and application operations.
LLM & AI Experience
* Practical experience working with Large Language Models (LLMs).
* Familiarity with applying LLMs to engineering or operational workflows is required.
Professional Attributes
* Strong desire to learn and deeply understand complex systems.
* Self-starter with the ability to take ownership and drive initiatives independently.
* Demonstrates leadership, accountability, and problem-solving mindset.
* Strong collaboration and communication skills
Similar jobs
- GD
Platform DevOps Administrator
NewGDH
United States🇺🇸$45 - $50/hrRemote21 hours agoDockerNode.jsAnsible+5Technology - AR
Engineer Sr I - Product Security (AI Platform Engineer) -remote
NewArthrex
United States🇺🇸Remote21 hours agoSSOAzureBash+5Technology - VC
Release Train Engineer
NewVish Consulting Services, Inc.
United States🇺🇸Hybrid21 hours agoSAFeScrumAgile+2Engineering - AM
NoSQL DevOps Engineer
NewAmmaluIT
Sunnyvale, CA🇺🇸Hybrid21 hours agoDockerAWSEncryption+6Technology - IG
Site Reliability Engineer with Security Clearance
NewInsight Global, Inc.
Arlington, VA🇺🇸Remote21 hours agoAWSAzureKubernetes+2Technology - VS
Devops Engineer with Harness || Charlotte, NC || Plano, TX || Phoenix, AZ || Johnston, RI || W2 & C2C
NewValue Spectrum Technologies LLC
Plano, TX🇺🇸Hybrid21 hours agoMicroservicesAWSAnsible+10Technology