Quick Overview
Job Description
SENIOR SITE RELIABILITY ENGINEER (SRE)
Java Microservices, Spring Boot & AKS Reliability Engineering
EMPLOYMENT TYPE Full-time / Contract | SENIORITY Senior | PRIMARY FOCUS Microservices & AKS |
Position Summary
We are seeking a Senior Site Reliability Engineer with strong hands-on software development experience in Java, Spring Boot, and cloud-native microservices. The ideal candidate is a backend software engineer who applies an SRE mindset to service design, production readiness, automation, observability, performance, and resiliency.
This role will focus on improving the reliability of distributed services deployed on Azure Kubernetes Service (AKS). The successful candidate must be comfortable reading and modifying application code, troubleshooting services across the application and Kubernetes layers, implementing telemetry, and engineering permanent solutions to production reliability risks.
ROLE EXPECTATION This is not a production support, Kubernetes administration, monitoring-tool administration, or ticket-driven operations role. Recent hands-on Java/Spring Boot microservices development is required. |
Primary Objectives
· Engineer reliable, scalable, and observable Java/Spring Boot microservices running on AKS.
· Improve service resiliency through code, architecture changes, automation, and production readiness standards.
· Build proactive detection and diagnostics using logs, metrics, traces, and service-level indicators.
· Reduce customer impact and operational toil by addressing reliability risks at the application and platform layers.
· Partner with development and platform teams to improve deployment safety, runtime stability, and service ownership.
Key Responsibilities
Java & Spring Boot Microservices Engineering
· Design, develop, review, and enhance Java-based microservices using Spring Boot and related frameworks.
· Build secure, scalable, and well-defined REST APIs and event-driven services.
· Apply software engineering practices including clean code, automated testing, code reviews, dependency management, and design patterns.
· Review service architecture and application code to identify reliability, performance, scalability, and supportability gaps.
· Troubleshoot complex failures across application logic, service dependencies, data stores, messaging components, and external integrations.
AKS & Cloud-Native Reliability
· Deploy, operate, and troubleshoot microservices running on Azure Kubernetes Service (AKS).
· Understand Kubernetes workloads, pods, deployments, services, ingress, ConfigMaps, secrets, health probes, resource limits, autoscaling, and rollout behavior.
· Partner with platform engineering teams on AKS networking, identity, security, node capacity, workload placement, and cluster-level dependencies.
· Analyze pod restarts, container failures, memory and CPU constraints, scaling behavior, readiness failures, and deployment issues.
· Improve deployment strategies, rollback readiness, workload stability, and safe production releases.
Reliability & Resiliency Engineering
· Define and manage SLIs, SLOs, Error Budgets, and service health indicators for critical microservices.
· Implement and validate timeouts, retries, circuit breakers, bulkheads, rate limiting, caching, idempotency, and graceful degradation.
· Conduct production readiness and dependency risk reviews for new services and major changes.
· Participate in load, performance, capacity, and resiliency testing, including controlled dependency failure scenarios.
· Identify systemic reliability risks and drive corrective application or platform engineering changes before customer impact occurs.
Observability & Diagnostics
· Implement structured logging, application metrics, distributed tracing, and correlation identifiers using OpenTelemetry, Micrometer, or comparable frameworks.
· Create actionable dashboards and alerts that reflect service health, dependency behavior, saturation, latency, throughput, and errors.
· Correlate telemetry across APIs, microservices, messaging, databases, AKS workloads, and third-party systems.
· Improve diagnostic detail in application code so failures can be identified quickly without relying on manual investigation.
· Partner with SRE and engineering teams to reduce noisy alerts and improve proactive detection.
Software Delivery & Automation
· Build and improve CI/CD pipelines for compile, test, scan, package, deploy, validate, and rollback activities.
· Automate health validation, release verification, diagnostics, scaling checks, and repetitive operational tasks.
· Build reusable libraries, templates, and standards for service instrumentation and reliability patterns.
· Use Infrastructure as Code and configuration management practices to support consistent environments.
· Mentor engineers on cloud-native design, application reliability, observability, and operational ownership.
Incident Management & Continuous Improvement
· Provide hands-on technical leadership during incidents involving Java services, AKS workloads, integrations, or shared dependencies.
· Use code, telemetry, Kubernetes events, deployment history, and dependency data to diagnose failures.
· Drive blameless post-incident reviews and ensure corrective actions result in permanent engineering improvements.
· Reduce mean time to detect, diagnose, recover, and prevent recurrence through automation and better system design.
Required Qualifications
Java & Microservices Development
· 7+ years of software engineering experience with significant recent hands-on backend development.
· Strong Java development experience and deep practical knowledge of Spring Boot.
· Experience designing, building, testing, and supporting production microservices and REST APIs.
· Strong understanding of distributed systems, synchronous and asynchronous communication, failure handling, and service ownership.
· Experience with automated unit, integration, contract, and component testing.
· Ability to read, debug, modify, and review production application code.
AKS & Kubernetes
· Hands-on experience deploying, operating, and troubleshooting applications on Azure Kubernetes Service (AKS).
· Strong understanding of containers, Docker, Kubernetes workload resources, service discovery, ingress, probes, secrets, configuration, and autoscaling.
· Experience diagnosing application failures using pod logs, Kubernetes events, resource metrics, and deployment status.
· Understanding of Kubernetes networking, identity, security, resource management, and rollout strategies.
· Experience working with Helm, Kubernetes manifests, GitOps, or comparable deployment methods.
Cloud, Integration & Data
· Experience with Azure services such as AKS, API Management, Service Bus, Key Vault, Azure Monitor, databases, and storage.
· Experience with messaging or event platforms such as Azure Service Bus, Kafka, RabbitMQ, or Event Hubs.
· Experience with relational and/or NoSQL databases and practical knowledge of query performance, connection management, and failure handling.
· Understanding of API security, OAuth 2.0, JWT, managed identity, secrets management, and certificate-based communication.
SRE & Observability
· Strong understanding of SRE principles, including SLIs, SLOs, Error Budgets, production readiness, incident response, and toil reduction.
· Hands-on experience with OpenTelemetry, structured logging, application metrics, distributed tracing, and alerting.
· Experience with an observability platform such as Splunk, Datadog, Dynatrace, Grafana, Prometheus, New Relic, or equivalent.
· Strong analytical, troubleshooting, communication, and cross-functional collaboration skills.
Preferred Qualifications
· Experience with Java 17 or later and modern Spring ecosystem components.
· Experience with Spring Cloud, Spring Security, Resilience4j, Micrometer, JPA, or comparable frameworks.
· Experience with Istio, service mesh, KEDA, or advanced AKS scaling patterns.
· Experience with Terraform, Bicep, or other Infrastructure as Code technologies.
· Experience supporting high-volume digital commerce, ordering, payment, loyalty, or customer-facing platforms.
· Experience with performance engineering, JVM tuning, chaos engineering, or fault injection.
· Experience leading technical designs and mentoring software engineers.
Success Measures
Similar jobs
- PT
Site Reliability Engineer (SRE)
NewPrudent Technologies and Consulting
Santa Clara, CA🇺🇸Hybrid17 hours agoMicroservicesBashChef+7Technology - LT
Site Reliability Engineer
NewLorven Technologies, Inc.
Minneapolis, MN🇺🇸Hybrid17 hours agoAWSELKLogstash+3Technology - MG
Senior DevOps Engineer – AWS/Kubernetes/GitHub Actions
NewMedinext Global LLC
Chicago, IL🇺🇸On-site17 hours agoDockerAWSNginx+9Technology - HS
API Gateway/Platform Engineer
NewHierarch Soft Technologies, Inc.
Wilmington, DE🇺🇸Hybrid17 hours agoDockerMicroservicesSpring+8Technology - IS
AWS Cloud DevOps Engineer at Berkeley Heights, NJ (Hybrid)
NewIRIS Software, Inc.
Berkeley Heights, NJ🇺🇸Hybrid17 hours agoAWSCDKCloudFormation+1Technology - UT
Senior DevSecOps / Cloud / Platform Engineer
NewUnicorn Technologies LLC
Atlanta, GA🇺🇸Hybrid17 hours agoDockerSpringAWS+7Technology