Haystack
← Back to Jobs
Manufacturing
HT

Hiring | SRE / Production Support | SFO, CA | Contract

Healthcare Triangle IncSan Francisco, CA🇺🇸United StatesPosted 11 Sept 2026

Why This Role Stands Out

This long-term contract SRE role offers a fantastic opportunity to leverage your expertise in Kubernetes, Java, and cloud technologies to ensure the smooth operation of critical production systems. You'll thrive here if you possess strong troubleshooting skills and a passion for building resilient, high-performing applications. Apply today to join a dynamic team and make a significant impact!

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
San Francisco, CA, United States
Posted
21 hours ago

Job Description

Role : SRE / Production Support
Location : SFO, CA
Duration : Long-term Contract
Preferred Skills/Experience :
  • A successful candidate will have at least 5 years of experience in software development or technical support or operations experience.
  • Excellent troubleshooting, debugging and problem-solving skills required.'
  • Expertise in Kubernetes, Docker, Jenkins, and Java stack-based production systems administration.
  • Proficiency in at least two of the technologies -Cassandra, Yugabyte, Kafka, Microservices, Spring Boot, Spark Streaming, Flink desired.
  • Experience with monitoring and logging tools such as Datadog, Kibana, and Splunk.
  • Experience in incident management preferred.
  • Strong experience with load balancing principles (F5, etc.) a plus.
  • Strong experience with infrastructure/cloud technologies (Google Cloud Platform, AWS, Azure) preferred.
  • Experience in systems and multi-tier application and network troubleshooting.
  • Strong experience with configuration management and automation (e.g., Ansible) a plus.
  • Experience programming in core java or python is required.
  • Experience with scripting languages (Shell & Python).
  • Experience with database concepts (SQL or NO SQL).
  • You know Linux. Even Kubernetes still runs on computers! This means you can debug most normal issues with performance, networking, kernel drivers, package management, etc. or have a good idea where to start
  • Demonstrable knowledge of Terraform, Jenkins, Artifactory, a strong plus.
  • A BA/BS/master's degree in the field of CS or related field is preferred but not required.
Required Skills :
  • Create and maintain automation scripts for deployment, scaling, and monitoring.
  • Develop and enhance production monitoring and management capabilities leveraging existing platforms and tools.
  • Work independently and within a team to triage and remediate production system and application incidents.
  • Handle escalations from USBank Consumer Domain partners about critical issues.
  • Work with Domain, Infra and other Support teams inside the USBank in troubleshooting, escalating, and resolving critical site incidents.
  • Identify recurring system and application issues and work with cloud teams, infra teams, product development, vendors, and other stakeholders in investigating and resolving causes.
  • Send communications regarding outages to other USBank teams, partners, and other customers.
  • Maintain accurate documentation of site incidents, including impact details, timelines, steps taken for mitigation/resolution.
  • Develop and maintain technical documentation for all USBank operations infrastructure and practices.
  • Remediate all P1, P2, P3,P4 Vulnerabilities within the approved US Bank SLA.
  • Schedule: TBD

Similar jobs