Haystack
← Back to Jobs
Technology

Site Reliability Engineer

Dice Talent & Staffing SolutionsSan Francisco, CA🇺🇸United StatesPosted 12 Aug 2026

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

On behalf of our client, Dice Talent Solutions is seeking a Site Reliability Engineer.

This full-time (direct hire) position located in San Francisco, CA. Work is performed 100% on-site.

Candidates must be able to work on our client s W2 (annual salary + benefits package). Due to government contract requirements, United States Citizenship is required. No visa sponsorship or transfers are available. No 3rd party Corp to Corp inquiries will be considered.

Summary: Our client is a cutting-edge startup focused on delivering AI-based weather forecasting solutions to enhance climate resilience and public safety. To deliver advanced weather forecasts, they use numerous cutting-edge computing environments, including cloud-based hyperscalers, on-premise GPU clusters, field-deployed computers, and large supercomputing centers. They work closely with the Department of Defense (DoD) and must adhere to strict compliance and security standards. The team thrives in a dynamic, fast-paced environment, and every member plays a critical role in driving our mission forward.

Responsibilities:

  • Scaling Production Environment:
    • Design and implement scalable infrastructure solutions to support growing business needs.
    • Optimize system performance and availability through capacity planning and performance tuning.
  • Monitoring Stack Improvement:
    • Develop and maintain whitebox (application-level) and blackbox (system-level) monitoring systems.
    • Ensure comprehensive observability through the integration of metrics, logging, and tracing.
    • Utilize tools such as Datadog and PagerDuty to establish reliable alerting and incident response processes.
  • ETL Pipeline Management:
    • Design, launch, and maintain robust ETL pipelines to support data-driven operations.
    • Collaborate with data teams to ensure data quality and pipeline reliability.
  • CI/CD and Build Systems:
    • Implement and manage continuous integration and continuous deployment pipelines.
    • Improve developer productivity by maintaining reliable build systems and workflows.
  • On-call Responsibilities:
    • Participate in on-call rotations to ensure high availability and timely incident resolution.
    • Develop and automate incident response playbooks to minimize downtime.

Required:

  • 6+ years of experience in SRE or DevOps roles.
  • Proficiency with infrastructure as code tools, particularly Terraform.
  • Experience with AWS &/or Google Cloud
  • Strong background in software engineering and CI/CD pipeline management.
  • Experience with on-call operations and incident management.
  • Familiarity with monitoring and alerting tools such as Datadog and PagerDuty.
  • Knowledge of cloud-native services, including SNS/SQS and Redis.
  • Experience with ML experiment tracking and GPU optimization is a plus.

Desired:

  • Expertise in managing and scaling distributed systems.
  • Strong understanding of networking, security, and Linux systems.
  • Experience in automating infrastructure and deployment processes.
  • Familiarity with message queues (SNS/SQS) and caching systems (Redis).
  • Knowledge of ML workflows and GPU resource management.

Skills

AWS
ETL
Edge Computing
Datadog
Google Cloud
PagerDuty
Redis
Terraform

Similar jobs