Site Reliability Engineer
Quick Overview
Job Description
On behalf of our client, Dice Talent Solutions is seeking a Site Reliability Engineer.
This full-time (direct hire) position located in San Francisco, CA. Work is performed 100% on-site.
Candidates must be able to work on our client s W2 (annual salary + benefits package). Due to government contract requirements, United States Citizenship is required. No visa sponsorship or transfers are available. No 3rd party Corp to Corp inquiries will be considered.
Summary: Our client is a cutting-edge startup focused on delivering AI-based weather forecasting solutions to enhance climate resilience and public safety. To deliver advanced weather forecasts, they use numerous cutting-edge computing environments, including cloud-based hyperscalers, on-premise GPU clusters, field-deployed computers, and large supercomputing centers. They work closely with the Department of Defense (DoD) and must adhere to strict compliance and security standards. The team thrives in a dynamic, fast-paced environment, and every member plays a critical role in driving our mission forward.
Responsibilities:
- Scaling Production Environment:
- Design and implement scalable infrastructure solutions to support growing business needs.
- Optimize system performance and availability through capacity planning and performance tuning.
- Monitoring Stack Improvement:
- Develop and maintain whitebox (application-level) and blackbox (system-level) monitoring systems.
- Ensure comprehensive observability through the integration of metrics, logging, and tracing.
- Utilize tools such as Datadog and PagerDuty to establish reliable alerting and incident response processes.
- ETL Pipeline Management:
- Design, launch, and maintain robust ETL pipelines to support data-driven operations.
- Collaborate with data teams to ensure data quality and pipeline reliability.
- CI/CD and Build Systems:
- Implement and manage continuous integration and continuous deployment pipelines.
- Improve developer productivity by maintaining reliable build systems and workflows.
- On-call Responsibilities:
- Participate in on-call rotations to ensure high availability and timely incident resolution.
- Develop and automate incident response playbooks to minimize downtime.
Required:
- 6+ years of experience in SRE or DevOps roles.
- Proficiency with infrastructure as code tools, particularly Terraform.
- Experience with AWS &/or Google Cloud
- Strong background in software engineering and CI/CD pipeline management.
- Experience with on-call operations and incident management.
- Familiarity with monitoring and alerting tools such as Datadog and PagerDuty.
- Knowledge of cloud-native services, including SNS/SQS and Redis.
- Experience with ML experiment tracking and GPU optimization is a plus.
Desired:
- Expertise in managing and scaling distributed systems.
- Strong understanding of networking, security, and Linux systems.
- Experience in automating infrastructure and deployment processes.
- Familiarity with message queues (SNS/SQS) and caching systems (Redis).
- Knowledge of ML workflows and GPU resource management.
Skills
Similar jobs
Senior Gaming Platform Engineer
Apple, Inc. · United States
5 minutes agoSite Reliability Engineer (DevOps)
Accenture LLP · Annapolis, United States
6 minutes ago$111.8k - $221.8k/yrJunior DevOps Engineer
Leidos · Chantilly, United States
6 minutes ago$69.5k - $125.7k/yrSAP DevOps Engineer
Accenture LLP · Washington, United States
6 minutes ago$100.2k - $203.4k/yrDevOps Cloud Engineer with Security Clearance
Woodside Staffing Solutions & Consulting · Co Spgs, United States
9 minutes agoConfiguration Manager with Security Clearance
SOSi · Pearl Harbor, United States
12 minutes ago