Why This Role Stands Out
Thrive in a mission-driven environment by shaping the reliability of cutting-edge AI weather forecasting technology, working closely with the Department of Defense. If you excel in fast-paced settings and are passionate about building scalable, highly available systems, this on-site role in San Francisco offers a unique opportunity to make a significant impact.
Quick Overview
Job Description
On behalf of our client, Dice Talent Solutions is seeking a Site Reliability Engineer.
This full-time (direct hire) position. Our client is looking to hire a candidate local to Boston, MA, Omaha, NE, or Washington DC who is able to travel when required to government buildings in these metro areas.
Candidates must be able to work on our client s W2 (annual salary + benefits package). Due to government contract requirements, United States Citizenship is required. No visa sponsorship or transfers are available. No 3rd party Corp to Corp inquiries will be considered.
Summary: Our client is a cutting-edge startup focused on delivering AI-based weather forecasting solutions to enhance climate resilience and public safety. To deliver advanced weather forecasts, they use numerous cutting-edge computing environments, including cloud-based hyperscalers, on-premise GPU clusters, field-deployed computers, and large supercomputing centers. They work closely with the Department of Defense (DoD) and must adhere to strict compliance and security standards. The team thrives in a dynamic, fast-paced environment, and every member plays a critical role in driving our mission forward.
Responsibilities:
- Scaling Production Environment:
- Design and implement scalable infrastructure solutions to support growing business needs.
- Optimize system performance and availability through capacity planning and performance tuning.
- Monitoring Stack Improvement:
- Develop and maintain whitebox (application-level) and blackbox (system-level) monitoring systems.
- Ensure comprehensive observability through the integration of metrics, logging, and tracing.
- Utilize tools such as Datadog and PagerDuty to establish reliable alerting and incident response processes.
- ETL Pipeline Management:
- Collaborate with data teams to ensure data quality and pipeline reliability.
- CI/CD and Build Systems:
- Implement and manage continuous integration and continuous deployment pipelines.
- Improve developer productivity by maintaining reliable build systems and workflows.
- On-call Responsibilities:
- Participate in on-call rotations to ensure high availability and timely incident resolution.
- Develop and automate incident response playbooks to minimize downtime.
Required:
- 6+ years of experience in SRE or DevOps roles.
- Proficiency with infrastructure as code tools, particularly Terraform.
- Experience with AWS &/or Google Cloud
- Strong background in software engineering and CI/CD pipeline management.
- Experience with on-call operations and incident management.
- Familiarity with monitoring and alerting tools such as Datadog and PagerDuty.
- Knowledge of cloud-native services, including SNS/SQS and Redis.
- Experience with ML experiment tracking and GPU optimization is a plus.
Desired:
- Expertise in managing and scaling distributed systems.
- Strong understanding of networking, security, and Linux systems.
- Experience in automating infrastructure and deployment processes.
- Familiarity with message queues (SNS/SQS) and caching systems (Redis).
- Knowledge of ML workflows and GPU resource management.
Similar jobs
- KP
ITSM Consultant V (Senior Agile Scrum - DevOps)
NewKaiser Permanente
Greensboro, North Carolina🇺🇸Hybrid24 minutes agoSAFeScrumAgileTechnology - SO
Lead AWS Cloud Platform Engineer - IaC / SRE – Preferred CNAPP / Fortinet / AWS Clean Room
NewSystem One
Reston, VA🇺🇸$109/hrOn-site9 hours agoDockerMicroservicesShell+10Technology - LE
DevOps Engineer
NewLeidos
Winter Garden, FL🇺🇸$87.1k - $157.4k/yrHybridYesterdaySAFeDockerMicroservices+12Technology - LE
DevOps Engineer
NewLeidos
Kissimmee, FL🇺🇸$87.1k - $157.4k/yrHybridYesterdaySAFeDockerMicroservices+12Technology - LE
DevOps Engineer
NewLeidos
Orlando, FL🇺🇸$87.1k - $157.4k/yrHybridYesterdaySAFeDockerMicroservices+12Technology - AT
DevOps Engineer [211923]
Aquent Talent
Phoenix, AZ🇺🇸Hybrid5 weeks agoAgileBashC#+5Technology