Why This Role Stands Out
This Site Reliability Engineer II role offers significant growth potential by allowing you to work with cutting-edge Google Cloud Platform data platforms and develop advanced automation skills. You'll thrive if you're passionate about ensuring system reliability and proactively solving complex technical challenges in a collaborative, remote-friendly environment. Apply now to contribute to a reputable tech company and advance your SRE career.
Quick Overview
Job Description
Site Reliability Engineer II
REMOTE
Full Time
Overview / Summary
We are seeking a Site Reliability Engineer to join an SRE team focused on observability, monitoring, and technical consulting across Google Cloud Platform-based data platforms. This role is responsible for ensuring the availability, reliability, and performance of cloud and network systems and services through automation, monitoring, troubleshooting, and continuous optimization.
Key Responsibilities
- Collaborate with infrastructure teams to implement critical solutions by automating routine tasks.
- Monitor and manage production environments, proactively identifying and resolving issues.
- Participate in building advanced tooling for system access monitoring, log session recording, and reliability administration across multiple geographically distributed data centers.
- Engage with engineering teams to improve on-call efficiencies, incident management, and post-mortem analysis.
- Perform capacity planning and optimization to support growing demands and traffic patterns.
- Maintain monitoring and alerting systems for proactive system health checks.
- Continuously improve system performance, stability, and security through data-driven analysis and optimization.
- Create and maintain comprehensive documentation and diagrams to facilitate knowledge sharing.
- Work hands-on with cloud infrastructure, BigQuery workloads, CI/CD pipelines, and enterprise monitoring tools to maintain critical systems at scale.
Required Qualifications
- Bachelor's degree.
- 4+ years of experience in IT.
- 3+ years of development experience.
- Practitioner-level experience with at least one coding language or framework.
- Hands-on experience with Google Cloud Platform (Google Cloud Platform).
- Experience with BigQuery.
- Experience with Dynatrace.
- Proficiency with monitoring and observability tools, ideally Dynatrace or comparable tools such as Datadog or New Relic.
- Familiarity with ITSM tools such as ServiceNow, including incident, problem, and change management.
Preferred Qualifications
- Experience with Google Cloud Platform Cloud Run.
- Experience with Python.
- Strong troubleshooting and problem-solving skills.
- Familiarity with AI tools, including agents, skills, LLMs, and copilots.
- Experience defining and tracking SLAs, SLOs, and SLIs.
Similar jobs
- HP
Azure DevOps Senior Technical Consultant
NewHPTech Inc.
United States🇺🇸Hybrid21 hours agoAzureGitHub ActionsTerraformTechnology - ST
Azure Network DevOps Engineer
NewSRI Tech Solutions
Plano, TX🇺🇸Hybrid21 hours agoDockerAWSAzure+6Technology - PG
Business Analyst Requirements, User Stories, Azure DevOps, Onsite - 69986
NewPRIMUS Global Services Inc.
TX🇺🇸On-site21 hours agoAzureTechnology - WS
DevOps Cloud Engineer with Security Clearance
NewWhite Sky Technologies
Annapolis Junction, MD🇺🇸$140k - $235k/yrHybrid2 days agoDockerRubyAWS+9Technology - OC
Systems Engineer SRE- TS/SCI + FS Poly with Security Clearance
Our client is a software and systems development firm, built by
Chantilly, VA🇺🇸On-site4 days agoDockerMySQLAWS+9Technology - TA
Staff Platform Engineer, Security with Security Clearance
NewTrue Anomaly. Inc.
Denver, CO🇺🇸$205k - $295k/yrHybrid21 hours agoAWSEncryptionAzure+6Technology