Haystack
← Back to Jobs
Technology

Infrastructure Support Engineer / SRE / Cloud Operations Engineer

GLOBAL IT CON LLCSan Jose, CA🇺🇸United StatesPosted 6 Aug 2026

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

Position Overview

  • Role: Infrastructure Support Engineer / SRE / Cloud Operations Engineer

  • Location: Sunnyvale, CA / San Jose, CA (Onsite / In-Person Interview)

  • Duration: Long Term

  • Experience Level: 10+ Years Overall Experience

Key Responsibilities

  • Production Operations & Monitoring: Monitor enterprise production environments using tools such as Datadog, Grafana, Prometheus, New Relic, or Splunk.

  • Incident & Outage Management: Lead triaging, escalation, and resolution for Sev1, Sev2, and Sev3 incidents; perform thorough Root Cause Analysis (RCA).

  • Infrastructure Troubleshooting: Identify and resolve operational issues across virtual machines, OS, storage (SAN/NAS), and networking components (DNS, Load Balancers, TCP/IP).

  • Cloud Operations: Support core cloud resources across AWS, Azure, or Google Cloud Platform (VMs, VPCs/Networking, IAM, Storage).

  • Stakeholder & Customer Communication: Act as the primary technical contact during critical incidents, ensuring clear, timely updates to internal teams and customers.

Key Requirements & Qualifications

  • 10+ years of progressive IT experience in Cloud Operations, System Administration, or Site Reliability Engineering (SRE).

  • Demonstrated experience handling high-severity incident management (Sev1/Sev2/Sev3) and writing post-incident RCA reports.

  • Hands-on proficiency with enterprise monitoring platforms (Datadog, Grafana, Prometheus, Splunk, Dynatrace, or New Relic).

  • Strong foundational knowledge of Compute (VMs, CPU/RAM utilization), Storage (SAN, NAS, Disk management), and Networking (TCP/IP, DNS, Firewalls, Load Balancers).

  • Solid operational grasp of at least one public cloud platform (AWS, Azure, or Google Cloud Platform).

  • Excellent communication and customer-handling skills during high-pressure outages.

Skills

AWS
New Relic
Splunk
TCP/IP
Azure
DNS
Datadog
Google Cloud
Grafana
Prometheus

Similar jobs