Haystack
← Back to Jobs
Technology

Site Reliability Engineer - Cloud Infrastructure

Fort Technologies Inc.San Jose, CA🇺🇸United StatesPosted 6 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Looking for Level 1 Infrastructure Support Engineer / Site Reliability Engineer (SRE) / Cloud Operations Engineer who can handle production environments and customer-facing incidents. 

 

Skills

  • Overall 8+ years of experience
  • Core platform Engineering 
  • Incidents, Alerts
  • SRE Mindset
  • Customer handling skills
  • Infra/Cloud basic concepts
  • Network
  • Storage

 

Primary Responsibilities:

  • Monitor production environments.
  • Handle incidents, alerts, and outages.
  • Perform root cause analysis (RCA).
  • Troubleshoot infrastructure issues.
  • Coordinate with customers and internal teams during critical incidents.
  • Ensure system reliability and availability.

Must-Have Skills

1. SRE Mindset

  • Understanding of reliability, availability, and performance.
  • Focus on automation and reducing manual efforts.
  • Experience with incident management and problem management.

2. Incident & Alert Management

  • Handling Sev1, Sev2, and Sev3 incidents.
  • Experience with monitoring tools such as:
    • Datadog
    • New Relic
    • Dynatrace
    • Splunk
    • Grafana
    • Prometheus
  • Performing RCA and post-incident reviews.

3. Infrastructure Fundamentals

Strong understanding of:

  • Compute: Virtual Machines, CPU, Memory utilization
  • Storage: SAN, NAS, Disk management, Storage troubleshooting
  • Network: TCP/IP, DNS, Load Balancers, Firewalls, Routing basics

4. Cloud Basics

Knowledge of at least one cloud platform:

  • AWS
  • Azure
  • Google Cloud Platform

Skills

AWS
New Relic
Splunk
TCP/IP
Azure
DNS
Datadog
Google Cloud
Grafana
Prometheus

Similar jobs