Haystack
← Back to Jobs
Technology

AMS / Site Reliability Engineer (SRE) – Telematics & Connected Car

Info Way SolutionsIrvine, CA🇺🇸United StatesPosted 6 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

*]:pointer-events-auto R6Vx5W_threadScrollVars scroll-mb-[calc(var(--scroll-root-safe-area-inset-bottom,0px)+var(--thread-response-height))] scroll-mt-[calc(var(--header-height)+min(200px,max(70px,20svh)))]" dir="auto" data-turn-id="request-6a678188-2d94-83e8-bd1c-0c95fa0561c9-18" data-turn-id-container="request-6a678188-2d94-83e8-bd1c-0c95fa0561c9-18" data-testid="conversation-turn-144" data-turn="assistant">

Job Summary

We are seeking an experienced Application Management Services (AMS) / Site Reliability Engineer (SRE) to support mission-critical applications within the Telematics and Connected Car ecosystem. The ideal candidate will have strong experience in Incident Management, Application Support, SRE practices, monitoring and observability tools, and production support for automotive connected vehicle platforms.

The candidate will be responsible for ensuring application availability, monitoring production environments, driving incident resolution, performing root cause analysis (RCA), and collaborating with cross-functional engineering teams to improve system reliability and operational excellence.


Key Responsibilities

  • Provide production support for enterprise applications within the Telematics and Connected Car ecosystem.
  • Monitor application health, system performance, and infrastructure using observability tools.
  • Manage the complete Incident Management lifecycle, including Detection, Triage, Resolution, Recovery, and Root Cause Analysis (RCA).
  • Act as the primary point of contact during production incidents and coordinate with cross-functional support teams.
  • Perform application troubleshooting, log analysis, and production issue resolution.
  • Analyze recurring issues and recommend preventive measures to improve system stability.
  • Create and maintain operational runbooks, SOPs, and knowledge base articles.
  • Collaborate with development, QA, infrastructure, cloud, and networking teams for incident resolution and deployment support.
  • Support production releases, change management activities, and post-deployment validation.
  • Drive continuous service improvements by identifying automation opportunities and improving monitoring capabilities.
  • Prepare incident reports, service health dashboards, and operational metrics.

Required Skills & Qualifications

  • Bachelor''s degree in Computer Science, Information Technology, Engineering, or a related field.
  • 8+ years of experience in Application Support, Production Support, AMS, or Site Reliability Engineering (SRE).
  • Hands-on experience working within an SRE or Incident Management team.
  • Strong understanding of the complete Incident Management Lifecycle:
    • Detection
    • Triage
    • Resolution
    • Recovery
    • Root Cause Analysis (RCA)
  • Experience with monitoring and observability tools such as:
    • Dynatrace
    • Grafana
    • ELK Stack (Elasticsearch, Logstash, Kibana)
    • Splunk
    • Prometheus
  • Experience supporting enterprise applications in production environments.
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Excellent communication and stakeholder management skills.
  • Experience working in Agile/Scrum environments.

Domain Expertise (Mandatory)

  • Experience in the Telematics or Connected Car ecosystem.
  • Understanding of Connected Vehicle architecture and production support.
  • Experience supporting vehicle connectivity platforms, telematics services, remote diagnostics, or connected mobility applications.
  • Knowledge of automotive communication concepts and connected vehicle technologies is highly preferred.

Preferred Skills

  • AWS, Azure, or Google Cloud Platform Cloud Platforms.
  • Kubernetes and Docker.
  • Linux/Unix Administration.
  • Python, Bash, or Shell Scripting.
  • CI/CD tools (Jenkins, GitHub Actions, Azure DevOps).
  • ITIL Foundation Certification.
  • ServiceNow Incident Management.
  • Automotive communication protocols (CAN, CAN FD, UDS, MQTT).
  • Knowledge of SLOs, SLIs, Error Budgets, and Reliability Engineering best practices.

Technology Environment

  • Application Management Services (AMS)
  • Site Reliability Engineering (SRE)
  • Incident Management
  • Production Support
  • Dynatrace
  • Grafana
  • ELK Stack
  • Splunk
  • Prometheus
  • ServiceNow
  • Linux / Unix
  • Python / Shell Scripting
  • AWS / Azure / Google Cloud Platform
  • Docker
  • Kubernetes
  • CI/CD
  • Git
  • Agile / Scrum
  • Telematics
  • Connected Car
  • Automotive Production Support
 
 
 

Skills

Docker
Shell
AWS
ELK
Logstash
Scrum
Splunk
Agile
Azure
Bash
GPT
Git
GitHub Actions
Google Cloud
Grafana
Jenkins
Kibana
Kubernetes
Prometheus
Python
SAFe
Stakeholder Management

Similar jobs