Why This Role Stands Out
This role offers a fantastic opportunity to gain deep expertise in connected car systems and hone your incident management skills within a reputable automotive player. You'll thrive here if you're a proactive problem-solver with a background in telematics, ready to contribute to critical 24/7 operations. Apply to become an integral part of this dynamic team!
Quick Overview
Seniority
Mid Senior
Work mode
On Site
Location
Irvine, CA, United States
Posted
Yesterday
ConfluenceJiraStakeholder ManagementTelematics
Job Description
Job Title: Incident Support Engineer (Automotive background)
Work Location : Irvine, CA (Onsite)
Contract duration: 12+ months Contract
Work Location : Irvine, CA (Onsite)
Contract duration: 12+ months Contract
Job type: Contract
Role Summary
We are seeking Support Engineer role for a 24/7 Incident Management (IM) operations program supporting a key player in the automotive industry's connected car space. This role will primarily focus on monitoring the system health using different data monitoring tools such as Datadog, Grafana, MaxGauge, Elastic, Dynatrace. Once an anomaly or unusual pattern is observed, escalate it to the respective teams, initiate an incident bridge, driving the bridge, taking incident timeline, facilitate the teams to bring an ongoing incident to closure and ensuring adherence to contractual SLAs. The IM Support Engineer will orchestrate end-to-end incident handling—from anomaly detection to bridge initiation, stakeholder communications, resolution, and post-incident governance—with a strong background in Telematics and Connected Vehicle systems (including Remote Services).
Key Responsibilities
· Monitor the health and availability of production systems supporting the Connected Car / Telematics ecosystem.
· Proactively monitor application, infrastructure, API, database, and service health using tools such as:
· Datadog
· Grafana
· MaxGauge
· Application and system logs
· Organization-specific monitoring and observability tools
· Identify abnormal system behavior, performance degradation, errors, latency, service failures, and potential outages.
· Analyze dashboards, alerts, metrics, logs, and application behavior to identify the initial scope and impact of incidents.
· Perform proactive monitoring to identify issues before they impact customers.
· Validate alerts and distinguish between genuine production incidents and false positives/noise.
· Adhere to the IM SLAs (MTTD, MTTA, MTTR, communication SLAs); ensure measurement, reporting, and continuous improvement.
· Should be proficient in handling the incidents ranging from high priority ones (P1, P2) to the low priority incidents (P3, P4).
· Initiate and run incident bridges/war rooms; coordinate cross-functional responders (application, infrastructure, network, OEM partners, and third parties).
· Oversee proactive detection through data monitoring tools (Datadog, Dynatrace, Grafana, MaxGauge) and ensure alert quality, runbooks, and signal-to-noise optimization.
· Establish and enforce SOPs for incident declaration, severity classification, response roles and decision logs.
· Drive disciplined stakeholder communications: timely updates to product, operations, OEM/customer contacts, leadership, and impacted regions; maintain comms cadence and channels.
· Ensure post-incident governance: facilitate RCA, document contributing causes, corrective and preventive actions (CAPA), and circulate the RCA across teams for sign-off.
· Define and maintain IM dashboards, KPIs, and executive reports; present weekly/monthly service reviews with trend analysis and action plans.
· Collaborate with Product/Engineering to influence reliability roadmaps (resiliency patterns, observability, capacity, release safeguards).
· Ensure compliance with information security, data privacy, and OEM contractual obligations during incident handling and communications.
· Continuously refine IM playbooks, runbooks, and training; conduct simulations/game days and readiness audits across onsite–offshore teams.
Domain Expertise: Telematics & Connected Car
· Hands-on knowledge of Telematics Control Unit (TCU), eSIM/OTA provisioning, backend telematics platforms, and data flows between vehicle, cloud, and mobile apps.
· Understanding on Internet of Things including but not limited to Messaging Queues, bulk provisioning, notification systems, API Gateways, Load Balancers, Mobile applications & web applications.
· Familiarity with Remote Services (e.g., remote lock/unlock, start/stop, charge control, climate pre-conditioning), geo-services, and safety/assist features.
· Understanding of service dependencies: identity/auth, messaging, device management, CAN bus signals, firmware/OTA update orchestration, and regional compliance.
· Experience coordinating incidents across OEM partners, Tier-1 suppliers, cloud providers, and customer support operations.
Required Qualifications
· 5+ years in Operations/Service Management with 4+ years leading Incident Management in large-scale, 24/7 environments.
· Demonstrated experience running bridges for P1/P0 incidents; proven incident commander skills and decision-making under pressure.
· Strong background in Telematics/Connected Car domain and vehicle remote services concepts.
· Proficiency with observability and monitoring tools: Dynatrace, Datadog, Grafana, MaxGauge (dashboards, alerting, traces, logs).
· Expertise in IM processes: detection → triage → severity assignment → bridge initiation → stakeholder updates → resolution → post-incident RCA.
· Working knowledge of ITIL practices (Incident, Problem, Change, Service Level Management).
· Should have good hands on experience in using tools like JIRA, Confluence, Jenkins, XMatters.
· Excellent communication (written/verbal), executive presence, and stakeholder management across onsite–offshore teams.
· Ability to analyze telemetry and time-series data to drive root cause hypotheses and corrective actions.
· Experience defining SLAs/SLOs and building KPI dashboards (MTTD/MTTA/MTTR, incident volume, recurrence, comms SLA, customer impact).
Tools & Technologies
· Datadog, Dynatrace , Grafana, MaxGauge (APM, logs, metrics, dashboards).
· Incident & Comms: ticketing (Jira/ServiceNow), chat/bridge tools (Teams/Zoom), status pages.
· Data: time-series analysis, log aggregation, tracing, alerting policies and noise reduction.
Similar jobs
- CA
Automotive Sales - Join Our Team
NewCrain Automotive
Jacksonville, Arkansas🇺🇸On-site30 minutes agoAutomotive - ED
Automotive Technician
NewElevate Digital
Rockville, MD🇺🇸$50/hrOn-site23 hours agoAutomotive - AT
Telecommunications Mechanic I with Security Clearance
NewAbacus Technology
Charleston, SC🇺🇸Hybrid23 hours agoAutomotive - DO
ELECTRONIC INTEGRATED SYSTEMS MECHANIC with Security Clearance
Department of the Air Force
Egg Harbor, NJ🇺🇸Remote7 weeks agoAutomotive - ST
Salesforce Technical Project Manager (Automotive Cloud)
NewSRI Tech Solutions
Atlanta, GA🇺🇸Hybrid23 hours agoScrumAgileStakeholder ManagementAutomotive - RD
Program Manager (Automotive industry experience)
NewRandstad Digital
Newark, CA🇺🇸$65 - $75/hrHybrid23 hours agoSAFeScrumAgile+3Automotive