Haystack
← Back to Jobs
Manufacturing
VS

Production Support Engineer

VBEST Software IncChandler, AZ🇺🇸United StatesPosted 28 Aug 2026

Why This Role Stands Out

This hybrid role at VBEST Software Inc. offers exciting challenges and significant growth potential for experienced Production Support Engineers who thrive on end-to-end incident ownership and root cause analysis. You'll hone your skills in Unix, SQL, and critical monitoring tools while contributing to a reputable company. Apply today to embrace this opportunity for impactful work and continuous development.

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Chandler, AZ, United States
Posted
3 days ago

Job Description

Job Summary
Team and role context
  • This is a live incident production support seat, not a pure runbook follow role. Team monitors batch, feeds, and application health across Unix, SQL, Dynatrace, Splunk, and ServiceNow, and owns incidents end to end including root cause.
  • Candidates need to be comfortable being handed a vague scenario, for example an application down two hours after a change went in, and walking an interviewer through live triage, not reciting a memorized process.
Required
  • 3 to 5 years hands on application production support
  • Strong Autosys
  • Strong Unix and Shell scripting
  • Working SQL, Oracle, Hadoop or similar DBMS
  • Ability to juggle and prioritize multiple concurrent issues
  • Fast independent learner
Required, confirmed from team member
  • Command level Unix fluency, not tool familiarity. Interviewers ask for the actual syntax: du and df for disk space, top and uptime for system load, find with -type and -mtime flags for locating files by age. A candidate who can describe what a command does but can't produce the flags will get caught here.
  • Ability to distinguish failure types precisely. Interviewers specifically probe whether a candidate conflates a data issue with a Unix file system issue with a database connection pool issue. These are three different diagnostic paths and the interviewer will correct and re-ask if a candidate blurs them.
  • SQL depth beyond basic querying: index behavior and why a query runs slow, truncate versus delete, join types used in production troubleshooting, not just writing selects.
  • Autosys job state knowledge beyond scheduling, specifically the difference between a job marked inactive versus on hold versus on ice, and why a team would use each.
  • Splunk and Dynatrace used together, and candidates should be able to explain why a team runs both rather than just one. Splunk gets tested as a log search and correlation tool, not described as monitoring.
  • Noisy alert management. Interviewers ask how a candidate would handle an alert threshold that's firing too often, for example tuning a CPU alert from 70 percent to 85 percent to cut false positives.
  • Incident severity fluency: candidates should be able to define P1 through P4 without hesitation and describe how communication changes at each level, including whether they'd post to a status page versus direct email during an outage.
  • Basic API and auth troubleshooting exposure, specifically JWT or token related login failures, came up as a possible scenario topic.
  • Shell scripting for automation of repetitive work, for example disk cleanup or log rotation scripts, should be a real example the candidate has done, not something on the resume without a story behind it.

Similar jobs