Why This Role Stands Out
This hybrid role at VBEST Software Inc. offers exciting challenges and significant growth potential for experienced Production Support Engineers who thrive on end-to-end incident ownership and root cause analysis. You'll hone your skills in Unix, SQL, and critical monitoring tools while contributing to a reputable company. Apply today to embrace this opportunity for impactful work and continuous development.
Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
Chandler, AZ, United States
Posted
3 days ago
Job Description
Job Summary
Team and role context
- This is a live incident production support seat, not a pure runbook follow role. Team monitors batch, feeds, and application health across Unix, SQL, Dynatrace, Splunk, and ServiceNow, and owns incidents end to end including root cause.
- Candidates need to be comfortable being handed a vague scenario, for example an application down two hours after a change went in, and walking an interviewer through live triage, not reciting a memorized process.
Required
- 3 to 5 years hands on application production support
- Strong Autosys
- Strong Unix and Shell scripting
- Working SQL, Oracle, Hadoop or similar DBMS
- Ability to juggle and prioritize multiple concurrent issues
- Fast independent learner
Required, confirmed from team member
- Command level Unix fluency, not tool familiarity. Interviewers ask for the actual syntax: du and df for disk space, top and uptime for system load, find with -type and -mtime flags for locating files by age. A candidate who can describe what a command does but can't produce the flags will get caught here.
- Ability to distinguish failure types precisely. Interviewers specifically probe whether a candidate conflates a data issue with a Unix file system issue with a database connection pool issue. These are three different diagnostic paths and the interviewer will correct and re-ask if a candidate blurs them.
- SQL depth beyond basic querying: index behavior and why a query runs slow, truncate versus delete, join types used in production troubleshooting, not just writing selects.
- Autosys job state knowledge beyond scheduling, specifically the difference between a job marked inactive versus on hold versus on ice, and why a team would use each.
- Splunk and Dynatrace used together, and candidates should be able to explain why a team runs both rather than just one. Splunk gets tested as a log search and correlation tool, not described as monitoring.
- Noisy alert management. Interviewers ask how a candidate would handle an alert threshold that's firing too often, for example tuning a CPU alert from 70 percent to 85 percent to cut false positives.
- Incident severity fluency: candidates should be able to define P1 through P4 without hesitation and describe how communication changes at each level, including whether they'd post to a status page versus direct email during an outage.
- Basic API and auth troubleshooting exposure, specifically JWT or token related login failures, came up as a possible scenario topic.
- Shell scripting for automation of repetitive work, for example disk cleanup or log rotation scripts, should be a real example the candidate has done, not something on the resume without a story behind it.
Similar jobs
- IC
On-call Software Engineer, Cloud Data Platform
NewICF Consulting Group, Inc.
Reston, VA🇺🇸$98.6k - $167.6k/yrOn-siteYesterdaySQLSQL ServerAWS+8Technology - UP
Senior Software Engineer - Machine Learning Platform
NewUpstart
United States🇺🇸$166.9k - $230k/yrRemoteYesterdayAWSMLOpsMLflow+6Technology - HA
Senior Specialist, Systems Engineer
Harris Corporation
Greenville, TX🇺🇸Hybrid2 weeks agoTechnology - PN
Sr Staff Engineer Software
NewPaloAlto Networks
Boston, MA🇺🇸$126k - $204.5k/yrOn-siteYesterdayExpressMongoDBNode.js+6Technology - AS
Software Engineer (Data & AI)
Apex Systems
MD🇺🇸Hybrid4 weeks agoSAFeDockerSpring+13Technology - HA
Senior Associate, Software Engineering
Harris Corporation
Rochester, NY🇺🇸$74.5k - $138.5k/yrHybrid2 weeks agoMATLABEncryptionFPGA+3Technology