SRE Site Reliability- (Support & Operations Engineer)
Quick Overview
Job Description
Title: SRE Site Reliability- (Support & Operations Engineer)
Location: BTC - 11450 Grooms Road, Blue Ash, OH, 45242, US
Number of Open Positions: 6 open / 6
Work Location Preferences: Onsite
Visa: H-Four and L2EAD & Independent visa
Job Description
- Level 3 (LV3) Support, Site Reliability, and Operations Engineer supporting Kroger's Next Generation Point of Sale (NGPOS) platform
- This team is embedded within the engineering organization, partnering directly with developers throughout the SDLC rather than sitting in a separate support silo
- We care more about problem-solving instinct, curiosity, and drive than a checklist of prior experience the items below describe a strong match, not a hard bar, and candidates early in their career or without every listed skill are encouraged to apply
- In office, Blue Ash, OH 5 days a week; willingness to travel and provide onsite support for new go-lives, pilots, and rollouts
- Participate in a scheduled on-call rotation (nights/weekends/holidays) for 24x7 support; PTO/leave may be restricted during peak retail periods and major go-lives
- Bachelor's degree in Computer Science, Information Systems, or a related field (or equivalent experience) preferred; prior experience in application support, operations, or engineering is a plus but not required
- Exposure to Point-of-Sale systems in enterprise environments is helpful; some experience with Java is preferred, Go is a plus
- Any exposure to SQL or NoSQL databases (e.g., MongoDB), scripting (Shell, PowerShell, or Python), or ITSM/monitoring tooling (ServiceNow, Jira, Confluence, Splunk, Grafana, or equivalent) is a plus willingness to learn these on the job matters more than prior mastery
- Interest in or willingness to learn payment card security standards (PCI-DSS)
- Excellent oral and written communication skills; able to translate between business users and Kroger Technology teams
- Familiarity with the Agile process is helpful; experience as a Scrum Master a plus
- Willingness to speak up and challenge developers on Agile best practices and definitions of done during refinement and QA
- Accountable for driving support tickets and emails to resolution
- Good judgment under pressure; able to prioritize across competing support, testing, and scrum-master duties
Key Responsibilities
- Partner with cross-functional teams to expedite issue resolution and manage to SLAs
- Support and maintain infrastructure and applications, including off-hours support (24 x 7) as required
- Test infrastructure and application changes
- Establish priorities and develop functional and programming specifications for application enhancements and modifications
- Participate in the application technical design process; design, code, and unit test application changes using SDLC best practices
- Complete estimates and work plans for design, development, implementation, and rollout tasks
- Monitor systems, consoles, and performance for service interruptions or delays
- Address and resolve infrastructure system failures
- Execute, log, and report break/fix changes, service requests, and support activities
- Maintain operational procedures, processes, and scripts; follow documented processes to ensure infrastructure stability
- Own incident and problem management: drive major incident response, root cause analysis, and post-incident reviews
- Champion engineering standards and continuously improve software delivery processes
- Build partnerships across application, business, and infrastructure teams
- Independently execute large projects and lead other analysts in completing projects
Additional Notes from the Manager:
Candidates who have just Windows systems experience and not Linux are still able to be considered from this role. Again, candidates do not need to check every single box to be considered, however, the priority skillsets must be fulfilled.
Top 3 skills:
- Major Incident Management experience has actually run or driven a bridge/war-room for a P1/P2 outage, not just "participated." This is the core of the role; everything else is supporting it. Ask for a specific incident they led end-to-end.
- RCA/Problem Management rigor can articulate a structured RCA methodology (5-whys, fishbone, timeline reconstruction) and, more importantly, examples where they drove a corrective action to closure (not just wrote a doc that nobody actioned).
- Production troubleshooting under pressure across a mixed stack comfortable reading logs/dashboards (Dynatrace-type tooling) and reasoning about cloud + on-prem + store/POS systems simultaneously, since a real incident here will span all three.
Additional Information
- Q: Is this position working hands-on with POS systems (the JD mentions maintaining infrastructure), or is it primarily working in the office to develop and support the POS platform?
- A: This is a Level 3 Support Engineer role within the NGPOS engineering organization, supporting the software and infrastructure behind the POS platform. There is a hands-on hardware component when working in our engineering lab or onsite supporting stores, but the core of the role is software/systems support, not hardware development.
- Q: Roughly what percentage of the role is production support/incident response versus software engineering/development?
- A: This is not a developer role. Prior development experience is not required, though it can be an asset when troubleshooting complex issues. We've had both developers and non-developers be very successful in this position.
- Site Reliability Engineers will be successful but must understand this is a support role
- Still lots of engineering involved, will be working in the code every day
- Looking for a solid and motivated technologist
- Candidate does not need to tick all skills boxes, but does need to be able to be onsite and travel
Are there any automatic disqualifiers?
- No on call /off-hours availability
- Pure DevOps/CI-CD background with zero incident-response exposure
Travel Expectations
- Lots of flexibility, it will ebb and flow with needs
- Currently more local travel (Cincy and Louisville Stores) but as new sites go live, travel will expand
- Will be traveling every 4 weeks or so
- Travel and on call is split between 6 people - and expecting team to grow
On Call/ Off Hours Expectations
- Split between 6 people and rotates weekly
- Team is currently splitting 2 shifts a day
- Periodic off hours during go lives
- This team is an escalation point - they are not fielding everything
- There is flexibility and team makes sure to balance personal lives
Work location
- In office 5-days a week at BTC location to work in lab
- Might be relaxed on onsite later down line once travel picks up
- Relocation is fine - Candidate can start remotely and relocate within 1 month of onboarding, will most likely still have to travel during that time
Skills
Similar jobs
Software Development Engineer - SRE, Medicare Sales
CVS Health · United States
21 minutes agoGitLab DevOps Architect | 100% Remote
Montek System · United States
44 minutes agoSenior DevOps Engineer
Raas Infotek LLC · Atlanta, United States
44 minutes agoDevOps Engineer
Technotopia Solutions LLC · United States
46 minutes agoLead, Software Engineer - DevOps Architect - TS/SCI with Poly with Security Clearance
L3Harris Technologies · Palm Bay, United States
47 minutes agoRelease Train Engineer (RTE)/Scrum Master/DevOps Concepts- Only W2
Info Dinamica Inc · Irving, United States
50 minutes ago