Haystack
← Back to Jobs
Other
CL

AI Engineers/Architects with experience on Resiliency & Observability Engineering

CloudiousTX🇺🇸United StatesPosted 10 Sept 2026

Why This Role Stands Out

This role offers a fantastic opportunity to drive enterprise-wide modernization and build cutting-edge AI-led resiliency capabilities, leveraging your expertise in Java full-stack development, SRE, and observability. If you're a mid-senior engineer passionate about enhancing system reliability and eager to shape the future of cloud services, you'll thrive in this impactful position. Explore this chance to grow your skills and contribute to significant initiatives within a reputable company.

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
TX, United States
Posted
5 days ago
Microservices.NETJavaJiraRoot Cause Analysis

Job Description

We have a priority requirement to identify 2 onshore ( Dallas / Austin )

Senior Agentic AI Engineers/Architects with experience on Resiliency & Observability Engineering

Senior Agentic AI Engineers/Architects Resiliency & Observability Engineering | Java Full Stack

Program Objective - Build AI-led resiliency capabilities and drive long-term modernization, observability, and reliability engineering initiatives across enterprise services.

Key Skill requirements:

  • Strong hands-on experience with Java, Full Stack Development, Microservices, and SDLC practices. .Net knowledge is desired too
  • Experience on AI/Agentic AI, SRE, Reliability Engineering, Observability, and Production Engineering
  • Experience identifying service resiliency gaps and driving reliability improvements.
  • Hands-on exposure to observability, production support, and incident analysis.
  • Performance engineering and Java multithreading experience is a plus.
  • Excellent troubleshooting and problem-solving skills.
  • Strong database programming and query optimization skills.
  • Ability to perform deep root cause analysis using application logs, source code, system events, and monitoring tools.
  • Ability to work independently with limited upfront requirements and perform technical analysis to identify improvement opportunities.
  • Ability to translate findings into actionable engineering initiatives and create Jira stories/backlogs to improve application resiliency, reliability, and operational efficiency.
  • Experience working closely with engineering teams to drive modernization and stability improvements across services.

Similar jobs