Haystack
← Back to Jobs
Manufacturing
KA

Resiliency AI Architect with .NET (AI/Agentic AI, Reliability Engineering, SRE, Observability, Production Engineering.NET Development Background) : Local to Texas

K Anand CorporationDallas, TX🇺🇸United StatesPosted Oct 5, 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
Dallas, TX, United States
Posted
19 hours ago
JiraRoot Cause Analysis

Job Description

Position      : Resiliency AI Architect with .NET (AI/Agentic AI, Reliability Engineering, SRE, Observability, Production Engineering.NET Development Background) : Only USA
Location      : Dallas, TX/ Austin, TX- Onsite from Day 1
Type of Job      : Contract : 12 Months
Experience     : 10+ Years

strong focus on Application Resiliency, AI-led modernization, and .NET architecture.

What the Client Is Looking For
The client is looking for someone who can guide the team as an architect while also being hands-on in identifying and implementing resiliency improvements in existing enterprise applications.
The environment is a high-transaction trading/financial environment, involving transactions such as wire transfers, where availability, reliability, and resiliency are critical.

Key Areas to Prepare
•    AI-led Application Resiliency – Explain how you can use AI to analyze existing applications, identify reliability issues, and recommend improvements.
•    Agentic AI / AI Tools – Be prepared to discuss how AI agents or tools can analyze multiple repositories, code, logs, incidents, and system dependencies.
•    .NET Architecture – Strong understanding of enterprise .NET applications, microservices, APIs, architecture patterns, and modernization.
•    Reliability & High Availability – Be ready to explain how you would improve application availability and move toward a higher SLA.
•    Observability – Logs, metrics, traces, monitoring, alerting, dashboards, distributed tracing, and identifying failure patterns.
•    SRE Concepts – SLI, SLO, SLA, error budgets, incident management, RCA, MTTR, fault tolerance, and graceful degradation.
•    Kafka – Be prepared to discuss Kafka reliability, consumer failures, retries, dead-letter topics, partitioning, replication, and recovery.
•    Database Resiliency – Availability, performance, connection failures, replication, failover, query optimization, and recovery strategies.
•    Performance & Scalability – Identify bottlenecks and explain how you would improve throughput and response time.
•    Production Engineering – Real-world examples of troubleshooting production issues and implementing permanent fixes.
•    Modernization – How you would assess an existing application and create a phased roadmap for improving resiliency.

Key Skills (Priority Order)
•    Strong hands-on experience in .NET development (primary skill), including enterprise applications, microservices, APIs, and SDLC best practices.
•    Experience with AI/Agentic AI, Reliability Engineering, SRE, Observability, and Production Engineering.
•    Proven ability to identify resiliency gaps and drive improvements in availability, reliability, scalability, performance, and operational efficiency.
•    Hands-on experience implementing resiliency patterns such as retries, timeouts, circuit breakers, failover, and idempotency.
•    Strong expertise in application performance engineering, troubleshooting, root cause analysis, and production support.
•    Experience with monitoring, alerting, logging, dashboards, telemetry, and service reliability metrics (SLIs/SLOs).
•    Experience with cloud-native architectures, Kafka or other messaging platforms, and CI/CD pipelines.
•    Ability to assess existing applications and drive modernization, maintainability, observability, security, and stability improvements.
•    Experience translating assessment findings into actionable engineering initiatives, Jira stories, and implementation backlogs.
•    Strong database, analytical, and problem-solving skills.
•    Java experience is desirable to support a mixed .NET and Java technology landscape.

Similar jobs