Why This Role Stands Out
This Senior Site Reliability Engineer role offers a fantastic opportunity to drive innovation in platform operations for a leading financial services technology organization, with a competitive hourly rate of $75-$80. You'll thrive here if you're passionate about automating complex systems, leveraging AI/ML for proactive problem-solving, and contributing to resilient, large-scale distributed architectures. Embrace this chance to elevate your skills in a hybrid environment that fosters impactful work and professional growth.
Quick Overview
Seniority
Mid Senior
Work mode
On Site
Location
Austin, TX, United States
Posted
Yesterday
MongoDBOracleSQLAWSScrumSplunkAgileAzureBashDNSDatadog.NETGitHub ActionsGoogle CloudGrafanaJavaJenkinsKafkaKubernetesPowerShellPythonRabbitMQ
Job Description
Location: Austin, TX Salary: $75.00 USD Hourly - $80.00 USD Hourly Description:
Role: Senior Site Reliability Engineer / DevOps Engineer (AIOps & Observability)
Work Arrangement: Hybrid (4 days weekly on-site in Austin, TX)
Position Type: 6+ Month Contract
About the Role
Our client, a premier financial services and technology organization, is seeking an experienced Site Reliability Engineer (SRE) / DevOps Engineer to modernize platform operations across enterprise cloud and authentication ecosystems. In this role, you will bridge software engineering and systems operations to enhance platform availability, eliminate operational toil, and advance AI/ML-driven reliability engineering.
You will champion an SRE mindset across mission-critical platforms, architecting end-to-end automation, predictive alerting, anomaly detection, and self-healing systems. Partnering with cross-functional engineering, infrastructure, and delivery teams, you will ensure high availability and resilient delivery across large-scale distributed architectures.
Key Responsibilities
Minimum Qualifications
Preferred Qualifications
Medical, dental, and vision insurance are available to qualified candidates who meet eligibility requirements.
By providing your phone number, you consent to: (1) receive automated text messages and calls from the Judge Group, Inc. and its affiliates (collectively "Judge") to such phone number regarding job opportunities, your job application, and for other related purposes. Message & data rates apply and message frequency may vary. Consistent with Judge's Privacy Policy, information obtained from your consent will not be shared with third parties for marketing/promotional purposes. Reply STOP to opt out of receiving telephone calls and text messages from Judge and HELP for help.
Contact:
This job and many more are available through The Judge Group. Please apply with us today!
Role: Senior Site Reliability Engineer / DevOps Engineer (AIOps & Observability)
Work Arrangement: Hybrid (4 days weekly on-site in Austin, TX)
Position Type: 6+ Month Contract
About the Role
Our client, a premier financial services and technology organization, is seeking an experienced Site Reliability Engineer (SRE) / DevOps Engineer to modernize platform operations across enterprise cloud and authentication ecosystems. In this role, you will bridge software engineering and systems operations to enhance platform availability, eliminate operational toil, and advance AI/ML-driven reliability engineering.
You will champion an SRE mindset across mission-critical platforms, architecting end-to-end automation, predictive alerting, anomaly detection, and self-healing systems. Partnering with cross-functional engineering, infrastructure, and delivery teams, you will ensure high availability and resilient delivery across large-scale distributed architectures.
Key Responsibilities
- SRE & Reliability Engineering: Advocate for SRE principles, systematize operational workflows, and own production automation frameworks to eliminate repetitive toil and boost operational efficiency.
- AIOps & Intelligent Observability: Design and deploy AI/ML-assisted observability, predictive alerting mechanisms, and anomaly detection models to preempt performance bottlenecks and outages across core platforms.
- Proactive Monitoring & Incident Response: Construct enterprise dashboards, metrics, and alert policies using tools such as Splunk and AppDynamics; triage mission-critical incidents, perform root-cause analysis, and participate in rotational on-call support.
- Automation & Scripting: Engineer robust automation scripts, tools, and runbooks using Python, Bash, PowerShell, Java, or .NET for operational remediation and infrastructure consistency.
- CI/CD & Delivery Orchestration: Develop automated software delivery pipelines, advance GitOps methodologies, and integrate AI/ML-driven checks to optimize deployment reliability.
- Infrastructure & Platform Operations: Administer, tune, and troubleshoot virtualized Windows Server (2019/2022) and Linux environments, messaging infrastructures, and cloud/platform runtimes.
- Capacity & Rollout Strategy: Direct capacity planning leveraging predictive analytics and telemetry trend models; ensure safe deployment validation across large-scale distributed platforms.
Minimum Qualifications
- 6-8 years of experience in enterprise systems administration, production operations, and platform reliability engineering.
- 6-8 years of experience building proactive observability dashboards, log aggregation pipelines, and alerting frameworks (e.g., Splunk, AppDynamics, Grafana, Datadog).
- 6-8 years of experience operating within formal SDLC processes, CI/CD workflows, and continuous improvement frameworks.
- Practical systems administration, troubleshooting, and performance tuning experience across both Linux and virtualized Windows Server (2019/2022) environments.
- Hands-on scripting and programming capabilities in one or more languages: Python, Bash, PowerShell, Java, or .NET.
- Working knowledge of enterprise messaging brokers (Kafka, RabbitMQ, Solace, or IBM MQ).
- Familiarity with database systems (SQL, Oracle, or MongoDB).
- Working knowledge of fraud prevention or compliance platforms, specifically Actimize.
- Solid understanding of core IP networking principles (DNS, DHCP, routing, firewalls) and high-availability distributed systems.
- Demonstrated experience implementing AI/ML, AIOps, or anomaly-detection approaches within production monitoring or operational automation.
- Bachelor's degree in Computer Science, Information Technology, or a related discipline (or equivalent practical experience).
Preferred Qualifications
- Prior experience supporting mission-critical platforms within the financial services or banking sector.
- Working experience configuring, deploying, and supporting applications on Google Cloud Platform (Google Cloud Platform) or Pivotal Cloud Foundry (PCF).
- Experience with enterprise container orchestration platforms (Kubernetes, OpenShift) or multi-cloud environments (AWS, Azure, Google Cloud Platform).
- Familiarity with modern CI/CD orchestration tools (Harness, Jenkins, GitHub Actions) and GitOps practices.
- Experience operating within Agile/Scrum delivery frameworks.
Medical, dental, and vision insurance are available to qualified candidates who meet eligibility requirements.
By providing your phone number, you consent to: (1) receive automated text messages and calls from the Judge Group, Inc. and its affiliates (collectively "Judge") to such phone number regarding job opportunities, your job application, and for other related purposes. Message & data rates apply and message frequency may vary. Consistent with Judge's Privacy Policy, information obtained from your consent will not be shared with third parties for marketing/promotional purposes. Reply STOP to opt out of receiving telephone calls and text messages from Judge and HELP for help.
Contact:
This job and many more are available through The Judge Group. Please apply with us today!
Similar jobs
- BA
Senior Blockchain DevOps Engineer - AVP - Digital Assets
NewBarclays
New Jersey🇺🇸Hybrid8 minutes agoEthereumBlockchainKubernetes+1Technology - OR
Hardware Platform Engineer - Data Center Systems & Rack Integration
NewOracle
United States🇺🇸$114.6k - $234.6k/yrHybrid8 minutes agoOracleTechnology - MA
DevSecOps Platform Engineer
NewMANTECH
Tampa, Florida🇺🇸Hybrid8 minutes agoDockerGCPShell+12Technology - IT
Senior Site Reliability Engineer / Backup Engineer
NewICON Technologies
Chicago, IL🇺🇸HybridYesterdayAWSELKEncryption+16Technology - JG
Site Reliability Engineer (SRE) - Mid-Level
NewJudge Group, Inc.
Southlake, TX🇺🇸On-siteYesterdayAWSSplunkAnsible+8Technology - MR
Site Reliability Engineer
NewMotion Recruitment Partners, LLC
Chicago, IL🇺🇸HybridYesterdayAWSAnsibleJenkins+2Technology