Reliability Engineer
Why This Role Stands Out
This hybrid Reliability Engineer role offers significant growth potential within a reputable technology company, allowing you to build cutting-edge automation and resilience solutions. You will thrive here if you possess strong infrastructure automation and reliability engineering expertise and enjoy solving complex challenges. Apply today to contribute to high-scale, high-availability services and enhance your skills in a dynamic environment.
Quick Overview
Job Description
Job Title: Reliability Engineer,
Location: Westlake, TX
Long term contract with possibility of conversation to Full time
As Senior Reliability Engineer, you blend deep Infrastructure automation experience and reliability engineering expertise with a passion for delivering results. Our Reliability Engineering group within Enterprise Infrastructure combines Operations Excellence with the Development Experience to deliver services at high scale, high availability with resilience by using automation and Infrastructure Code. We build reliability into our ecosystem by applying standards in Resiliency Engineering, Automation, Observability & Chaos Testing.
Additionally, this role contributes to enterprise backup and recovery capabilities including automation of recovery workflows, rehoused recovery into alternate datacenters, and testing of recovery processes.
We are looking for a systems thinking, reliability engineer who has helped teams scale through production insight, data and backup recovery, operational automation, developer guidance, real-time metrics, automation, automation, automation.
- Strong background in several of the following: Go, Angular, Python, JavaScript, AWS, RESTful services, Ruby, MVC, Jenkins CI/CD, Configuration Automation (Chef, Ansible).
- Preferred background in: Bootstrap, HTML/CSS, Shell Scripting, messaging frameworks (MQ), Service Oriented/Micro-service Architectures, OpenStack, Relational Databases (PostgreSQL).
- Comfortable working in both Public and private cloud environments.
- Crafting scalable solutions and automation to monitor the health and establish signals to drive understanding of our Container Platform environments.
- Strengthening operational processes with client's support and incident management teams for our cloud ecosystem
- Working with client and cloud service provider product teams and driving ongoing reliability improvements in their Kubernetes service offerings.
- Anticipating, discovering through ongoing interaction with, and prioritizing client / partner needs to serve as their voice and guide execution of the team.
The Expertise You Have
- Bachelor's Degree or equivalent experience in a technology related field (e.g. Computer Science, Engineering, etc.) required.
- Production experience running Cloud and on-prem Storage workloads at scale
- Experience managing and maintaining Kubernetes Clusters on EKS/AKS and RKS.
- Demonstrates a drive for continuous improvement and enjoys tackling complex problems.
- Experience managing and interpreting large datasets using query languages and visualization tools(PowerBI/tableau),
- Experience in software development with Python, NodeJS, or Java with a focus on SDLC and automation
- 5 -7 years of hands-on experience deploying and/or supporting highly distributed multi-tiered systems at scale.
- Experience building and deploying Docker images including Docker Compose
- Hands-on experience with Jenkins Core, including authoring and maintaining declarative CI/CD pipelines and libraries
- Experience with distributed version control systems, Git preferred
- Experience crafting and maintaining logging, monitoring, and alerting capabilities using tools like Datadog and Splunk
- Practical experience in building cloud hosted and native applications for the enterprise. Maintains a deep understanding of a wide variety of AWS/Azure services that support reliability, observability, and automation/orchestration.
- Experience in incident/crisis management and supporting critically important applications
The Skills You Bring
- Hands on experience with one or more observability tools (Prometheus, Grafana, ELK/OpenSearch, OpenTelemetry, Datadog, etc.)
- Ability to automate with various scripting languages (Python, Shell scripting, etc.)
- Experience managing systems using infrastructure as code tools (IAM, ARM, Terraform, Chef)
Additional Value in Backup & Recovery:
- Advance enterprise resiliency through improved recovery capabilities.
- Reduce recovery time via automation.
- Enable rehoused recovery into new datacenters.
- Strengthen platform reliability through data protection design.
Dexian stands at the forefront of Talent + Technology solutions with a presence spanning more than 70 locations worldwide and a team exceeding 10,000 professionals. As one of the largest technology and professional staffing companies and one of the largest minority-owned staffing companies in the United States, Dexian combines over 30 years of industry expertise with cutting-edge technologies to deliver comprehensive global services and support.
Dexian connects the right talent and the right technology with the right organizations to deliver trajectory-changing results that help everyone achieve their ambitions and goals. To learn more, please visit .
Dexian is an Equal Opportunity Employer that recruits and hires qualified candidates without regard to race, religion, sex, sexual orientation, gender identity, age, national origin, ancestry, citizenship, disability, or veteran status.
Skills
Similar jobs
Node.js Developer
Everest Global Solutions · Dallas, United States
11 minutes agoFull Stack Developer
Applied Thought Auditors & Consultants Inc. · Alpharetta, United States
11 minutes agoSoftware Engineer - DSP (Full Scope Poly) with Security Clearance
Shadowgate Partners, Inc. · Fairfax, United States
11 minutes ago$105k - $170k/yrSalesforce Developer
Innosoul inc · Raleigh, United States
11 minutes agoPython Developer
Sincera Technologies, Inc. · New York, United States
11 minutes agoSolution Architect(Telecom OSS, Mainframe, Gen AI)
Yochana IT Solutions · United States
11 minutes ago