Why This Role Stands Out
This role offers significant growth potential as you architect and manage scalable cloud infrastructure, championing DevOps and SRE best practices to drive rapid software delivery. You'll thrive here if you're a collaborative technologist passionate about building robust CI/CD pipelines and bridging development and operations. Apply today to contribute to Qentelli's innovative technology solutions.
Quick Overview
Job Description
Title: Sr. DevOps Engineer
Location: Richardson, TX (5 days onsite)
Duration: 12+ months
Objectives of this Role
Design, build, and maintain scalable, secure, and resilient CI/CD pipelines that support rapid, reliable software delivery across all engineering teams.
Architect and manage cloud infrastructure on AWS and Azure using Infrastructure as Code (IaC) principles, ensuring environments are reproducible, auditable, and cost-optimized.
Establish and enforce DevOps best practices, standards, and guardrails across the software development lifecycle — from code commit through production deployment.
Serve as a technical bridge between development and operations, understanding application architecture, dependencies, and requirements to design optimal deployment strategies.
Champion a Site Reliability Engineering (SRE) mindset: define SLOs, SLIs, and error budgets, and drive improvements in system availability, performance, and incident response.
Collaborate with security teams to embed DevSecOps practices into pipelines, ensuring vulnerability scanning, compliance validation, and secrets management are integral — not afterthoughts.
Mentor and provide technical leadership to development and operations teams on cloud-native patterns, automation strategies, operational excellence, and best practices.
Drive platform engineering initiatives that reduce toil, improve developer experience, and accelerate time-to-market for new features and services.
Daily and Monthly Responsibilities
CI/CD & Automation
Design, implement, and maintain end-to-end CI/CD pipelines using GitHub Actions, Jenkins, or equivalent tools, supporting multi-environment deployments (dev, staging, production).
Automate build, test, security scan, and deployment workflows to eliminate manual intervention and reduce deployment risk.
Develop and maintain pipeline-as-code templates and reusable workflow libraries for use across all engineering teams.
Implement automated rollback mechanisms, blue/green and canary deployment strategies to minimize downtime and deployment risk.
Infrastructure & Cloud Engineering
Develop, deploy, and manage cloud infrastructure on AWS and Azure using Terraform and Ansible, ensuring environments are version-controlled and auditable.
Design and operate highly available, fault-tolerant architectures leveraging AWS services (EC2, EKS, ECS, Lambda, RDS, S3, VPC, Route53, CloudFront) and Azure equivalents.
Manage Kubernetes clusters (EKS/AKS), including cluster provisioning, node group management, RBAC configuration, network policy enforcement, and autoscaling.
Build and manage serverless architectures using AWS Lambda, API Gateway, and event-driven patterns.
Govern cloud cost management through resource tagging strategies, right-sizing recommendations, and FinOps practices.
Security & Compliance (DevSecOps)
Integrate security tooling (SAST, DAST, SCA, container scanning) directly into CI/CD pipelines to enforce security gates before code reaches production.
Manage secrets and credentials using AWS Secrets Manager, HashiCorp Vault, or Azure Key Vault — enforcing zero hard-coded secrets policies.
Implement and maintain identity and access management (IAM) policies following the principle of least privilege across all cloud environments.
Ensure infrastructure and application compliance with regulatory standards (SOC 2, PCI-DSS, FFIEC) through automated policy-as-code tools such as OPA or AWS Config Rules.
Conduct regular security audits of cloud environments, pipeline configurations, and container images.
Monitoring, Observability & Incident Management
Build and maintain comprehensive observability stacks covering metrics, logs, and distributed traces using tools such as Datadog, Prometheus, Grafana, ELK/OpenSearch, or AWS CloudWatch.
Define and manage alerting strategies, escalation policies, and on-call runbooks to ensure rapid incident detection and resolution.
Lead post-incident reviews (PIRs), drive root cause analysis, and implement preventive measures to eliminate repeat incidents.
Establish and track SLOs and error budgets for critical services, using data to prioritize reliability investments.
Containerization & Microservices
Design, build, and maintain containerized application environments using Docker and Kubernetes, ensuring optimal resource utilization, security hardening, and operational reliability.
Implement and manage service mesh solutions (Istio, AWS App Mesh) to govern inter-service communication, traffic management, and mTLS.
Maintain Helm charts and Kubernetes manifests as version-controlled artifacts across environments.
Support teams in migrating monolithic workloads to cloud-native, microservices-based architectures.
Collaboration & Platform Engineering
Partner with software engineers, architects, and product teams to design systems that meet functional, performance, and reliability requirements from day one.
Review application code and architecture to understand deployment requirements, dependencies, and optimization opportunities.
Build and maintain internal developer platforms (IDPs) and self-service tooling that reduce cognitive load and accelerate delivery.
Participate in architecture reviews, sprint planning, and retrospectives as a key engineering contributor with both operational and development perspectives.
Educate and coach development teams on cloud-native patterns, DevOps best practices, operational responsibilities, and performance optimization.
Collaborate on code review processes to ensure deployability, security posture, and operational considerations are addressed early.
Required Skills and Qualifications
Education & Experience
Bachelor''s degree in Computer Science, Information Technology, Software Engineering, or a related technical discipline.
5+ years of hands-on experience in a DevOps, Platform Engineering, or Site Reliability Engineering role.
3+ years of direct experience designing and managing cloud infrastructure on AWS (primary) and/or Microsoft Azure.
CI/CD & Automation
Proven expertise in building and managing CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI, or Azure DevOps.
Strong proficiency in scripting and automation using Python, Bash, and/or Go.
Deep experience with Infrastructure as Code using Terraform (required) and configuration management with Ansible.
Solid understanding of GitOps workflows and tools such as ArgoCD or Flux.
Software Development Experience
Solid hands-on experience developing applications in Python (required) — including libraries, frameworks (Django, Flask, FastAPI), and package management.
Working knowledge of Java or Node.js (JavaScript/TypeScript) — ability to understand application architecture, dependencies, and deployment requirements.
Familiarity with REST APIs, microservices design patterns, and application lifecycle management.
Ability to read, review, and contribute to application code to understand build/deployment requirements and collaborate effectively with development teams.
Experience with version control workflows (Git), code review processes, and pull request management.
Cloud & Infrastructure
Hands-on expertise with core AWS services: EC2, EKS, ECS, Lambda, RDS, Aurora, S3, VPC, IAM, CloudFormation, Secrets Manager, CloudWatch, and Route53.
Experience managing Kubernetes clusters in production, including networking (CNI), storage (CSI), autoscaling (HPA/VPA/KEDA), and RBAC.
Proficiency with Docker and container image lifecycle management, including multi-stage builds and image security hardening.
Experience with networking protocols and concepts: TCP/IP, DNS, TLS/SSL, HTTP/S, load balancing, VPN, and VPC peering/Transit Gateway.
Familiarity with the AWS Well-Architected Framework and its five pillars: Operational Excellence, Security, Reliability, Performance Efficiency, and Cost Optimization.
Security & Compliance
Demonstrated experience implementing DevSecOps practices including SAST/DAST integration, secrets management, and container vulnerability scanning.
Working knowledge of IAM policy design, RBAC, and cloud security posture management (CSPM) tools.
Familiarity with compliance frameworks relevant to financial services: SOC 2, PCI-DSS, FFIEC, or equivalent.
Observability & Reliability
Experience building observability platforms using tools such as Datadog, Prometheus, Grafana, ELK Stack, or AWS CloudWatch/X-Ray.
Proven ability to define SLOs, build alerting frameworks, and lead incident response and post-mortem processes.
Experience with chaos engineering principles and tools (e.g., AWS Fault Injection Simulator, Gremlin) is a plus.
Collaboration & Communication
Strong written and verbal communication skills, with the ability to convey complex technical concepts to both technical and non-technical stakeholders.
Demonstrated experience working in Agile/Scrum delivery models with cross-functional engineering teams.
Ability to produce and maintain high-quality technical documentation including runbooks, architecture diagrams, and post-incident reports.
Preferred Qualifications
AWS Certified DevOps Engineer – Professional or AWS Solutions Architect – Associate/Professional certification.
Microsoft Certified: Azure DevOps Engineer Expert or Azure Administrator Associate certification.
Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD).
HashiCorp Certified: Terraform Associate.
Experience with service mesh technologies (Istio, Linkerd, AWS App Mesh).
Exposure to FinOps practices and cloud cost governance frameworks.
Experience in regulated industries (financial services, healthcare) with demonstrated ability to meet audit and compliance requirements.
Contributions to open-source projects or internal developer tooling/platform initiatives.
Familiarity with data engineering concepts: ETL pipelines, data lakes, streaming platforms (Kafka, Kinesis), and big data tooling.
Experience with API management platforms and gateway technologies (Kong, AWS API Gateway, Apigee).
Skills
Similar jobs
DevOps Engineer
ClearBridge Technology Group · San Francisco, United States
1 hour ago$140k - $180k/yrSite Reliability Engineer III
Apex Systems · Plano, United States
2 hours agoAI Platform Engineer
MassMutual · New York, United States
5 hours agoAI Platform Engineer
MassMutual · Boston, United States
5 hours agoAI Platform Engineer
MassMutual · Springfield, United States
5 hours agoSRE Engineer
Balin Technologies LLC · Sunnyvale, United States
7 hours ago