Why This Role Stands Out
This hybrid DevOps Engineer role offers a fantastic opportunity to develop your troubleshooting and automation skills in a high-impact production environment, directly contributing to service availability and operational stability. If you're a proactive problem-solver eager to take ownership and excel in a dynamic cloud infrastructure setting, this role is for you. Apply today to join a team dedicated to robust cloud operations.
Quick Overview
Job Description
Cloud Infrastructure Operations Engineer
Description:
- You will work in a highly critical production environment as a frontline Cloud Infrastructure Engineer, you will be the first point of contact for cloud-related escalations reported by customers and application teams. This is a high-impact, “eye-on-glass” role where your primary mission is to identify, isolate, and resolve infrastructure issues before they impact business operations. You are expected to be an “active troubleshooter” and not just a ticket router who uses deep technical knowledge and modern AI tools to expediate resolutions and support our senior team.
- Ability to safely execute automation scripts for cloud resource and OS-level operations, understanding scripts inputs and outputs, workflow, and troubleshooting basic execution issues.
- We value engineers who take full responsibility for their work. Take prompt action on production issues, follow established troubleshooting and escalation procedures, and maintain a strong focus on service availability, uptime, and operational stability.
RESPONSIBILITIES:
- Act as the primary responder for incoming cloud infrastructure issues. Perform rapid triage to understand the scope and business impact of reported problems.
- Collect and analyze system logs, network traces, and filesystem information. Apply Red Hat Linux skills to perform basic LVM operations, analyze syslog, and troubleshoot OS-level processes and services.
- Contact application owners or customers directly to gather missing details. Isolate issues between the application layer, OS layer and Cloud Infrastructure layer.
- When needed, open and manage support cases with cloud providers (AWS/Azure).
- If an issue exceeds defined SLAs, package all collected logs, evidence, and preliminary troubleshooting into a detailed and defined hand-off template for the senior Infrastructure team.
- Leverage AI assistants and AIOps tools to analyze and summarize logs, identify potential issues, and accelerate troubleshooting and evidence gathering.
- Monitor AWS and Azure Backup status report and perform VM/image-based restores such as AWS AMI and Azure image restores following established procedures.
- Develop and maintain small-scale automation using Bash, Python, and Ansible to streamline data collection, filtering, and routine troubleshooting and evidence-gathering tasks.
- Available to provide support when needed during off-hours, nights, or weekends when critical infrastructure issues or escalation require immediate attention.
SKILLS
Cloud Platforms: Hands-on experience with AWS and Azure core services:
- Compute: EC2, Virtual Machine, AMI and Managed image
- Networking: VPC/Vnet, Subnets, Security Groups, NSGs, VPN, Direct Connect, Express Route, and Load Balancers
- Storage: S3, EBS, EFS, Azure Blobs, Managed disks and Vault
- Security: IAM policies, RBAC, Secret keys
- Backup and Restore: Backup configuration, monitoring, and restore operations for cloud resources and virtual machines
Linux & OS-Level: Strong administration skills in Red Hat Linux are mandatory:
- LVM, filesystem administration – creation, expansion
- Managing system services, daemons, os-level logging
- Performance tuning and resource troubleshooting (CPU/Memory/IO etc.)
Scripting/Automation: Proficiency in programming/scripting languages and automation tools such as Python, Bash, or Ansible
AI for support: Familiarity with AI tools for troubleshooting, log analysis, and infrastructure support
Similar jobs
- DU
Senior Site Reliability Engineer
NewDuolingo
New York🇺🇸Hybrid2 hours agoDockerMySQLJava+5Technology - DI
Principal, Staff Site Reliability
NewDigitalBridge
Boca Raton🇺🇸Hybrid2 hours agoAWSELKSOC 2+14Technology - SA
Cloud Platform Engineer
NewStaffRight Associates, LLC
United States🇺🇸$70 - $75/hrRemote18 hours agoAngularAzureC#+6Technology - LT
URGENT NEED Search Platform Engineer
NewLorven Technologies, Inc.
Sunnyvale, CA🇺🇸On-site18 hours agoDockerAWSDatadog+7Technology - TE
Platform Engineer (Kubernetes & Observability)
NewTEKEngineersInc
Irvine, CA🇺🇸On-site18 hours agoFastAPINode.jsHelm+2Technology - TE
DevOps Engineer - Dallas, TX, Austin, TX, Houston, TX, San Antonio, TX.
NewTechniPros, LLC
Dallas, TX🇺🇸Hybrid18 hours agoDockerAWSAnsible+6Technology