← Back to Jobs
Technology
Site Reliability Engineer (SRE) ( Need Visa Candidates TN,E3 Only)
Incorporan IncSchaumburg, IL🇺🇸United StatesPosted 12 Aug 2026
Quick Overview
Work Type
Hybrid
Level
Mid Senior
Job Description
Job Role - Site Reliability Engineer (SRE)
Location - Schaumburg, IL or Secaucus, NJ (Hybrid)
Location - Schaumburg, IL or Secaucus, NJ (Hybrid)
Mandatory Skills: Python/R and ML libraries (scikit-learn, TensorFlow, PyTorch), Data analysis and visualization (Pandas, NumPy, Power BI/Tableau), SQL and database management
Note - We are seeking a highly motivated Site Reliability Engineer (SRE) to ensure the reliability, scalability, performance, and availability of critical production systems. The ideal candidate will combine software engineering and operations expertise to build automation, improve system resilience, reduce operational toil, and enhance service reliability
Key Responsibilities
• Monitor, maintain, and improve the reliability, availability, and performance of production systems.
• Design and implement monitoring, alerting, logging, and observability solutions.
• Establish and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
• Automate operational tasks and repetitive processes using scripting and Infrastructure as Code (IaC).
• Lead incident response activities, troubleshooting, root cause analysis (RCA), and post-incident reviews.
• Collaborate with development, infrastructure, and platform teams to improve system reliability and resilience.
• Perform capacity planning, performance tuning, and scalability assessments.
• Support CI/CD pipelines and deployment automation initiatives.
• Implement high-availability, disaster recovery, and failover strategies.
Required Skills
• Strong experience with Linux/Unix administration.
• Proficiency in scripting languages such as Python, Shell, or PowerShell.
• Hands-on experience with cloud platforms (AWS, Azure, or Google Cloud Platform).
• Experience with containerization technologies such as Docker and Kubernetes.
• Knowledge of monitoring and observability tools such as Prometheus, Grafana, ELK, Splunk, Dynatrace, or Datadog.
• Understanding of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
• Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation.
• Strong troubleshooting, debugging, and problem-solving skills.
• Understanding of networking, security, and distributed systems concepts.
Key Responsibilities
• Monitor, maintain, and improve the reliability, availability, and performance of production systems.
• Design and implement monitoring, alerting, logging, and observability solutions.
• Establish and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
• Automate operational tasks and repetitive processes using scripting and Infrastructure as Code (IaC).
• Lead incident response activities, troubleshooting, root cause analysis (RCA), and post-incident reviews.
• Collaborate with development, infrastructure, and platform teams to improve system reliability and resilience.
• Perform capacity planning, performance tuning, and scalability assessments.
• Support CI/CD pipelines and deployment automation initiatives.
• Implement high-availability, disaster recovery, and failover strategies.
Required Skills
• Strong experience with Linux/Unix administration.
• Proficiency in scripting languages such as Python, Shell, or PowerShell.
• Hands-on experience with cloud platforms (AWS, Azure, or Google Cloud Platform).
• Experience with containerization technologies such as Docker and Kubernetes.
• Knowledge of monitoring and observability tools such as Prometheus, Grafana, ELK, Splunk, Dynatrace, or Datadog.
• Understanding of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
• Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation.
• Strong troubleshooting, debugging, and problem-solving skills.
• Understanding of networking, security, and distributed systems concepts.
Experience
• 5–10+ years of overall IT experience.
• 5+ years of hands-on experience in Site Reliability Engineering, Production Support, DevOps, or Cloud Operations roles.
• 5–10+ years of overall IT experience.
• 5+ years of hands-on experience in Site Reliability Engineering, Production Support, DevOps, or Cloud Operations roles.
Skills
Docker
SQL
Shell
AWS
ELK
NumPy
Scikit-learn
Splunk
Tableau
Ansible
Azure
CloudFormation
Datadog
GitHub Actions
GitLab CI
Google Cloud
Grafana
Jenkins
Kubernetes
Pandas
Power BI
PowerShell
Prometheus
PyTorch
Python
TensorFlow
Terraform
Similar jobs
Content Delivery Network (CDN) Platform Engineer
Apex Systems · Dearborn, United States
5 minutes agoPrincipal Network Engineer, onsite, Tucson, AZ with Security Clearance
RTX · Tucson, United States
5 minutes ago$107.5k - $204.5k/yrJr. SRE Engineer
Balin Technologies LLC · Sunnyvale, United States
6 minutes agoDevops Engineer
Nexylum Global LLC · Dallas, United States
8 minutes agoDevOps Engineer Staff
Lockheed Martin Corporation · Fort Meade, United States
22 minutes ago$128.2k - $226.0k/yrAzure Cloud Platform Engineer, Onsite - 69673
PRIMUS Global Services Inc. · United States
1 hour ago