Why This Role Stands Out
This hybrid role offers a fantastic opportunity to leverage your AWS, AIOps, and full-stack expertise to build and operate highly available cloud applications, driving innovation through AI-driven automation. You'll thrive here if you're eager to enhance system reliability and performance across the entire tech stack, contributing to impactful projects within a forward-thinking company.
Quick Overview
Job Description
We are seeking a highly skilled AWS Site Reliability Engineer (SRE) with strong AIOps and Full Stack development experience to design, build, automate, and operate highly available cloud applications and platforms.
The ideal candidate will have a strong background in AWS infrastructure, Site Reliability Engineering, DevOps, observability, AIOps, automation, and full-stack application development. This role requires someone who can work across the entire technology stack—from cloud infrastructure and CI/CD pipelines to backend services, APIs, frontend applications, monitoring, and AI-driven operational automation.
The engineer will be responsible for improving system reliability, scalability, performance, and operational efficiency while leveraging AI/ML and AIOps capabilities to proactively identify incidents, predict failures, automate remediation, and improve application observability.
Key Responsibilities
AWS & Site Reliability Engineering
- Design, deploy, and maintain highly available, scalable, and fault-tolerant applications and infrastructure on AWS.
- Implement SRE principles including SLIs, SLOs, SLAs, error budgets, availability, reliability, and capacity planning.
- Manage AWS services including EC2, EKS, ECS, Lambda, S3, RDS, DynamoDB, API Gateway, CloudFront, Route 53, IAM, VPC, CloudWatch, SNS, SQS, and EventBridge.
- Troubleshoot complex production issues involving application, infrastructure, networking, database, and cloud components.
- Participate in incident management, root-cause analysis, problem management, and post-incident reviews.
- Develop automation to reduce manual operational activities and eliminate repetitive tasks.
- Perform capacity planning, performance tuning, disaster recovery, and business continuity activities.
DevOps & Infrastructure Automation
- Build and maintain CI/CD pipelines using tools such as Jenkins, GitHub Actions, GitLab CI/CD, AWS CodePipeline, or Azure DevOps.
- Implement Infrastructure as Code using Terraform, CloudFormation, or AWS CDK.
- Automate infrastructure provisioning, configuration management, deployments, and operational processes.
- Implement containerized workloads using Docker and Kubernetes/Amazon EKS.
- Develop automated deployment strategies including blue-green, canary, and rolling deployments.
- Integrate security, compliance, testing, and quality checks into CI/CD pipelines.
AIOps & AI-Driven Operations
- Implement AIOps solutions to improve monitoring, incident detection, event correlation, root-cause analysis, and automated remediation.
- Leverage AI/ML and Generative AI capabilities to analyze logs, metrics, traces, alerts, and operational data.
- Develop intelligent alerting and anomaly-detection mechanisms to identify potential production issues before they impact customers.
- Build AI-assisted incident investigation and troubleshooting workflows.
- Integrate LLM/GenAI capabilities into SRE and DevOps workflows for automated log analysis, incident summarization, knowledge retrieval, and remediation recommendations.
- Develop automated runbooks and self-healing mechanisms using event-driven AWS services and AI-assisted decision making.
- Integrate AIOps platforms and observability tools such as Datadog, Dynatrace, New Relic, Splunk, CloudWatch, Grafana, and Prometheus.
- Develop or integrate AI agents/workflows that can assist with incident response, operational diagnostics, and infrastructure management.
- Monitor AIOps/AI solutions for accuracy, reliability, security, and operational effectiveness.
Observability & Monitoring
- Implement comprehensive metrics, logs, traces, dashboards, and alerting across cloud and application environments.
- Work with Prometheus, Grafana, CloudWatch, OpenTelemetry, Datadog, Dynatrace, Splunk, or similar observability platforms.
- Establish meaningful service-level indicators and operational dashboards.
- Implement distributed tracing and application performance monitoring.
- Tune alerts to reduce false positives and alert fatigue.
- Build proactive monitoring and predictive health checks.
Full Stack Development
- Develop and maintain scalable backend services, APIs, and web applications.
- Build RESTful APIs and microservices using technologies such as Java/Spring Boot, Python/FastAPI, Node.js, or similar.
- Develop responsive frontend applications using React, Angular, TypeScript, JavaScript, HTML, and CSS.
- Integrate frontend applications with REST/GraphQL APIs and cloud-native backend services.
- Design and optimize database interactions using PostgreSQL, MySQL, MongoDB, DynamoDB, or similar databases.
- Implement authentication and authorization using OAuth 2.0, OpenID Connect, JWT, AWS IAM, or similar technologies.
- Develop automated unit, integration, API, and end-to-end tests.
- Troubleshoot application performance and scalability issues across frontend, backend, database, and infrastructure layers.
Required Technical Skills
Cloud
- AWS
- EC2, S3, VPC, IAM, RDS, DynamoDB
- EKS/ECS, Lambda
- CloudWatch, Route 53, API Gateway
- SQS, SNS, EventBridge
- AWS networking and security
SRE / DevOps
- Site Reliability Engineering principles
- Incident Management & Root Cause Analysis
- CI/CD
- Terraform / CloudFormation / AWS CDK
- Docker
- Kubernetes / EKS
- Jenkins / GitHub Actions / GitLab CI
- Linux administration and troubleshooting
- Bash/Shell scripting
- Python
AIOps / AI
- AIOps and intelligent event management
- Generative AI / LLM concepts
- AI-assisted incident management
- Anomaly detection and predictive monitoring
- Automated remediation / self-healing systems
- LLM APIs and AI agent workflows
- RAG/vector database concepts are a plus
- Experience integrating AI with DevOps/SRE workflows
Observability
- Datadog / Dynatrace / New Relic
- Prometheus
- Grafana
- Splunk
- AWS CloudWatch
- OpenTelemetry
- Application Performance Monitoring
Full Stack
- React / Angular
- JavaScript / TypeScript
- HTML / CSS
- Node.js / Python / Java
- REST APIs / GraphQL
- Microservices
- SQL / NoSQL databases
Preferred Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field.
- 7+ years of experience in software engineering, DevOps, Cloud Engineering, or SRE.
- 4+ years of hands-on AWS experience.
- Strong experience supporting production environments and distributed systems.
- Experience implementing AIOps or AI-powered operational solutions.
- Experience with Kubernetes and cloud-native architectures.
- Experience developing full-stack applications.
- Strong programming and scripting skills in Python, Java, Node.js, or similar languages.
- Experience working in Agile/Scrum environments.
- Strong troubleshooting, analytical, communication, and problem-solving skills.
Key Competencies
- Cloud-native architecture
- Site Reliability Engineering
- AIOps & Generative AI
- Infrastructure automation
- Full-stack development
- Observability & monitoring
- Production support
- Incident response
- Automation & self-healing
- Performance optimization
- Security and reliability
- Continuous improvement
Similar jobs
- CO
Python Full Stack Gen AI Lead/Architect
NewCompunnel Inc.
Jersey City, NJ🇺🇸Hybrid17 hours agoAWSMLOpsMachine Learning+9 - SU
Board Developer
NewSumasEdge Corporation
Toronto, ON🇺🇸Hybrid17 hours agoTechnology - ST
Actimize Developer --- Mount Laurel, NJ - onsite
NewSavi Technologies
Mount Laurel Township, NJ🇺🇸On-site17 hours agoDockerSQLAgile+8Technology - JT
Full Stack Developer IV (Locals to MN only)
NewJaven Technologies, Inc
Richfield, MN🇺🇸Hybrid17 hours agoOracleSpringAgile+10Technology - SI
Full Stack Developer (W2 Only)
NewStellent IT LLC
United States🇺🇸Hybrid17 hours agoDockerExpressMicroservices+20Technology - TE
Salesforce Developer -W2 Position
NewTeknosys Inc
Raleigh, NC🇺🇸Hybrid17 hours agoAgileTechnology