Quick Overview
Job Description
Job Title: Senior Platform / Site Reliability Engineer (SRE)
Location: Remote - Australia
Employment Type: Full-Time.
ABOUT THE OPPORTUNITYWe are looking for an experienced Senior Platform / Site Reliability Engineer to take ownership of a growing cloud-based SaaS environment.
This is a hands on technical role for someone who enjoys solving infrastructure challenges, improving system reliability, and working independently. The selected candidate will be responsible for maintaining a secure, stable, and scalable AWS environment while supporting ongoing product development and business growth.
KEY RESPONSIBILITIES- Manage AWS production environments across Australia and the United States using Terraform.
- Improve monitoring, logging, alerting, and observability to identify and resolve issues early.
- Maintain CI/CD pipelines and improve deployment processes, including automated health checks and rollbacks.
- Ensure PostgreSQL and Amazon RDS databases remain reliable, secure, and optimized for performance.
- Manage database backups, restoration procedures, query optimization, and scaling.
- Develop, document, and regularly test disaster recovery procedures.
- Monitor infrastructure capacity and prepare systems for increasing customer demand.
- Collaborate with software development teams to improve application reliability and performance.
- Manage infrastructure maintenance, server updates, security patches, and vulnerability fixes.
- Investigate production incidents, identify root causes, and implement long term solutions.
- Improve infrastructure automation and operational processes without disrupting ongoing development.
- Minimum 5 years of experience in Platform Engineering, Site Reliability Engineering, DevOps, or Cloud Infrastructure.
- At least 3 years of hands on experience managing AWS production environments for SaaS applications.
- Strong experience with Terraform and Infrastructure as Code (IaC).
- Hands on knowledge of CI/CD pipelines, automated deployments, and rollback procedures.
- Experience with monitoring, alerting, logging, and observability tools.
- Strong PostgreSQL and Amazon RDS experience, including performance tuning, backups, restores, and scalability.
- Experience developing and testing disaster recovery plans, including RTO and RPO.
- Good understanding of AWS security, networking, access management, and infrastructure management.
- Experience with capacity planning, performance optimization, and production incident resolution.
- Ability to independently manage production infrastructure and take full technical ownership.
- Strong problem solving and communication skills.
- Familiarity with SOC 2, ISO 27001, or similar compliance standards.
- Experience working with managed-service providers or external support teams.
- Background supporting enterprise SaaS applications, particularly in industrial or supply chain environments.
- Previous experience as the primary or sole Platform/SRE Engineer.
- Experience managing cloud infrastructure across multiple regions.
We are looking for a proactive, hands on engineer who can independently manage and improve a production SaaS environment.
The ideal candidate will have strong AWS, Terraform, PostgreSQL/RDS, and CI/CD experience, along with proven skills in monitoring, disaster recovery, infrastructure security, and platform scalability.
This position is suitable for someone who takes ownership, works independently, and focuses on practical improvements that strengthen reliability and support business growth.
Similar jobs
- AC
Platform Engineer: Cloud, Kubernetes & IaC Lead
NewAkuna Capital
Sydney🇦🇺Hybrid1 hour agoKubernetesTerraformTechnology - PL
Senior Site Reliability Engineer - Linux & Cloud
NewPlayStation
Adelaide🇦🇺Hybrid1 hour agoTechnology - CP
Senior Business Analyst - Agile & DevOps, Onsite
NewCompas Pty Ltd
Canberra, Australian Capital Territory🇦🇺On-site1 hour agoTechnology - PL
Staff Platform Engineer, Network Automation
NewPlayStation
Adelaide🇦🇺Hybrid1 hour agoTechnology - SO
Senior Data Platform Engineer AI-ready, Multi-tenant OLAP
NewSoCode
Western Australia🇦🇺Hybrid1 hour agoTechnology - TI
Security SRE Engineer - Build Reliable Security Platforms
NewTikTok
Sydney🇦🇺Hybrid1 hour agoTechnology - TI
Site Reliability Engineer, Security Engineering
NewTikTok
Sydney🇦🇺Hybrid2 hours agoDjangoExpressSpring+8Technology - FG
Junior C# Developer - AWS, DevOps & AI Automation
NewFTI Group
Sydney, New South Wales🇦🇺Hybrid2 hours agoSQLSQL ServerAWS+7Technology - GO
SRE & Software Engineering Manager - Scale & Reliability
NewGoogle LLC
Sydney🇦🇺Hybrid2 hours agoTechnology - CA
Site Reliability Engineer
NewCareers
Sydney🇦🇺Hybrid2 hours agoDockerGCPAWS+7Technology - IT
Azure DevOps Engineer - Greenfield Build (4-Week Contract)
NewIterate
Melbourne, Victoria🇦🇺Hybrid2 hours agoAzureTechnology - ZO
Senior Platform SRE - Remote AU, Scale & Reliability
NewZohorecruit
Brisbane, Queensland🇦🇺Remote2 hours agoAWSTechnology