SRE - AI Infrastructure & GPU HPC Ops
Why This Role Stands Out
This on-site SRE role offers an exciting opportunity to build expertise in cutting-edge AI infrastructure and GPU HPC operations within a reputable firm, where you'll collaborate closely with diverse engineering teams. You'll thrive here if you enjoy hands-on hardware and Linux troubleshooting and are eager to contribute to the reliability and scalability of advanced systems. Apply now to join a dynamic, fast-paced Melbourne environment and grow your career in a critical tech domain.
Quick Overview
Job Description
Firmus Technologies seeks a Site Reliability Engineer for AI Infrastructure to support and maintain our AI HPC infrastructure. You will work with Field Service Engineers, HPC and Network teams, and assist the Global Operations Centre to ensure reliability and scalability.
The role emphasizes hardware and Linux troubleshooting, scripting, and collaboration across engineering teams in a fast-paced, on-site Melbourne environment.
Similar jobs
Site Reliability Engineer III - Chief Technology Office
J.P. Morgan · Glasgow, United Kingdom
12 minutes agoHead of Application Operations & Site Reliability Engineering
HSBC · London, United Kingdom
12 minutes agoDevOps Engineer
Anson Mccade · Newcastle Upon Tyne, United Kingdom
26 minutes ago£45k/yrAutomation Tester - Agile & DevOps Focus, Canberra
IT Alliance Australia · Canberra, Australia
1 hour agoNetwork Engineer
Xcellink Pte Ltd · Watsons Creek, Australia
1 hour agoSenior DevOps Engineer
IT Alliance Australia · Canberra, Australia
1 hour ago