Site Reliability Engineer
Quick Overview
Job Description
As an NVIDIA Cloud and Engineering partner in Asia Pacific, you will gain skills, experience, and exposure across the AI industry and be part of shaping what this industry looks like for decades to come.
We are founder-led, not a big corporate. Decisions happen fast, our leaders are accessible, and there's minimum bureaucracy between you and the work.
Ownership comes early. Whatever your role, you will have a direct line to outcomes, helping shape how the business grows as we scale nationally across a long-term, large-scale roadmap.
Work alongside founders and experts in AI infrastructure, energy systems and next-generation compute.
What we build here has impact beyond the business. Our AI Factories are designed to operate as assets to the energy grid to actively strengthen the communities and regions they operate in
rather than drawing from them.
Considering applying? You don't need a perfect background to join our team. If you're driven and curious, there's a path for you. We back our people to grow into new domains and take on challenges beyond their previous experience.
ROLE SUMMARYFirmus Technologies is seeking a skilled Site Reliability Engineer to join our Operations team, supporting the daily operations and maintenance of our AI-accelerated high-performance computing (HPC) infrastructure. This role will work closely with Field Service Engineers, HPC and Network Engineering teams, and assist the Global Operations Centre (GOC). This is a unique opportunity to contribute directly to the stability and growth of cutting-edge AI infrastructure.
KEY RESPONSIBILITIES- Support in the deployment, configuration, and maintenance of various high-end GPU servers, storage servers, networkingequipmentand software components in highly secure environments.
- Perform hardware diagnostics, systems functionality and firmware updates asrequired.
- Collaborate with engineering teams toassistin tailored customer environments deployment (eg: bare-metal systems, HPC Clusters, Kubernetes,Slurmetc).
- Serve as first line of engineering support for onsite operational issues, including troubleshooting hardware,networkand software problems, and firmware compliance.
- Troubleshoot incidents, escape criticalissuesand provide feedback toappropriate teamsfor improvements.
- Participate in an on-call rotation to ensure 24/7 availability and responsiveness to critical issues.
- Provide technical support to the GOC Support Specialist team in troubleshootingcompute infrastructurerelated problems.
- Document incident details, resolutions, and lessons learned to enhance future problem-solving.
- Maintain clear,accurate, and up-to-date documentation to promote effective knowledge sharing across the team.
- Communicate effectively with GOC, HPC Engineers, internal teams, stakeholders, and end-users to ensure alignment on issue resolution.
- Take part in team meetings and knowledge-sharing sessions to foster collaboration and continuous learning.
- Bachelor's degree in computer engineering, computer science, or a related technical field.
- 5+ years of experience in field service technical areas.
- Strong understanding of server hardware technology, firmware lifecycle, Linux environments and troubleshooting hardware problems, with adherence to physical and system-level security standards.
- Experience with scripting languages ( eg : Bash, Python)
- Familiarity with using configuration management, CICD tools, workload manager and cluster softwares ( eg : Slurm , Kubernetes, Nvidia BCM) and Observability tools ( eg : Prometheus, Grafana, ELK, etc)
- Excellent problem-solving and analytical skills.
- Ability to work independently and as part of a team.
- Strong communication skills, both written and verbal.
This role is based in Launceston, Tasmania, Australia.
Employment BasisFull-time
We are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.
Skills
Similar jobs
SRE Engineer
Balin Technologies LLC · Sunnyvale, United States
12 minutes agoAWS Cloud Platform Engineer
Lumen Solutions Group Inc. · Reston, United States
15 minutes agoDevOps Engineer with Security Clearance
Shadowgate Partners, Inc. · Fairfax, United States
26 minutes ago$125k - $215k/yrCybersecurity Platform Engineer with Security Clearance
Booz Allen Hamilton · Hanscom AFB, United States
26 minutes ago$77.6k - $176k/yrDevOps Cloud Engineer with Security Clearance
Power3 Solutions · Annapolis Junction, United States
47 minutes ago$140k - $235k/yrSenior DevOps Cloud Engineer with Security Clearance
Power3 Solutions · Annapolis Junction, United States
47 minutes ago$140k - $235k/yr