Quick Overview
Job Description
Firmus Technologies
Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.
Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.
At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.
Firmus AI Cloud
Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.
It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.
AI FactoryOS Operations
AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system.
AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against.
The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability, and builds and operates the shared services the estate's own operation depends on. The function works closely with the engineering teams that build the platform, supplying the production evidence that shapes what they fix and what they build next.
Role Summary
Firmus runs large-scale, state-of-the-art AI infrastructure built on the latest generation of GPU rack-scale systems and operated as one estate to power the next generation of AI innovation. The Service Delivery Manager owns the incident and change management practices this operation runs on: the severity model and major incident command, the change calendar and change records, the runbook programme, production readiness review, and the operational reporting that keeps leadership and customers informed of service health, so that our services remain reliable and secure.
The role sits at the fusion of ITIL and SRE practices. Incident, problem and change are managed with the rigour of formal service management, and delivered with the methods of reliability engineering: measured against service level objectives, automated wherever automation makes response faster and safer, and continuously improved from what incidents reveal.
The role owns the processes, not the technical decisions inside them. This role owns the standard, the record and the discipline that make those decisions consistent, visible and auditable, and it owns the authorisation path that turns a technically complete release into an authorised production change.
Key Responsibilities
- Own and run the incident management practice for the function: the severity model, major incident command, and the standards every team operates to during an incident.
- Hold the incident record during major incidents, including the communications cadence, stakeholder updates and the customer commitment, and coordinate blameless post-incident review through to closed actions.
- Own the change calendar and change enablement process across the estate, including approvals, scheduling, conflict management and freeze periods, and maintain the change record as the definitive account of what changed.
- Act as product owner for the runbook programme: prioritise which faults are converted into guarded, tested automation, hold authors to a standard the first-response team can execute unaided, and track whether the programme is reducing escalation volume.
- Own the problem record: recurring faults, their root-cause status, and the engineering work required to close them out permanently.
- Own operational acceptance as the gate every new or materially changed service passes through before it is declared supported. Coordinate the domain specialists and testers each acceptance needs, hold the technical sign-off from the accountable engineer and the operational acceptance from the service owner, and refer residual risk to the Head of AI FactoryOS Operations for the declared production risk position.
- Report service level and error budget performance across the portfolio, and make error budget breaches visible as a reliability obligation on the engineering backlog rather than a number in a monthly pack.
- Produce operational reporting for leadership covering incident trends, change success rate, service health and the state of the runbook programme, and own customer communication on service health, incidents and planned change.
- Own the collation of access, change and incident evidence for ISO 27001, SOC 2 and enterprise customer due diligence, drawing on every team as an evidence source.
- Provide continuous cross-region coverage of the incident and change practice alongside peer Service Delivery Managers, and coach engineers on following the practices consistently.
Skills & Experience
- Significant experience in service delivery, service management or IT operations management in a large cloud provider, hyperscaler or infrastructure service provider context, including ownership of an incident, change and problem management practice in a 24/7 environment.
- Proven experience running major incident management: severity classification, incident command, stakeholder communication and post-incident review.
- Experience owning a change management or change enablement process for a technical environment, including a change calendar, approvals and change records.
- Experience with service transition and acceptance into production, ensuring new or changed services are supportable before go-live.
- Comfortable working closely with technical teams and technical detail, with the credibility to hold engineering teams to an agreed operational standard.
- Experience producing operational reporting for leadership, covering incident trends, service level performance and service health.
- Experience operating under formal compliance frameworks such as ISO 27001 or SOC 2, including contributing to audit and compliance evidence production.
- Strong stakeholder management and communication skills, including customer-facing communication on service health, incidents and planned change, with the ability to hold a consistent standard across multiple technical teams.
- Experience with an ITSM or incident management platform (for example ServiceNow, Jira Service Management or PagerDuty).
Preferred Experience
- Experience in a data centre, cloud, HPC or AI infrastructure environment.
- Familiarity with GitOps or infrastructure-as-code change workflows, sufficient to review a change record without authoring the change.
- ITIL certification or an equivalent formal service management qualification.
- A Bachelor's degree in computer science, engineering or a related discipline, or an equivalent combination of relevant experience and training.
Location & Reporting
Location: Based in Australia or Singapore, with travel to Australian AI Factory sites as required.
On-call: The function runs 24/7. This role provides major incident command cover across regions alongside peer Service Delivery Managers, on a published roster.
Reporting to: Reports to the Head of AI FactoryOS Operations while the function is being established, working under broad direction with a high degree of autonomy and direct access to the decision makers. As the function reaches its planned structure, the role will report to the Service Reliability Manager, with the Head of AI FactoryOS Operations remaining accountable for the function. The scope, level and remit of the role do not change under either arrangement.
Employment Basis
Permanent full-time
Diversity
At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.
Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.