Haystack
← Back to Jobs
Full time
Technology

Senior DevOps Engineer, AI & Applications

Firmus Technologies Pty Ltd.Melbourne, Victoria🇦🇺AustraliaPosted 19 Jul 2026

Quick Overview

Work Type
Hybrid
Schedule
Full Time
Level
Mid Senior

Job Description

Senior DevOps Engineer, AI & Applications

Firmus Technologies

Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting edge technology with a steadfast commitment to sustainability.

We design, build, and operate a new class of digital infrastructure - the AI Factory. Our model-to-grid technology pushes the boundaries of multi generational liquid cooling systems, energy management, AI software orchestration, and construction to deliver low cost AI tokens globally.

Firmus AI Cloud

Our large scale GPU cloud platform is purpose built to deliver energy efficient AI compute at scale to customers.

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings.

Role Summary

Every AI feature we ship touches thousands of GPUs. The Senior DevOps Engineer will build the release engineering backbone - CI/CD pipelines, automated testing gates, one click deployments with instant rollback to let Firmus scale fast and responsibly.

You are the bridge between engineering and operations, setting Firmus standards for how code gets to production, mentoring the team on deployment safety, and driving a blameless culture when things go wrong.

Key Responsibilities
  • Design and maintain team wide CI/CD pipelines (Jenkins, GitHub Actions, ArgoCD, or equivalent) with automated testing gates, artifact management, and deployments aligned with GPU cluster standards.
  • Implement release engineering best practices: repeatable releases, GitOps workflows, automated rollback, and change management procedure.
  • Build and manage test infrastructure: environment provisioning, data seeding, long running job validation for distributed training templates and multi node job submissions.
  • Establish engineering protocols and standards: repo organization, PR templates, code quality gates, dependency scanning, static analysis.
  • Partner with infra teams to ensure AI product features deployment practices meet compliance and security standards for massive GPU clusters.
  • Mentor team on testing strategies, deployment safety, and incident response procedures.
Skills & Experience
  • 5-7 years of CI/CD engineering, release engineering, or DevOps experience.
  • Deep expertise in GitHub Actions, GitLab CI, ArgoCD, or Jenkins with multi stage pipeline design and testing gate implementation.
  • Strong automation scripting (Python, Go, or Bash) for build orchestration and environment templating.
  • Strong Kubernetes fundamentals with hands on experience understanding Pod lifecycle, deployments, jobs, services, and load behavior.
  • Config & secret management: experience designing and operating ConfigMaps and Secrets, including rotation patterns and least privilege hygiene.
  • Safe rollout patterns: experience implementing rolling updates, canary, blue/green, readiness/liveness probes, PodDisruptionBudgets, and rollback procedures to ensure zero/low downtime.
  • Deployment safety & debugging: ability to debug Kubernetes rollout issues and derive automated CI/CD gates.
  • Familiarity with artifact management, versioning, and rollback procedures.
  • Integrating testing frameworks into CI pipelines (unit, integration, end to end).
  • Track record of improving engineering velocity and time to release quarter over quarter while maintaining release standards.
  • Ensuring platform reliability and customer trust with rare incidents and fast recovery.
  • Improving developer productivity and team scale by reducing CI/CD friction.
  • Maintaining cost efficiency and resource optimization for CI/CD and test infrastructure.
  • Promoting releases and reliability practices as corporate default and reducing repeat incidents.
Success Metrics
  • Engineering velocity and time to release improve quarter over quarter while release standards remain consistent.
  • Platform reliability and customer trust remain strong with rare incidents and fast recovery.
  • Developer productivity and team scale improve as CI/CD friction decreases.
  • Cost efficiency and resource optimization improve with controlled or reduced costs per unit of output.
  • Release and reliability practices become default across the org and repeat incidents decrease.
Location & Reporting
  • Melbourne, Australia
  • Reporting to Head of AI & Applications
Employment Basis

Full time

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.

Skills

ArgoCD
Bash
GitHub Actions
GitLab CI
Jenkins
Kubernetes
Python
SAFe

Similar jobs