Platform Engineer - GPU Cloud Infrastructure
Quick Overview
Job Description
Location: Parramatta, NSW - office-basedEmployment type: Full-time permanentRemuneration: $95,000-$105,000 total package inclusive of superannuationAdditional benefits: ESOP participation may be consideredStart date: Within four weeks
About OmegaAIOmegaAI is an Australian technology company building infrastructure for AI and compute-intensive workloads.
We are seeking a Platform Engineer to help evaluate, integrate and support specialist GPU cloudenvironments. You will work alongside experienced engineers, contribute to practical proof-of-conceptprojects and help develop reliable automation and operational processes.
Responsibilities- Investigate and evaluate specialist GPU cloud services.
- Configure and maintain development and testing environments.
- Support proof-of-concept integrations and technical evaluations.
- Deploy and run containerised workloads.
- Collect and analyse logs, operational results and performance data.
- Develop infrastructure automation using tools such as Terraform, Ansible, Python or shell scripting.
- Monitor resource availability, health, utilisation and cost.
- Investigate failed workloads and escalate complex issues when necessary.
- Contribute to automated testing, code reviews and deployment processes.
- Produce clear technical documentation, runbooks and investigation reports.
- Follow secure practices for credentials, secrets and infrastructure access.
- Ensure temporary resources are tracked and removed when no longer required.
- Collaborate with engineering and product teams to deliver dependable platform capabilities.
- Practical experience in platform engineering, DevOps, systems administration, cloud infrastructure or software engineering.
- Strong working knowledge of Linux.
- Experience with Python, Bash or another scripting language.
- Familiarity with Git-based development workflows.
- Experience using Docker or other container technologies.
- Understanding of at least one public cloud or infrastructure platform.
- Exposure to infrastructure-as-code or configuration-management tools such as Terraform or Ansible.
- Fundamental understanding of networking, DNS, TLS and storage.
- Ability to work with APIs and troubleshoot integration problems.
- Strong attention to detail and an ability to operate controlled infrastructure safely.
- Clear written and verbal communication.
- An interest in GPU infrastructure, AI platforms or high-performance computing.
Experience in any of the following would be advantageous but is not essential:
- Specialist GPU cloud or compute platforms.
- NVIDIA GPUs, CUDA or GPU-enabled containers.
- Kubernetes, Slurm or other workload-scheduling systems.
- Prometheus, Grafana or similar monitoring tools.
- CI/CD pipelines and automated testing.
- Object storage and large-scale data movement.
- REST API integrations.
- Secrets management and infrastructure security.
- Distributed computing, AI infrastructure or HPC projects.
Applicants must:
- Be an Australian citizen with current Australian working rights.
- Be eligible to obtain and maintain an Australian Government NV1 security clearance.
- Have a checkable personal and employment background.
- Be able to work from our Parramatta office.
- Be available to commence within four weeks.
An existing security clearance is desirable but not mandatory.
Our Working CultureThis is not a role for someone seeking a comfortable position, minimal accountability or a place to simply collect a salary.
We are building a high-performance team where initiative, ownership, intellectual curiosity, reliability and a willingness to go beyond the minimum are essential. You will be expected to solve problems, take responsibility, continuously improve and contribute meaningfully without needing to be pushed.
If you are looking for a routine role where you can remain passive or simply follow instructions, thisopportunity is unlikely to be the right fit.
What Success Looks LikeWithin your first three months, you should be able to:
- Understand and follow our engineering and operational processes.
- Configure a controlled test environment using approved automation.
- Deploy a containerised workload and collect meaningful operational results.
- Monitor infrastructure health, utilisation and cost.
- Diagnose common workload and infrastructure failures.
- Contribute tested and reviewed automation changes.
- Produce clear documentation and operational runbooks.
- Demonstrate ownership by tracking tasks through to completion.
- Operate responsibly, including cleaning up temporary infrastructure and protecting credentials.
- Practical experience with emerging GPU and AI infrastructure.
- Exposure to cloud platforms, automation and high-performance computing.
- Mentoring and support from experienced engineers.
- A clear pathway for professional and technical growth.
- Direct involvement in meaningful infrastructure projects.
- An office-based, collaborative engineering environment.
- ESOP participation may be considered for the right candidate.
- Structured performance reviews and development opportunities.
Skills
Similar jobs
Software Engineer 1 (DevOps) with Security Clearance
Avid Technology Professionals · Linthicum Heights, United States
23 minutes agoDevOps SME
Leidos · Springfield, United States
39 minutes ago$154.1k - $278.5k/yrDevOps SME
Leidos · Chantilly, United States
39 minutes ago$154.1k - $278.5k/yrDevOps SME
Leidos · Reston, United States
39 minutes ago$154.1k - $278.5k/yrDevOps SME
Leidos · Sterling, United States
39 minutes ago$154.1k - $278.5k/yrDevOps SME
Leidos · Falls Church, United States
39 minutes ago$154.1k - $278.5k/yr