Skip to main content
Back to jobs

Senior AI Infrastructure Engineer, Kubernetes

Security Engineer, Application - Greenhouse

Sydney, Australia 9/13/2026 Full-time
Apply on source ↗ Save

&lt;p&gt;&lt;strong&gt;Firmus Technologies&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;&lt;span data-contrast=&quot;none&quot;&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot; data-ccp-parastyle-defn=&quot;{&amp;quot;ObjectId&amp;quot;:&amp;quot;baa2b379-a6d3-56f4-af2b-6a08de9434e5|1&amp;quot;,&amp;quot;ClassId&amp;quot;:1073872969,&amp;quot;Properties&amp;quot;:[469777841,&amp;quot;Aeonik&amp;quot;,469777842,&amp;quot;Aeonik&amp;quot;,469777843,&amp;quot;Aeonik&amp;quot;,469777844,&amp;quot;Aeonik&amp;quot;,469769226,&amp;quot;Aeonik&amp;quot;,201342446,&amp;quot;1&amp;quot;,201342447,&amp;quot;5&amp;quot;,201342448,&amp;quot;1&amp;quot;,201342449,&amp;quot;1&amp;quot;,201341986,&amp;quot;1&amp;quot;,268442635,&amp;quot;20&amp;quot;,335551500,&amp;quot;197122&amp;quot;,335559740,&amp;quot;264&amp;quot;,201341983,&amp;quot;0&amp;quot;,335559738,&amp;quot;145&amp;quot;,469775450,&amp;quot;FIR Body&amp;quot;,201340122,&amp;quot;2&amp;quot;,134234082,&amp;quot;true&amp;quot;,134233614,&amp;quot;true&amp;quot;,469778129,&amp;quot;FIRBody&amp;quot;,335572020,&amp;quot;1&amp;quot;,469778324,&amp;quot;Body Text&amp;quot;]}&quot;&gt;Firmus Technologies is a global&amp;nbsp;&lt;/span&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;leader&lt;/span&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt; pioneering the development &lt;/span&gt;&lt;/span&gt;&lt;span data-contrast=&quot;none&quot;&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;and operation of efficient AI infrastructure across Asia Pacific.&lt;/span&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;&amp;nbsp;&lt;/span&gt;&lt;/span&gt;&lt;span data-ccp-props=&quot;{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559738&amp;quot;:145,&amp;quot;335559740&amp;quot;:264}&quot;&gt;&amp;nbsp;&lt;/span&gt;&lt;/p&gt; &lt;p&gt;&lt;span data-contrast=&quot;none&quot;&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by &lt;/span&gt;&lt;/span&gt;&lt;span data-contrast=&quot;none&quot;&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;combining&amp;nbsp;&lt;/span&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;cutting-edge&lt;/span&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;&amp;nbsp;technology with a steadfast commitment to sustainability.&lt;/span&gt;&lt;/span&gt;&lt;span data-ccp-props=&quot;{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559738&amp;quot;:145,&amp;quot;335559740&amp;quot;:264}&quot;&gt;&amp;nbsp;&lt;/span&gt;&lt;/p&gt; &lt;p&gt;&lt;span data-contrast=&quot;none&quot;&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;At Firmus, we are unique in our approach. We design, build, and&amp;nbsp;&lt;/span&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;operate&lt;/span&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt; a new class of digital &lt;/span&gt;&lt;/span&gt;&lt;span data-contrast=&quot;none&quot;&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed &lt;/span&gt;&lt;/span&gt;&lt;span data-contrast=&quot;none&quot;&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;the boundaries of multi-generational liquid cooling systems, energy management, AI software &lt;/span&gt;&lt;/span&gt;&lt;span data-contrast=&quot;none&quot;&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;orchestration, and construction. For our customers, this approach allows us to make every watt &lt;/span&gt;&lt;/span&gt;&lt;span data-contrast=&quot;none&quot;&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;count and deliver low-cost AI tokens globally.&lt;/span&gt;&lt;/span&gt;&lt;span data-ccp-props=&quot;{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559738&amp;quot;:145,&amp;quot;335559740&amp;quot;:264}&quot;&gt;&amp;nbsp;&lt;/span&gt;&lt;/p&gt; &lt;p&gt;&amp;nbsp;&lt;/p&gt; &lt;p&gt;&lt;strong&gt;Firmus AI Cloud&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;&lt;span data-contrast=&quot;none&quot;&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot; data-ccp-parastyle-defn=&quot;{&amp;quot;ObjectId&amp;quot;:&amp;quot;baa2b379-a6d3-56f4-af2b-6a08de9434e5|1&amp;quot;,&amp;quot;ClassId&amp;quot;:1073872969,&amp;quot;Properties&amp;quot;:[469777841,&amp;quot;Aeonik&amp;quot;,469777842,&amp;quot;Aeonik&amp;quot;,469777843,&amp;quot;Aeonik&amp;quot;,469777844,&amp;quot;Aeonik&amp;quot;,469769226,&amp;quot;Aeonik&amp;quot;,201342446,&amp;quot;1&amp;quot;,201342447,&amp;quot;5&amp;quot;,201342448,&amp;quot;1&amp;quot;,201342449,&amp;quot;1&amp;quot;,201341986,&amp;quot;1&amp;quot;,268442635,&amp;quot;20&amp;quot;,335551500,&amp;quot;197122&amp;quot;,335559740,&amp;quot;264&amp;quot;,201341983,&amp;quot;0&amp;quot;,335559738,&amp;quot;145&amp;quot;,469775450,&amp;quot;FIR Body&amp;quot;,201340122,&amp;quot;2&amp;quot;,134234082,&amp;quot;true&amp;quot;,134233614,&amp;quot;true&amp;quot;,469778129,&amp;quot;FIRBody&amp;quot;,335572020,&amp;quot;1&amp;quot;,469778324,&amp;quot;Body Text&amp;quot;]}&quot;&gt;Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built &lt;/span&gt;&lt;/span&gt;&lt;span data-contrast=&quot;none&quot;&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;to deliver energy-efficient AI&amp;nbsp;&lt;/span&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;compute&lt;/span&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;&amp;nbsp;at scale to customers.&lt;/span&gt;&lt;/span&gt;&lt;span data-ccp-props=&quot;{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559738&amp;quot;:145,&amp;quot;335559740&amp;quot;:264}&quot;&gt;&amp;nbsp;&lt;/span&gt;&lt;/p&gt; &lt;p&gt;&lt;span data-contrast=&quot;none&quot;&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;It empowers developers, enterprises, educational institutions, and government users to train and &lt;/span&gt;&lt;/span&gt;&lt;span data-contrast=&quot;none&quot;&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services &lt;/span&gt;&lt;/span&gt;&lt;span data-contrast=&quot;none&quot;&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;and applications, we are committed to delivering a cloud experience that is market-leading, &lt;/span&gt;&lt;/span&gt;&lt;span data-contrast=&quot;none&quot;&gt;&lt;span data-ccp-parastyle=&quot;FIR Body&quot;&gt;proprietary, and built to scale.&lt;/span&gt;&lt;/span&gt;&lt;span data-ccp-props=&quot;{&amp;quot;201341983&amp;quot;:0,&amp;quot;335559738&amp;quot;:145,&amp;quot;335559740&amp;quot;:264}&quot;&gt;&amp;nbsp;&lt;/span&gt;&lt;/p&gt; &lt;p&gt;&lt;strong&gt;Role Summary&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;The Senior Kubernetes Engineer, AI Infrastructure owns the technical design and delivery of the backend infrastructure that powers the Firmus Kubernetes platform. This is a hands-on principal-level individual contributor role, responsible for building production-grade cluster lifecycle, control-plane, networking, storage, security, observability, and automation capabilities across GPU-accelerated bare-metal environments.&lt;/p&gt; &lt;p&gt;They solve the hardest platform engineering problems, set Kubernetes engineering standards, and provide domain-level technical sign-off for platform designs. They work across AI Platforms, Solutions Architecture &amp;amp; Delivery, networking, security, and operations to create a secure, resilient, multi-tenant platform that can be deployed and operated consistently at AI-factory scale.&lt;/p&gt; &lt;p&gt;&lt;strong&gt;Key Responsibilities&lt;/strong&gt;&lt;/p&gt; &lt;ul&gt; &lt;li&gt;Define and own the Kubernetes platform reference architecture across management and workload clusters, including control-plane topology, cluster lifecycle, multi-tenancy, workload isolation, and failure-domain design.&lt;/li&gt; &lt;li&gt;Build and maintain the backend services, APIs, controllers, operators, and automation required to provision, configure, upgrade, scale, and retire Kubernetes clusters reliably.&lt;/li&gt; &lt;li&gt;Engineer repeatable bare-metal Kubernetes deployment and lifecycle workflows using infrastructure-as-code and automated provisioning technologies such as Cluster API, kubeadm, Redfish, PXE, Ironic, or Metal3.&lt;/li&gt; &lt;li&gt;Design and operate cluster networking across CNI, ingress, service discovery, DNS, load balancing, network policy, and service mesh; integrate Multus, SR-IOV, BGP, InfiniBand, or RoCE where required for high-performance AI workloads.&lt;/li&gt; &lt;li&gt;Define persistent-storage and data-service patterns using CSI, Ceph, local NVMe, object storage, backup and restore, and disaster-recovery mechanisms appropriate for stateful platform and AI workloads.&lt;/li&gt; &lt;li&gt;Integrate and productionise NVIDIA GPU and Network Operators, device plugins, drivers, DCGM telemetry, scheduling, quotas, and topology-aware placement for multi-node accelerated workloads.&lt;/li&gt; &lt;li&gt;Establish GitOps and CI/CD patterns for platform software, configuration, policy, and release management, with safe testing, progressive rollout, rollback, and upgrade practices.&lt;/li&gt; &lt;li&gt;Build platform security into the architecture through identity and access control, RBAC, secrets management, policy-as-code, image and software-supply-chain controls, tenant isolation, and auditable change management.&lt;/li&gt; &lt;li&gt;Define service-level objectives and engineer observability for metrics, logs, traces, events, capacity, and performance; lead diagnosis of complex distributed systems failures and eliminate recurring operational toil.&lt;/li&gt; &lt;li&gt;Set engineering standards, design patterns, review practices, and operational readiness criteria; mentor senior engineers and resolve cross-team technical decisions while remaining directly involved in implementation.&lt;/li&gt; &lt;/ul&gt; &lt;p&gt;&amp;nbsp;&lt;/p&gt; &lt;p&gt;&lt;strong&gt;Skills &amp;amp; Experience&lt;/strong&gt;&lt;/p&gt; &lt;ul&gt; &lt;li&gt;7+ years of progressive infrastructure, systems, or platform engineering experience, including substantial ownership of production Kubernetes platforms and at least 3 years operating at senior staff, principal, or equivalent level.&lt;/li&gt; &lt;li&gt;Deep knowledge of Kubernetes internals, including the API server, etcd, scheduler, controller manager, kubelet, admission, CRI, CNI, CSI, reconciliation patterns, cluster performance, upgrades, and control-plane failure modes.&lt;/li&gt; &lt;li&gt;Demonstrated experience designing, building, and operating highly available, large scale and multi-cluster Kubernetes platforms on bare metal, private cloud, or hybrid infrastructure.&lt;/li&gt; &lt;li&gt;Strong software engineering ability in Go and/or Rust, with practical Python and Bash skills; experience building Kubernetes operators, controllers, admission webhooks, CLIs, or platform services.&lt;/li&gt; &lt;li&gt;Expert Linux systems knowledge, including namespaces, cgroups, systemd, kernel, host networking and container runtime behaviour, performance analysis, and low-level troubleshooting.&lt;/li&gt; &lt;li&gt;Strong Kubernetes networking expertise across Cilium, Calico, or equivalent CNI implementations, plus load balancing, DNS, ingress, BGP, network policy, and multi-network architectures.&lt;/li&gt; &lt;li&gt;Strong infrastructure automation and GitOps experience with tools such as Terraform, Ansible, Argo CD, Flux, GitHub Actions, GitLab CI, or Jenkins.&lt;/li&gt; &lt;li&gt;Practical experience with Kubernetes security and governance, including RBAC, OPA Gatekeeper or Kyverno, secrets management, certificate lifecycle, image security, and workload isolation.&lt;/li&gt; &lt;li&gt;Experience implementing production observability with Prometheus, Grafana, OpenTelemetry, Loki, Elasticsearch, or equivalent technologies, and using telemetry to manage reliability, capacity, and performance.&lt;/li&gt; &lt;li&gt;Experience with GPU-enabled Kubernetes infrastructure, NVIDIA GPU Operator, accelerator scheduling for AI workloads at large scale, RDMA networking, and distributed AI workload requirements.&lt;/li&gt; &lt;li&gt;Experience with distributed storage and data services such as Ceph, CSI-backed storage, object storage, backup and restore, and disaster recovery.&lt;/li&gt; &lt;li&gt;CKA-level expertise is expected; CKA, CKS, or relevant cloud-native certifications are strongly preferred.&lt;/li&gt; &lt;li&gt;Bachelor’s degree in computer science, engineering, or a related discipline, or equivalent depth of practical engineering experience.&lt;/li&gt; &lt;li&gt;Clear technical judgement and communication, with a record of influencing architecture across software, networking, security, platform, and operations teams.&lt;/li&gt; &lt;/ul&gt; &lt;p&gt;&amp;nbsp;&lt;/p&gt; &lt;p&gt;&lt;strong&gt;Location &amp;amp; Reporting&lt;/strong&gt;&lt;/p&gt; &lt;ul&gt; &lt;li&gt;Australia (Sydney, NSW or Launceston, TAS)&lt;/li&gt; &lt;li&gt;Reporting to Head of AI Platform&lt;/li&gt; &lt;/ul&gt; &lt;p&gt;&amp;nbsp;&lt;/p&gt; &lt;p&gt;&lt;strong&gt;Employment Basis&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;Full-time&lt;/p&gt; &lt;p&gt;&amp;nbsp;&lt;/p&gt; &lt;p&gt;&lt;strong&gt;Diversity&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.&lt;/p&gt; &lt;p&gt;Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.&lt;/p&gt;<p>Find more <a href="https://www.arbeitnow.co.uk/english-speaking-jobs">English Speaking Jobs in United Kingdom</a> on Arbeitnow</a>

AI Infrastructure Platform

Related on GOH