Infrastructure Automation Engineer, AI Cluster Commissioning
Description
Infrastructure Automation Engineer, AI Cluster Commissioning
Firmus Technologies
Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.
Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.
At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.
Firmus AI Cloud
Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.
It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.
Why Firmus?
As an NVIDIA Cloud and Engineering partner in Asia Pacific, you will gain skills, experience, and exposure across the AI industry and be part of shaping what this industry looks like for decades to come.
We are founder-led, not a big corporate. Decisions happen fast, our leaders are accessible, and there's minimum bureaucracy between you and the work. Ownership comes early. Whatever your role, you will have a direct line to outcomes, helping shape how the business grows as we scale nationally across a long-term, large-scale roadmap.
Work alongside founders and experts in AI infrastructure, energy systems and next-generation compute.
What we build here has impact beyond the business. Our AI Factories are designed to operate as assets to the energy grid to actively strengthen the communities and regions they operate in rather than drawing from them.
Considering applying? You don't need a perfect background to join our team. If you're driven and curious, there's a path for you. We back our people to grow into new domains and take on challenges beyond their previous experience.
Role Summary
Firmus Technologies is seeking a skilled Infrastructure Automation Engineer to join our Commissioning team. The position will play a crucial role in the development, maintenance, and improvement of the software tools involved in the commissioning process from initial power-on through integration, phased bring-up, component and system testing, benchmarking, and acceptance testing to hand-over to operations. These tools are used by internal teams, partners, and vendors working together to transform physically built infrastructure into a production-ready platform. This role involves custom tool development, integration with a wide variety of other closed-source and open-source tools, cloud services, and SaaS platforms. We are deploying leading edge AI Factory technology and the tools supporting the commissioning process are critical for successful on-time delivery. This role offers an exciting opportunity to work at the forefront of AI technology and contribute to the growth of AI infrastructure.
Key Responsibilities
Hardware bring-up
In collaboration with internal teams, partners, and vendors, develop custom tools and automate processes for initial system bring-up, hardware configuration, firmware updates, system inventory, testing, and issue tracking. Create accompanying documentation to support the team throughout the commissioning process, for the operations team once the system goes into production, and to accelerate future projects.
System and status monitoring
Work with vendors to implement, configure, and adapt system and status monitoring tools suitable for implementation early in the commissioning phase. Implement health-checks and diagnostics tools which assist remediation teams in finding root cause and resolving issues. Adapt these tools for new platforms, new features, new metrics, and prepare for integration into production systems.
Network tools
Commissioning of the network also requires suitable tools and processes, which need to be automated and streamlined. In close collaboration with the networking team, vendors, and partners implement suitable tools and an efficient process from initial power-on and OS deployment to firmware updates and network configuration. In addition, develop and implement suitable diagnostics and monitoring tools which are critical to resolve issues from the physical layer through to the application layer.
Storage
Develop tooling and automation frameworks supporting deployment, configuration, benchmarking, and validation of storage platforms.
Communication tools
Efficient and effective communication between various teams involved in the commissioning process needs to be supported by suitable tools. Integrate messaging tools, inventory databases, component trackers, issue trackers, time and resource trackers, reporting platforms, dashboards, and operational analytics to avoid duplication and support a fast, reliable, and repeatable commissioning process.
Testing and benchmarking
Develop test and benchmarking tools, customise them for the specific project and customer requirements, and execute the tests to confirm the successful completion of the commissioning process.
Automation and documentation
Leverage modern Infrastructure as Code (IaC) tools to simulate, automate and accelerate the commissioning process. Implement CI/CD pipelines for tool and configuration changes to improve speed, consistency, and auditability. Automate repeatable tasks, document the tools and processes, and prepare them for hand-over to operations teams.
Security
In all tools and throughout the deployment process implement best security practices including security updates and patches, secure handling of credentials, keys, and secrets, secure system setup, network and storage configurations.
Success Measures
Success in this role is measured by:
- Reduced commissioning time through automation.
- High reliability and adoption of commissioning tools.
- Successful automation of repeatable deployment activities.
- Effective monitoring and diagnostic coverage.
- Accurate documentation and operational handover.
- Reduction in manual effort and deployment errors.
Skills & Experience
- Bachelor’s degree in computer science, engineering, or a related technical field.
- 5+ years of experience developing software, automation platforms, operational tooling, or infrastructure automation solutions for large-scale Linux, cloud, HPC, or AI environments.
- Solid understanding of advanced hardware technologies, particularly high performance GPU infrastructure, high end network environments, and high performance storage solutions.
- Experience in Linux systems including server, network, and storage configuration.
- Experience developing production software tools and APIs, consuming and integrating REST APIs, applying modern software development practices, building automated testing frameworks.
- Experience using scripting, automation, and IaC tools including bash, Python, Ansible.
- Experience using infrastructure tools and APIs including Redfish, IPMI, DCGM, Prometheus, Grafana, NetBox, etc.
- Familiarity with AI and HPC tooling such as Slurm, NVIDIA DCGM, NCCL, Kubernetes, and distributed system validation frameworks.
- Excellent problem-solving and analytical skills.
- Ability to work independently and as part of a team.
- Strong communication skills, both written and verbal.
- Willingness to undertake international and/or domestic travel for on-site deployments and commissioning as required.