UX Engineer (AI and Applications)
Description
Firmus Technologies
Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.
Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.
At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.
Firmus AI Cloud
Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.
It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.
Why Firmus?
As an NVIDIA Cloud and Engineering partner in Asia Pacific, you will gain skills, experience, and exposure across the AI industry and be part of shaping what this industry looks like for decades to come.
We are founder-led, not a big corporate. Decisions happen fast, our leaders are accessible, and there's minimum bureaucracy between you and the work. Ownership comes early. Whatever your role, you will have a direct line to outcomes, helping shape how the business grows as we scale nationally across a long-term, large-scale roadmap.
Work alongside founders and experts in AI infrastructure, energy systems and next-generation compute.
What we build here has impact beyond the business. Our AI Factories are designed to operate as assets to the energy grid to actively strengthen the communities and regions they operate in rather than drawing from them.
Considering applying? You don't need a perfect background to join our team. If you're driven and curious, there's a path for you. We back our people to grow into new domains and take on challenges beyond their previous experience.
Role Summary
The UX Engineer will define and deliver intuitive, coherent, and technically credible user experiences across the AI & Applications team’s internal and external product portfolio. The role will shape how users discover, provision, operate, optimize, and troubleshoot AI infrastructure and AI-powered services through products such as the AI Cloud Portal, Global Operations Console, Bare Metal as a Service, Kubernetes as a Service, and future X-as-a-Service offerings.
This role sits at the intersection of user experience design, front-end engineering, product thinking, and AI-platform understanding. The UX Engineer will translate complex infrastructure and AI capabilities—including GPU capacity, bare-metal provisioning, Kubernetes clusters, workload scheduling, benchmark results, model recipes, inference services, Model-to-Grid optimization, and agentic operations—into clear workflows, understandable interfaces, and actionable experiences for customers, developers, AI researchers, support teams, operators, and internal delivery teams.
A central focus will be making sophisticated AI-factory capabilities usable. The UX Engineer will help users understand and act on model-to-grid recommendations, scheduler decisions, workload placement, cluster health, benchmark outcomes, capacity availability, performance bottlenecks, inference cost, energy or power signals, and agent-generated operational insights. The role will ensure that advanced automation remains transparent, controllable, auditable, and trustworthy—particularly where agentic systems recommend or execute actions.
The UX Engineer will report to the Head of AI & Applications, working closely with other AI engineers, DevOps and scheduler engineers, inference and optimization engineers, platform,infrastructure, security, global operations, and customer-facing teams. The role will create a scalable design system and product experience that supports both immediate platform needs and future AI-factory and XaaS product expansion.
Key Responsibilities
- Own end-to-end UX design and front-end experience quality for the AI Cloud Portal, Global Operations Console, Bare Metal as a Service, Kubernetes as a Service, and future XaaS products.
- Conduct user discovery and workflow analysis with key personas, including AI researchers, ML engineers, application developers, cloud administrators, tenant administrators, infrastructure operators, support engineers, data-centre teams, and external customers.
- Translate product requirements, user needs, operational workflows, technical constraints, and system telemetry into user journeys, information architectures, wireframes, prototypes, interaction patterns, and production-ready interfaces.
- Design self-service experiences for discovering, requesting, provisioning, configuring, accessing, monitoring, scaling, updating, and decommissioning bare-metal GPU capacity and Kubernetes environments.
- Design clear workflows for GPU and AI workload management, including job submission, job templates, model and workload recipes, queue selection, quota visibility, priority requests, capacity reservations, workload status, failures, retries, and support escalation.
- Create usable interface patterns that explain proprietary scheduler decisions, including why a job is queued, admitted, deferred, placed on a particular topology, pre-empted, rescheduled, or unable to run.
- Design Model-to-Grid experiences that help users and operators understand the relationship between model or workload requirements, benchmark results, GPU resources, network topology, storage, capacity, power, thermal conditions, and scheduling outcomes.
- Create user-facing workflows for selecting and applying validated model, training, fine-tuning, and inference recipes based on workload goals such as latency, throughput, accuracy, cost, energy efficiency, GPU availability, or time-to-results.
- Design intuitive benchmarking experiences, including benchmark configuration, execution tracking, result exploration, comparison of baselines and variants, reproducibility metadata, bottleneck identification, and actionable optimizationrecommendations.
- Shape the Global Operations Console experience for monitoring and managing fleet-level AI-factory operations, including cluster health, capacity, workload demand, infrastructure events, service status, scheduler health, resource utilization, maintenance activities, and operational incidents.
- Make complex infrastructure status and telemetry actionable through dashboards, drill-down views, alerts, guided troubleshooting, root-cause context, recommended actions, and role-appropriate escalation paths.
- Design agentic application experiences that clearly distinguish between agent observations, recommendations, planned actions, approved actions, completed actions, failures, and human intervention points.
- Create safe and trustworthy human-in-the-loop workflows for agentic operations, including approval queues, permission scopes, action previews, confidence or evidence views, audit trails, rollback options, and incident escalation.
- Design user experiences for AI platform identity, access, tenancy, quotas, approvals, notifications, billing or consumption visibility, support, and administrative controls in partnership with Security and platform teams.
- Develop and maintain a scalable design system, component library, interaction standards, accessibility guidelines, content patterns, and visual language suitable for cloud, AI-platform, and operations-console products.
- Implement high-quality front-end experiences or work closely with