Performance Engineering Architect
Description
Performance Engineering Architect
This role has been designed as ‘’Onsite’ with an expectation that you will primarily work from an HPE office.
Who We Are
Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.
Job Description
In the HPE Hybrid Cloud, we lead the innovation agenda and technology roadmap for all of HPE. This includes managing the design, development, and product portfolio of our next-generation cloud platform, Green Lake. Working with customers, we help them reimagine their information technology needs to deliver a simple, consumable solution that helps them drive their business results. Join us redefine what’s next for you.
What you’ll do
We are seeking a Performance Expert to join our Performance Engineering team. The engineer will characterize and optimize the performance of AI solutions across GPUs, servers, storage, networking, system software, and AI frameworks.
The role combines hands-on benchmarking, system-level analysis, workload optimization, automation, and collaboration with engineering, product management, field, sizing, and partner teams. The engineer will turn product features, emerging technologies, and customer questions into reproducible performance evidence that supports product design, release readiness, solution sizing, and customer decisions.
What you need to bring:
- Design and execute performance studies for generative AI, machine learning, and other AI workloads.
- Characterize inference and training performance across models, GPUs, precision formats, deployment profiles, and scale configurations.
- Measure latency, throughput, time to first token, inter-token latency, concurrency, scaling efficiency, resource utilization, and cost/performance.
- Evaluate single-GPU, multi-GPU, and multi-node configurations.
- Analyze GPU communication, host-to-GPU bandwidth, PCIe topology, NUMA placement, memory bandwidth, storage, and networking behavior.
- Characterize LLM serving, retrieval-augmented generation, vector database, object detection, speech, multimodal, and agentic AI workloads.
- Develop custom benchmark suites and automation when standard benchmarks do not represent the required workload.
- Identify performance bottlenecks and recommend changes to hardware configuration, BIOS settings, operating systems, drivers, runtimes, frameworks, and applications.
- Establish repeatable test methodologies, baselines, acceptance criteria, and release-performance gates.
- Automate benchmark execution, telemetry collection, result processing, comparison, and reporting.
- Produce clear performance reports, sizing evidence, engineering recommendations, and field-facing guidance.
- Evaluate emerging AI technologies, models, GPU platforms, storage architectures, and distributed inference techniques.
- Collaborate with solution engineering, product management, quality engineering, storage, compute, networking, solution-sizing, field, and technology-partner teams.
Mandatory skills: PCAI, GenAI, and AI workloads
- Strong understanding of generative AI and machine learning performance concepts , particularly inference performance.
- Hands-on experience deploying and benchmarking AI or GenAI workloads on GPU-accelerated systems.
- Understanding of key inference metrics like End-to-end latency, Time to first token, Inter-token latency, Token throughput, Request throughput, Concurrency, Batch size, GPU utilization and memory consumption, Scaling efficiency
- Experience with LLM serving frameworks or inference runtimes such as NVIDIA NIM, vLLM, NVIDIA Triton Inference Server, TensorRT-LLM, Hugging Face, PyTorch, or equivalent technologies.
- Understanding of model precision and quantization formats such as FP32, FP16, BF16, FP8 and NVFP4
- Working knowledge of tensor parallelism, pipeline parallelism, data parallelism, and multi-GPU or distributed execution.
- Experience with containers and container orchestration, including Docker or Podman and Kubernetes or OpenShift.
- Strong Linux administration, troubleshooting, and performance-analysis skills.
- Proficiency in Python and shell scripting for benchmark automation, workload orchestration, telemetry collection, and result processing.
- Experience using AI performance tools or equivalent benchmark frameworks. Relevant examples include: AIPerf, LLMPerf, NVIDIA Perf Analyzer, MLPerf, DLIO, MLPerf Storage, LMCache Bench, VectorDBBench
- Experience monitoring GPU systems using NVIDIA DCGM, nvidia-smi, PyNVML, Prometheus, or equivalent telemetry platforms.
- Ability to correlate application-level performance with GPU, CPU, memory, storage, and network telemetry.
- Sound understanding of statistical measurement, repeatability, run-to-run variation, baseline comparison, and performance-regression analysis.
- Ability to document methodology, configurations, results, limitations, and recommendations clearly.
Domain skills: server hardware and architecture
- Strong understanding of modern server architecture, including:
- CPU sockets, cores, threads, and cache hierarchy
- NUMA architecture and processor affinity
- System memory, DIMM population, channels, bandwidth, and latency
- PCIe generations, lanes, switches, topology, and device placement
- GPU architecture, GPU memory, interconnects, and host-to-GPU data movement
- Local and shared storage
- Ethernet and high-speed networking
- Hands-on experience configuring, testing, and troubleshooting enterprise servers.
- Ability to analyze interactions among CPUs, GPUs, memory, PCIe, storage, networking, operating systems, and applications.
- Understanding of BIOS and firmware settings that affect performance, including power profiles, CPU governors, memory configuration, NUMA settings, and I/O options.
- Experience with server performance tools or equivalent utilities, such as fio, vdbench, perf/netperf, memory and GPU stream benchmarks, MLC, etc
- Ability to isolate system bottlenecks and distinguish application limitations from hardware, firmware, operating-system, storage, or network constraints.
- Understanding of performance, price/performance, scalability, efficiency, and performance-per-watt considerations.
- Experience developing reproducible server configurations, test procedures, and performance baselines.
- Familiarity with enterprise hardware qualification, firmware and driver compatibility, and controlled configuration management.
Nice-to-have skills: virtualization
- Experience with one or more virtualization platforms:
- VMware ESXi and vSphere
- KVM
- Red Hat OpenShift Virtualization
- SUSE Harvester
- Understanding of virtual CPU, virtual NUMA, memory overcommitment, device passthrough, SR-IOV, and GPU virtualization.
- Experience comparing virtualized and bare-metal performance.
- Familiarity with VMware tools such as esxtop, vSAN Observer, and VMmark.
- Understanding of virtualized storage and networking performance.
- Experience diagnosing contention, noisy-neighbor effects, resource scheduling, and virtualization overhead.
- Familiarity with Kubernetes scheduling, GPU operators, MIG, and accelerator allocation in containerized or virtualized environments.
Additional preferred experience
- Experience with performance benchmarks such as SPEC CPU, SPECjbb, HammerDB, TPC, SAP, or similar industry benchmarks.
- Exp