Member of Technical Staff, Performance Engineering
Description
OVERVIEW
Physical Superintelligence is a startup with roots at Google, NVIDIA, Harvard, Meta, MIT, Oxford, Johns Hopkins, Cambridge, and the Perimeter Institute building AI systems to discover new physics at scale. We are seeking a performance engineer who owns a workload's time end to end, from the application code and the data path through the scheduler to the GPU kernels, across AI and HPC, to make one cluster produce more physics.
Our mission is to discover and commercialize transformative physics breakthroughs at scale with artificial superintelligence, safely, verifiably, and for broad public benefit.
The last century's golden age of physics gave us transistors, lasers, and nuclear energy. We believe artificial superintelligence will unlock the next one. We're creating the infrastructure to industrialize scientific discovery and usher in this new era.
We have one product: new physics, at scale.
ROLE AND RESPONSIBILITIES
- Own a workload's time end to end. Your unit of work is a workload, not a box or a layer: a reinforcement-learning training loop, a multi-node physics simulation, an open-weight model we serve, a numerical campaign on an open physics problem, a coupled multiphysics engine. Its time is spent in application code, data loading and preprocessing, memory, storage and checkpoint I/O, container and queue overhead, service round trips, network collectives, and GPU kernels. All of it is yours. When it runs slow you find which layer and you fix it there. Where a researcher owns the code you land the fix in their repository and leave a benchmark behind.
- Write the kernels. Physics campaigns, simulation engines and learned-model work need custom GPU code: exact and extended-precision numerics, sparse and dense linear algebra, adaptive meshes and stencils, attention and cache paths for serving. You write them in CUDA C++ or Triton, profile them, and turn the good ones into a library other people call.
- Make simulations run at scale. Fluid, thermal, electrical, molecular and coupled multiphysics solvers, both established codes and the ones physicists write themselves, run on the cluster as multi-node MPI jobs: domain decomposition and strong scaling, GPU ports of solver kernels where the measurements say they pay, I/O and checkpointing, and continuous data generation for learned models. You decide with numbers when a solver belongs on GPUs and when it does not.
- Make training fast on our nodes. Distributed and reinforcement-learning training on the cluster: sampler and trainer placement, collective communication, checkpoint cost, resume from failure. You know where step time goes, what MFU means for an RL loop as against a dense pretraining run, and how to close the gap.
- Stand up and tune serving on our own capacity. Deploy open-weight models on our GPUs, choose the serving engine and the quantization, and make the endpoint cheaper per token than the frontier API it replaces. You own the throughput and latency numbers and the decision of which requests move in-house.
- Measure the cluster instead of guessing. Per-workload utilization and goodput, taken on real jobs, so that queue policy, preemption classes and the next capacity decision rest on numbers somebody observed. You supply the measured GPU-hours per project that the quarterly rent-or-expand call is made on.
WHAT WE'RE LOOKING FOR
- Whole-stack performance ownership. You have taken a workload from slow to fast across more than one layer, with a trace that shows where the time went before and after, and at least one of those layers was not the GPU: a data loader, a filesystem, a scheduler, a container start, a service call, a Python hot loop. You work from measurements, instrument systems you did not write, and verify a win on a real workload before you claim it.
- Kernels you wrote yourself. Five or more years making GPU code fast, including CUDA C++ or Triton kernels that other people's systems depend on. You read Nsight Compute output the way a physicist reads a plot, and you can say where a kernel sits on the roofline and why.
- Scale on at least one side, and the appetite for the other. Either real multi-node AI training experience (distributed training throughput, collective-communication behavior, InfiniBand topology, where step time goes and why MFU disappoints) or HPC simulation at scale (a CFD, molecular dynamics, multiphysics or comparable code taken across nodes with MPI, scaling measured, solver kernels ported or tuned on GPUs, FP64 and extended-precision numerics). Half the demand on this cluster is simulation and physics numerics and half is AI; you bring one and learn the other on our hardware.
- You can explain a performance result to a researcher who will never open a profiler, and your numbers are usable by people who make decisions with them.
NICE TO HAVE
- Both sides of the scale requirement: multi-node AI training and HPC MPI simulation.
- Inference serving for large open-weight models: vLLM, SGLang or comparable, multi-GPU placement of mixture-of-experts models, KV-cache and quantization trade-offs.
- Source-level work in OpenFOAM, LAMMPS, GROMACS, Nek5000, or comparable codes; AMR or stencil frameworks; PETSc, Trilinos, MAGMA or comparable libraries.
- CUTLASS, CuTile, or comparable kernel frameworks; compiler-level work in Triton or MLIR.
- Kueue or Slurm queue policy on a shared cluster: priority classes, preemption, bin-packing, and the before-and-after that justified a policy change.
- You have owned a capacity decision with money attached and been held to the number.
- Experience turning research code into a maintained library that a team of scientists depends on.
HOW WE WORK
We hold a high technical bar and give people full ownership of their work, from spec to ship to on-call. We write contracts before logic, test against real systems instead of mocks, and favor simple designs that ship over clever ones that do not. Our development process is AI-native: we work with agentic coding tools daily, write specs that are legible to humans and agents alike, and lead with leverage.
LOCATION AND COMPENSATION
This role is based in Boston. We will consider remote candidates on a case-by-case basis. We offer competitive compensation including salary, benefits, and meaningful early-stage equity. We evaluate on whole-stack performance ownership, GPU depth across AI and HPC, measurement discipline, and shipping velocity. We are an equal opportunity employer and value diverse perspectives in building platforms for AI-driven discovery.