Platform Engineer

CompanyMako
LocationLondon
CategoryInfrastructure
DepartmentFront Office Infrastructure
Seniority-
Workplace-
Posted2026-09-24
Viagreenhouse

Description

At Mako, we believe in the power of collaboration to drive innovation in pursuit of our collective ambition; excellence in trading. Our diverse community is connected through a commitment to being the best we can be with the highest standards of integrity.

We're looking for a Platform Engineer to help design, build, and operate our on-premises Kubernetes platform running on a self-managed Linux VM fabric (KVM/libvirt-based hypervisor layer). You'll own the infrastructure that sits beneath our application workloads — from the hypervisor and VM provisioning up through Kubernetes cluster lifecycle, networking, storage, and the golden-path tooling that application teams use to ship software.

This is a hands-on, deeply technical role for someone who enjoys operating infrastructure at the systems level — not a cloud-managed-service consumer role. The VM fabric itself is still being designed and built out, so you'll have real influence over its architecture, not just its day-to-day operation. You'll be responsible for keeping the fabric and the clusters running on top of it healthy, secure, and performant, without the safety net of a hyperscaler's managed control plane.

What you’ll be involved in

  • Design the Linux VM fabric underpinning the platform from the ground up: hypervisor architecture (KVM/libvirt), host networking topology, storage backing for VM disks, and how the fabric will scale as workload demand grows
  • Operate and maintain the hypervisor layer across multiple physical hosts, including host patching, live migration/evacuation, and failure response with minimal workload disruption
  • Design and maintain a distributed shared storage system such as Ceph, or an alternative
  • Design and maintain VM templating and golden-image pipelines so Kubernetes nodes are provisioned consistently and can be rebuilt or rotated on demand
  • Automate the VM lifecycle end-to-end — provisioning, scaling, patching, decommissioning — via infrastructure-as-code
  • Manage compute, memory, and storage capacity planning across the fabric, including host-level oversubscription strategy and headroom for node failure or maintenance
  • Own virtual networking within the fabric — host networking, VLANs/overlay networks — and design how it hands off cleanly into the Kubernetes CNI layer above it
  • Design, build, and maintain the full lifecycle of on-prem Kubernetes clusters: bootstrapping, version upgrades, node scaling, and decommissioning, using tooling such as kubeadm, Cluster API, or Kubespray
  • Manage the control plane end-to-end, including etcd operations (backup/restore, performance tuning, disaster recovery), since there's no managed control plane to fall back on
  • Configure and tune cluster networking: CNI selection, network policy enforcement, and on-prem load balancing
  • Stand up and manage ingress and internal DNS for workloads across environments
  • Own persistent storage integration for stateful workloads via CSI drivers
  • Define and enforce multi-tenancy patterns across both layers — tenant isolation and resource allocation on the VM fabric (compute, storage, network) as well as namespace/resource quota strategy, RBAC, and policy enforcement (OPA/Gatekeeper or Kyverno) at the Kubernetes layer
  • Build and maintain GitOps-based delivery for both cluster configuration and workloads (ArgoCD or Flux), treating cluster and infrastructure state as code
  • Harden hosts and clusters against security baselines (CIS benchmarks for Linux and Kubernetes), manage secrets (Vault, sealed-secrets), and keep the container runtime and node OS patched
  • Build observability across the full stack — from hypervisor/host health up through cluster metrics and logs (we currently use Prometheus, Grafana, OpenSearch, and Checkmk; open to alternatives) — with particular focus on the capacity and failure signals a managed cloud provider would normally surface for you
  • Plan and execute Kubernetes version upgrades and node OS/kernel upgrades with minimal workload disruption
  • Design and maintain disaster recovery and backup strategy spanning both layers — VM snapshots/backups and etcd/cluster state — so the platform can be rebuilt from bare infrastructure if required
  • Troubleshoot incidents across the entire stack — from a misbehaving pod, down through kubelet, container runtime, and CNI, into the underlying VM and hypervisor layer when needed
  • Participate in an on-call rotation for platform-level incidents; drive root-cause analysis and post-incident reviews
  • Partner with application teams to define and support a smooth developer experience (self-service namespaces, CI/CD integration, internal developer platform tooling)
  • Coordinate with datacentre/facilities and network teams on physical host provisioning, rack capacity, and hardware refresh cycles
  • Contribute to the platform roadmap: capacity growth, tooling upgrades, and reducing operational toil through automation

What we need from you

Essential

  • Solid production experience running Kubernetes in a self-managed, on-premises context (not just EKS/GKE/AKS) — you understand what breaks when there's no managed control plane, and you've operated etcd and the control plane yourself
  • Hands-on experience designing and operating a Linux KVM/libvirt-based VM fabric as the foundation for Kubernetes — host architecture, templating, and provisioning automation, ideally from relatively early stage rather than just inheriting a mature environment
  • Strong Linux systems administration background (networking, storage, service management, kernel tuning, troubleshooting under pressure)
  • Practical, in-depth knowledge of Kubernetes networking (CNI internals, service meshes a plus) and storage (CSI drivers, distributed storage systems such as Ceph/Longhorn)
  • Experience with infrastructure-as-code and configuration management (Terraform, Ansible, Packer)
  • Experience with GitOps workflows and CI/CD pipelines
  • Comfortable with observability stacks (Prometheus/Grafana, ELK/Loki) and using them to diagnose infra issues without cloud-native tooling
  • Security-conscious: familiar with hardening standards, RBAC, network segmentation, and secrets management
  • Strong troubleshooting instincts across the full stack — hypervisor, OS, network, container runtime, Kubernetes control plane
  • Good written and verbal communication; comfortable working with distributed/hybrid teams

Desirable

  • Experience with bare-metal Kubernetes provisioning
  • Background in a regulated or air-gapped/restricted-network environment
  • Contributions to open-source infrastructure tooling
  • Experience using AI tooling (e.g. AI coding assistants, LLM-based automation) to accelerate development — we're keen to use AI to speed up our development cycle
  • Experience running AI infrastructure.

Why This Role

You'll have real ownership over infrastructure end-to-end — from the hypervisor to the pod — with no black-box managed services standing between you and root cause. If you like understanding systems all the way down and want to shape how a platform team operates outside the public cloud, this is that role.

We're a FOSS-first company for our on-prem infrastructure: we build our platform on open-source tooling where it fits, rather than defaulting to proprietary or vendor-locked products. That means fewer licensing constraints on how you design solutions and direct access to the source when something needs to be understood or fixed at depth.

##

We are Mako

At Mako, we are welcoming, inclusive and collaborative. We work fast and smart in a supportive and dress-down environment that allows colleagues to be themselves and achieve great things. We uphold the principles of a flat structure that offers unrivalled engagement with senior leadership and career development opportunities. We have a comprehensive benefits package, including:

  • Flexible leave and hybrid working policies
  • Private health and dental insurance
  • Generou