Senior SRE Engineer

SeniorHybrid
CompanyCodeway
LocationBarcelona
CategorySoftware Engineering
DepartmentHQ - Tech
SenioritySenior
WorkplaceHybrid
Posted2026-07-15
Viaashby

Description

ABOUT CODEWAY

Codeway builds category-leading consumer AI apps on mobile and web, at global scale.

600M+ downloads. 60+ apps. 400+ builders across Barcelona and İstanbul. Completely bootstrapped. No board. Profitable since year two.

Small teams move faster than big ones. But they need big resources to win at scale. Nobody starts from zero. The problem decides the solution, not us.

We turn category wins into six focused vertical companies: Wishlabs (AI Creativity), Wellture (Wellness), Learna (Education), Sparked (Media), Cosmic Jam (Entertainment), and Catalyst Apps (Utilities & Productivity). Each has full ownership and the depth to go further than anyone else in its space.

Codeway HQ powers them all. One platform. One team that finds and grows builders. One set of guardrails. Every idea born here starts with an unfair advantage: the platform, the people, and the profits of every product before it. That's the system.

Turns out you can do both.

Move like an indie. Hit like a giant.

POSITION

We’re looking for a Senior Site Reliability Engineer to own and mature reliability, performance, and security across our growing platform. This role sits at the intersection of Engineering, Infrastructure, and Security, helping design, operate, and continuously improve the systems that keep dozens of consumer apps running for users around the world.

You’ll work closely with product engineering teams to make reliability measurable rather than assumed. That means defining and enforcing SLIs, SLOs, and error budgets; operating and hardening a multi-cluster Kubernetes environment; building the observability that catches problems before users feel them; and leading incident response when things break. It’s a hands-on role with real ownership over how reliability and security evolve as we scale.

Several parts of our reliability practice are still early in their maturity. We’re looking for someone who enjoys building the standards, processes, tooling, and automation that will form the foundation of our SRE function — not someone waiting for a playbook to already exist.

We welcome applicants from all backgrounds and experiences. If you’re excited about running systems at consumer scale and believe you could be a strong fit, we encourage you to apply, even if your experience doesn’t align perfectly with every qualification listed below.

WHAT YOU’LL BE DOING

Reliability, SLOs, Observability

  • Define, instrument, and report on SLIs, SLOs, and error budgets across critical services, so reliability decisions are driven by data rather than opinion.
  • Own observability end-to-end — metrics, logs, traces, dashboards, and alerting — and drive measurable reductions in detection and resolution times.
  • Reduce alert noise and false positives so on-call engineers can trust what wakes them up.
  • Run reliability reviews and an error-budget policy that shapes how teams prioritize between shipping and stability.

Kubernetes & Platform Operations

  • Operate, scale, and upgrade our multi-cluster Kubernetes (GKE) environment: cluster lifecycle, autoscaling, networking, ingress, and resource management.
  • Act as the deep-expertise escalation point for cluster and platform issues across dozens of services.
  • Own capacity planning, performance, and cloud cost efficiency, balancing spend against reliability targets.
  • Build self-service platform tooling that lets product teams move quickly without needing to become infrastructure experts.

Security & Resilience

  • Embed security into the platform through RBAC and least-privilege, secrets management, image and dependency scanning, network policies, and a disciplined patching cadence.
  • Partner with the security function on vulnerability remediation, audit readiness, and secure-by-default infrastructure.
  • Own disaster recovery: define and regularly validate RTO/RPO targets through DR drills and failure testing.
  • Contribute to architecture and production-readiness reviews so reliability and security are designed in, not bolted on.

Incident Response & Automation

  • Lead the on-call rotation and act as incident commander during production incidents.
  • Run blameless postmortems, quantify impact, and track corrective actions through to closure so the same failure doesn’t recur.
  • Build and maintain Infrastructure as Code (Terraform) and CI/CD pipelines, enforcing GitOps and progressive delivery with automated rollbacks.
  • Systematically identify, measure, and eliminate operational toil through automation, protecting engineering time for high-leverage work.

WHAT YOU’LL BRING?

  • Experience operating high-traffic, always-on production systems at meaningful scale, typically gained over 5–8 years in SRE, Platform, or DevOps roles.
  • Hands-on production Kubernetes experience — you’ve run clusters day to day, through upgrades, autoscaling, and real troubleshooting under load, not just deployed to them.
  • A strong cloud engineering background, along with solid Linux and networking fundamentals.
  • A track record of defining and operating with SLOs and error budgets, and comfort being measured on reliability outcomes.
  • Experience with Infrastructure as Code and CI/CD pipeline design — you treat infrastructure and delivery as code.
  • Depth in observability tooling: instrumentation, dashboarding, and alert design.
  • A genuine security-first mindset, where least-privilege, secrets hygiene, and vulnerability management are habits rather than afterthoughts.
  • Scripting and automation fluency in at least one language, used to build tooling and remove toil.
  • Incident-command experience: owning on-call, running blameless postmortems, and driving resolution times down over time.
  • Ability to communicate clearly with both engineers and leadership, especially under pressure.

NICE TO HAVE

  • Experience with high-scale consumer or mobile app backends, or with AI/ML inference workloads and their scaling characteristics.
  • Experience with GitOps and progressive-delivery patterns such as canary and blue-green rollouts.
  • Familiarity with service mesh, API gateways, or multi-region and multi-cluster topologies.
  • Cloud cost optimization and FinOps discipline at scale.
  • Exposure to compliance initiatives (SOC 2, ISO 27001, GDPR) and broader DevSecOps practice.
  • Chaos engineering or resilience testing experience.
  • Relevant certifications in Kubernetes, cloud, or DevOps disciplines.
  • Experience supporting many independent services and teams concurrently in a fast-shipping, product-led environment.

OUR ENVIRONMENT

You’ll help operate and improve a modern, cloud-native environment built around:

  • Kubernetes (GKE) and containerized workloads
  • Google Cloud Platform (GCP)
  • Terraform and Infrastructure as Code
  • CI/CD and GitOps tooling
  • Modern observability (metrics, logs, traces, alerting)
  • Cloudflare CDN and edge

Experience with these exact platforms is beneficial but not required. We value strong fundamentals, curiosity, and the ability to quickly learn new technologies and environments.

WHAT SUCCESS LOOKS LIKE

Within your first 12 months, you’ll help establish and mature key reliability capabilities, including:

  • Clear SLIs, SLOs, and error budgets live and reported for our most critical services.
  • A measurable reduction in detection and resolution times for production incidents.
  • A consistent, blameless incident response practice, with postmortem actions tracked to closure.
  • A hardened Kubernetes fleet, with key security gaps closed across access control, secrets, scanning, and patching.
  • Expanded Infrastructure as Code and observability coverage across the platform.
  • Disaster recovery drills that reliably pass agreed RTO/RPO targets.
  • Operational toil trending down against an explicit target, with automation replacing manual work.
  • Product teams self-serving on standardized, secure, well-instrumented platform tooling.

Just as importantly, the measurement remit here doesn’t stop at the engineeri