MLOps Engineer โ€“ Azure & AI/ML Platforms (all genders)

Hybrid
CompanyKI group
LocationCologne, Nordrhein-Westfalen, Germany
CategorySoftware Engineering
DepartmentKI performance
Seniority-
WorkplaceHybrid
Posted2026-01-07
Viarecruitee

Description

๐Ÿš€ Become our new MLOps Engineer (all genders)

At KI Performance , we move AI from experimentation to production. As a MLOps Engineer , you will design, build, and operate highly scalable, secure Azure-based AI platforms used in production environments. This role sits at the intersection of cloud infrastructure, DevOps, and AI delivery , with a strong focus on enabling iterative AI development , reliable model deployment , and platform scalability .

You will work closely with AI engineers, data teams, and product stakeholders to ensure that AI use cases can be developed, deployed, and operated efficiently at scale โ€” with production-grade reliability, security, and observability.

Your responsibilities

Cloud Infrastructure & Platform Engineering

-
Design, implement, and operate scalable Azure infrastructure for AI and data-intensive platforms using Terraform

-
Build and maintain secure Azure networking architectures (VNETs, subnets, NSGs, Private Endpoints)

-
Implement access control and governance using Azure RBAC, Key Vault, and Azure Policies

-
Ensure infrastructure is production-ready with a focus on performance, reliability, and scalability

CI/CD & Release Engineering

-
Design and operate modern CI/CD pipelines using GitHub Actions

-
Enable fast, safe, and repeatable deployments for infrastructure, services, and AI models

-
Support iterative development with strong versioning, testing, and rollback strategies

MLOps & AI Platform Enablement

-
Operationalize AI use cases using MLflow (experiment tracking, model registry, deployment workflows)

-
Support the full AI lifecycle from experimentation to production deployment

-
Deploy and operate model inference services exposed via REST APIs (FastAPI preferred)

-
Collaborate closely with AI engineers to ensure models are production-ready

Observability & Reliability

-
Implement end-to-end observability using OpenTelemetry

-
Set up monitoring and logging using Azure Application Insights (or equivalent tooling)

-
Proactively improve system reliability, performance, and incident response

Engineering & Automation

-
Use Python for automation, AI integration, backend services, and tooling

-
Support platform self-service capabilities for engineering teams

-
Continuously improve infrastructure and operational maturity