MLOps Engineer โ Azure & AI/ML Platforms (all genders)
Description
๐ Become our new MLOps Engineer (all genders)
At KI Performance , we move AI from experimentation to production. As a MLOps Engineer , you will design, build, and operate highly scalable, secure Azure-based AI platforms used in production environments. This role sits at the intersection of cloud infrastructure, DevOps, and AI delivery , with a strong focus on enabling iterative AI development , reliable model deployment , and platform scalability .
You will work closely with AI engineers, data teams, and product stakeholders to ensure that AI use cases can be developed, deployed, and operated efficiently at scale โ with production-grade reliability, security, and observability.
Your responsibilities
Cloud Infrastructure & Platform Engineering
-
Design, implement, and operate scalable Azure infrastructure for AI and data-intensive platforms using Terraform
-
Build and maintain secure Azure networking architectures (VNETs, subnets, NSGs, Private Endpoints)
-
Implement access control and governance using Azure RBAC, Key Vault, and Azure Policies
-
Ensure infrastructure is production-ready with a focus on performance, reliability, and scalability
CI/CD & Release Engineering
-
Design and operate modern CI/CD pipelines using GitHub Actions
-
Enable fast, safe, and repeatable deployments for infrastructure, services, and AI models
-
Support iterative development with strong versioning, testing, and rollback strategies
MLOps & AI Platform Enablement
-
Operationalize AI use cases using MLflow (experiment tracking, model registry, deployment workflows)
-
Support the full AI lifecycle from experimentation to production deployment
-
Deploy and operate model inference services exposed via REST APIs (FastAPI preferred)
-
Collaborate closely with AI engineers to ensure models are production-ready
Observability & Reliability
-
Implement end-to-end observability using OpenTelemetry
-
Set up monitoring and logging using Azure Application Insights (or equivalent tooling)
-
Proactively improve system reliability, performance, and incident response
Engineering & Automation
-
Use Python for automation, AI integration, backend services, and tooling
-
Support platform self-service capabilities for engineering teams
-
Continuously improve infrastructure and operational maturity