Senior/Staff DevOps Engineer
Description
MEDvidi is an AI-powered mental healthcare platform setting a new standard for safe, effective, and scalable psychiatric care in the United States.
We combine licensed providers with proprietary AI tools to deliver consistent, outcomes-driven treatment for conditions like ADHD, anxiety, depression, and more. MEDvidi's technology automates charting, follow-ups, and treatment planning, freeing providers to focus on patient care while improving efficiency and clinical quality.
As our Senior/Staff DevOps Engineer , you'll decide where our infrastructure goes next and take the rest of engineering there. This is a hands-on role in the truest sense: you set the technical direction, ship it with your own hands, and own the outcomes for reliability, cost, security, and developer experience across the company. You'll partner directly with our product engineering teams, embedded in what they're building, removing infrastructure friction, and shaping infrastructure around real product needs.
Why this role
- Real ownership, end to end. You own outcomes, not tasks, including your metrics. Reliability, cost, performance, deployment health, and developer velocity belong to you: you set the targets and bring the numbers to the table proactively.
- Technical leadership without the management overhead. Lead by expertise and initiative where the most senior engineers are also the ones in the code. You set direction and raise the bar while staying hands-on.
- AI-native by default. Agentic AI is a first-class part of how we work. You'll delegate to autonomous agents, integrate their output into production changes, and help shape how the whole org builds with AI.
- 6+ years in DevOps/infrastructure engineering, with strong systems fundamentals and solid Linux administration and troubleshooting (performance analysis, resource management, process debugging).
- Hands-on AWS (EC2, EKS, RDS, ElastiCache, Lambda, SQS, EventBridge, API Gateway, ALB, S3) and production Kubernetes/EKS (cluster management, node scaling, policy enforcement; Karpenter, Kyverno, or similar).
- Strong Infrastructure as Code (Terraform and AWS CDK in TypeScript) and CI/CD ownership (GitLab CI/CD: reusable/shared templates, OIDC id_tokens, self-managed GitLab).
- Monitoring and observability in practice (Prometheus, Grafana, OpenTelemetry, OpenSearch, CloudWatch; log-shipping and error tracking/APM).
- Practical security engineering (secrets rotation, short-lived credentials, leak scanning, PHI-aware logging) and HashiCorp Vault as code (KV, JWT/OIDC auth for CI, policy design).
- Blue-green deployments with automated, health-gated rollback; PostgreSQL zero-downtime schema migrations (expand/contract) and migration gating in CI.
- Containers (Docker, ECR, immutable tags, image lifecycle) and network/protocol fundamentals (load balancing, TLS, DNS).
- Hands-on agentic AI workflows (Claude Code or similar): delegating to autonomous agents and integrating their output into production.
- A developer-focused mindset, strong problem-solving for complex system issues, and strong technical writing (docs-as-code, ADRs, design docs via MRs).
- Fluent Russian and English (B1).
- Experience working effectively in remote, distributed teams.
Would be a plus
- Experience in a regulated/compliance-heavy environment (HIPAA, SOC 2, or similar).
- Configuration management (Ansible) for VM fleet management.
- Node.js application operations (pm2, npm), our stack is Node.js + TypeScript.
- GitOps tooling (ArgoCD, Flux) and deeper PostgreSQL database administration.
- AWS certifications.