Nebius is a technology company building a full-stack AI-native cloud platform for AI practitioners. We provide large-scale GPU infrastructure, managed services, and tools for training and inference workloads.
About the role
We are looking for a Senior Cluster Operations Engineer to help operate and scale Nebius AI Cloud GPU clusters. You will work closely with infrastructure, platform, and customer-facing teams to ensure reliable, high-performance compute for AI training and inference.
What you’ll do
- Operate and maintain large-scale GPU clusters and supporting infrastructure (networking, storage, schedulers).
- Monitor cluster health, troubleshoot incidents, and drive root-cause analysis and remediation.
- Automate operational tasks and improve observability, alerting, and runbooks.
- Partner with SRE and platform engineering on capacity planning, upgrades, and rollouts.
- Support internal and external customers on performance, stability, and best practices for distributed workloads.
- Contribute to on-call rotation and incident response for production environments.
What we’re looking for
- 5+ years of experience in production infrastructure, SRE, or HPC/cluster operations.
- Strong Linux administration skills and experience with automation (Python, Go, or similar).
- Hands-on experience with Kubernetes, Slurm, or other workload schedulers in production.
- Familiarity with GPU servers, high-speed networking (InfiniBand/RoCE), and distributed storage.
- Experience with monitoring stacks (Prometheus, Grafana, ELK, or similar) and incident management.
- Clear communication skills and ability to work across engineering and customer teams.
Nice to have
- Experience supporting ML training or inference at scale.
- Background in cloud providers or bare-metal AI infrastructure.
- Contributions to internal tooling for fleet management and diagnostics.
What we offer
- Competitive compensation and benefits.
- Opportunity to work on cutting-edge AI infrastructure at global scale.
- Collaborative engineering culture with strong technical ownership.
About Nebius
Nebius builds AI-native cloud infrastructure designed for the full lifecycle of AI development—from experimentation to large-scale training and production inference.