Nebius is a technology company building a full-stack cloud infrastructure designed for the explosive growth of the AI industry. We provide AI-centric cloud services, purpose-built for AI innovators and their workloads.
About the role
We are looking for a Software Engineer to join our Observability team. You will help build and operate the monitoring, logging, and tracing systems that keep Nebius AI Cloud reliable at scale for our customers’ training and inference workloads.
What you’ll do
- Design, implement, and maintain observability platforms and tooling (metrics, logs, traces, profiling)
- Build reliable data pipelines and storage for high-cardinality telemetry at cloud scale
- Partner with engineering teams to improve service reliability, incident response, and SLO/SLI practices
- Develop automation and self-service capabilities for dashboards, alerts, and runbooks
- Contribute to performance analysis and capacity planning using observability data
- Participate in on-call rotation and help drive improvements from incidents and postmortems
What we’re looking for
- Strong software engineering background with experience building production systems
- Experience with observability stacks and cloud-native infrastructure (e.g., Prometheus, Grafana, OpenTelemetry, ELK/EFK, Jaeger/Tempo)
- Solid knowledge of Linux, networking, and distributed systems fundamentals
- Experience with Kubernetes and running services in production environments
- Proficiency in at least one backend language (e.g., Go, Python, Java, C++)
- Ability to communicate clearly and collaborate across teams
Nice to have
- Experience supporting GPU/ML workloads or large-scale batch/training jobs
- Experience with SRE practices, incident management, and reliability engineering
- Contributions to open-source observability projects
What we offer
- Competitive compensation and benefits
- Opportunity to work on cutting-edge AI infrastructure at global scale
- Collaborative, engineering-driven culture