Quick Overview
Job Description
About Nebius:
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
Infrastructure Engineer
Independently operates, maintains, troubleshoots, and takes end-to-end ownership of critical power, critical cooling, liquid cooling, and DCIM/BMS infrastructure systems; participates in incident management and commissioning activities; and mentors Associate Infrastructure Engineers.
Role Overview
The Infrastructure Engineer is a critical role within Nebius data center operations. A fully proficient Infrastructure Engineer who independently manages and troubleshoots critical power and critical cooling systems across the site and takes ownership of assigned engineering tasks from start to finish. Participates in incident management activities, supports site commissioning and build reviews, and contributes to ensuring all engineering work meets Nebius SLAs and standards. Enforces safe working practices, collaborates with facilities, network, and technician teams, and mentors Associate Infrastructure Engineers toward independent operation.
Core Responsibilities
- Monitor and contribute to the tracking of critical environment maintenance and repair for assigned service lines to Nebius SLAs; identify and escalate developing faults before they impact service uptime.
- Independently operate, monitor, and troubleshoot critical power distribution infrastructure including UPS systems, PDUs, RPPs, and in-rack busbar power systems.
- Manage and maintain critical cooling infrastructure including CRAHs, CRACs, CDUs, and RDHx; perform capacity checks and leakage inspections.
- Participate in incident management for infrastructure-impacting events; support root-cause analyses, document findings, and contribute to CAPA execution under the direction of the Senior Infrastructure Engineer.
- Operate and administer DCIM and BMS platforms; build and maintain dashboards, alerts, capacity reports, and infrastructure records.
- Support deployment and commissioning of GB-scale liquid-cooled rack infrastructure including direct liquid cooling (DLC) systems, manifolds, and CDU connections.
- Perform and own preventive maintenance tasks for all critical power and critical cooling systems; maintain accurate maintenance logs and compliance records.
- Participate in site reviews, design reviews, and commissioning activities; prepare and review technical reports to document findings and communicate results.
- Enforce Nebius critical power and critical cooling safety procedures; act as safety lead for engineering work orders on live infrastructure.
- Collaborate with facilities, network, and technician teams on infrastructure changes, capacity expansions, and major deployments.
- Support adherence to Nebius standards and policies through documentation review and active participation in commissioning and design activities.
- Contribute to vendor and contractor coordination by supporting scheduling, site access, and execution of work per Nebius expectations and safe-working practices.
- Actively mentors and trains Associate Associate Infrastructure Engineers; guides them through systems operation, safe working practices, and structured skill development toward independent operation.
Expected Capabilities / Expectations
- Working proficiency across all critical power systems: UPS, PDU, RPP, in-rack busbar, and generator interfacing.
- Hands-on experience with critical cooling systems: CRAH/CRAC, CDU, RDHx, and water-side infrastructure including leak detection.
- Competent DCIM/BMS administration: alert management, capacity modeling, and dashboard reporting.
- Developing knowledge of GB rack liquid cooling technology, DLC manifolds, and heat rejection infrastructure.
- U.S. Data Center Career Framework - Technical Career Framework Ability to participate in and support incident management activities including RCA and CAPA documentation.
- Strong documentation discipline: change records, maintenance logs, capacity data, and incident write-ups.
- Clear and confident communication with facilities, network, and vendor teams during changes and incidents.
Core Physical Requirements
- Ability to stand and remain on your feet for extended periods, typically four or more consecutive hours during active shift operations.
- Ability to safely lift, carry, and position equipment weighing up to 50 pounds unassisted, and heavier loads with appropriate team-lift protocols or mechanical aids.
- Comfortable working at heights including ascending and descending ladders, raised platform equipment, and elevated data center
- infrastructure.
- Ability to work in confined spaces such as under raised floors, within enclosed rack enclosures, and in cable management pathways.
- Comfortable working in environments with variable temperatures, including cold aisle containment zones and active cooling infrastructure.
- Manual dexterity sufficient to handle small form-factor components, precision cabling, and fine connector installations.
- Visual acuity sufficient to read equipment labels, small-form-factor interface indicators, and detailed wiring diagrams in variable lighting conditions.
- Ability to push or pull equipment carts, server sleds, and wheeled infrastructure weighing up to 500 pounds on level surfaces.
On-Call Requirement
- This role includes on-call participation to respond to after-hours critical power and critical cooling events, infrastructure alarms, and urgent maintenance requiring prompt engineering response.
Pay Transparency
We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law.
Benefits & Perks:
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
What's it like to work at Nebius:
Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI
Equal Opportunity Statement:
Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.
Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.
If you need accommodations during the application process, please let us know.