HireFT
Browse JobsHow it worksPricingAboutSuccess Stories
    Back to jobs
    CR

    Crux AI

    Data Center Operations Manager (Site, Facility Operations)

    Palo Alto, United StatesHybridFull-time10+ yrs experiencePosted 1w ago
    All Crux AI jobs

    Job description

    Built to set the gold standard for integrated AI infrastructure

    Crux AI is a newly formed, U.S.-based integrated AI infrastructure company created to remove the physical and operational constraints on consequential AI ambitions. Crux brings together power, high-density data centers, TPU silicon, networking, orchestration software, and ongoing operations as one integrated system.

    Crux is being capitalized to plan every layer together, develop each one to demanding standards, and operate the whole system with efficiency and reliability. That gives hyperscalers, frontier AI labs, sovereign customers, enterprises, and AI-native companies greater freedom to pursue the AI they are here to create.

    WHAT YOU'LL DO

    You will own critical facility operations for one Crux site or campus — the electrical, mechanical, fire/life-safety, and controls infrastructure that keeps high-density, liquid-cooled TPU clusters continuously available. You will lead the 24/7 site facility team — chief engineer, shift technicians, and O&M contractors — and run the maintenance, procedures, drills, and change control that produce uptime, whether the building is Crux-owned or operated by a colocation partner you hold to SLA. In the early phase you will help bring the site to life: commissioning participation, acceptance, and the transition from construction to disciplined steady-state operations.

    In this role, you will:

    • Establish and direct the site facility operations team — chief engineer, critical facility technicians, and contractors — with a 24/7 coverage model you design, staff, and train, maintaining highly available operations for all critical and non-critical infrastructure.

    • Operate and maintain the site’s electrical systems (utility interface through switchgear, generators, UPS, and distribution), mechanical plant (chilled water, direct-to-chip liquid cooling loops, CDUs, heat rejection), fire/life-safety, and BMS/EPMS controls.

    • Run the site maintenance program in CMMS — preventive, predictive, and corrective — with vendor and OEM service contracts managed to scope, and zero tolerance for overdue critical PMs.

    • Enforce operating discipline: SOPs/MOPs/EOPs current and rehearsed, change management on every energized system, drill program run on schedule, and monthly self-assessments presented to fleet leadership.

    • Lead site incident response for facility events end to end — detection, stabilization, escalation, root cause, corrective action — including communication when customer workloads are at risk.

    • On partner sites, manage the colocation operator to contracted SLAs: audit their maintenance records, witness their critical work, and enforce remedies when performance slips.

    • Support site bring-up: participate in commissioning through Level 5/IST, define and enforce ops acceptance criteria, tune the cooling plant to design intent in the first year, and manage warranty claims.

    • Own the site facility budget, utility coordination, and efficiency performance (PUE, WUE, energy cost), and run a safety-first culture (LOTO, energized work controls, working at height) with zero-compromise standards.

    SIGNALS OF SUCCESS

    After 60 days in this role:

    • Site staffing plan and 24/7 coverage model will be approved with hiring in motion.

    • Ops acceptance criteria will be agreed upon with construction/commissioning.

    After 6 months:

    • CMMS is loaded with the site asset registry and PM schedules.

    • Procedures library is drafted for every critical system (or partner program audited, on a colo site).

    After 1 year:

    • First site is operating with 24/7 coverage under the full standards library.

    • Preventive maintenance program is live in CMMS with zero overdue critical PMs.

    • Incident and change management program running.

    • Colo partner SLA and audit program in force at partner sites.

    EXPERIENCES, ATTRIBUTES AND MINDSET THAT INDICATE A GOOD MATCH

    Experiences

    • 10+ years in critical facility or data center operations — including power generation, HVAC, or mission-critical military/industrial environments — with 5+ years managing technical teams and vendors operating, maintaining, and troubleshooting electrical, mechanical, and controls infrastructure in a 24/7 environment.

    • Comprehensive working knowledge of electrical topologies (N+1/2N, UPS, generators, switchgear), mechanical and liquid cooling systems, fire/life-safety, and building automation — plus the safety programs (LOTO, hazardous energy control) that govern work on them.

    • Site bring-up or major expansion experience: commissioning participation, acceptance testing, and the construction-to-operations handoff on a mission-critical facility.

    Attributes

    • Owner of the building. When the plant alarms at 2am it is yours until the site is stable and the root cause is written down. You are comfortable being the most senior Crux person on site and making the call under pressure.

    • Discipline without bureaucracy. You run tight MOPs, drills, and change control because they prevent outages, and you kill process that does not — and you can tell the difference.

    Mindset

    • High tolerance for ambiguity. The site will energize before every procedure exists. You write the SOP, run it, fix it, and hand it to the next site rather than waiting for a standard to arrive.

    • AI-agentic first. Fluent with AI agents — or committed to becoming so quickly — and you embed them as first principles in how you and your team work, defaulting to agentic workflows before adding headcount or process.

    Nice to have (preferred, not required)

    • 7+ years in Tier III/IV or hyperscale data center environments (Meta, Google, AWS, Microsoft, or a major colo operator).

    • Direct-to-chip liquid cooling plant operations experience (CDUs, water chemistry, loop commissioning and tuning).

    • Trade certifications (Electrical, HVAC, Controls), stationary engineer license, or military/nuclear power background.

    • CMMS/EAM administration experience and comfort with energy and reliability analytics.

    • Experience managing a colocation operator to SLA from the tenant side.

    Salary Range Information

    The annual salary range for this position has been estimated based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

    About Crux

    • We offer generous base, bonus and additional incentive based compensation

    • Health, dental, and vision coverage for you and your dependents

    • Company-paid life insurance and disability

    • Full suite of other optional benefits

    • 401(k) Plan with 4% company match

    • Hybrid schedule offering four days in-office collaboration paired with one remote workday for focused, individual work

    Job details are sourced from the employer's original posting.

    Open job posting
    CR

    About the company

    Crux AI

    Built to set the gold standard for integrated AI infrastructure   Crux AI is a newly formed, U.S.-based integrated AI infrastructure company created to remove the physical and operational constraints on the world’s most consequential AI ambitions. Crux will finance, develop, and operate the full chain required to deliver large-scale TPU compute—from power and high-density data centers to silicon, networking, orchestration software, and ongoing operations—as one integrated system.   Those ambitions extend far beyond frontier models. They include faster drug discovery, earlier diagnoses, personalized treatment, mathematical breakthroughs, adaptive education, robotics for disaster response, energy optimization, national security, and data sovereignty. Their progress depends on reliable access to large blocks of accelerated compute—and on the power, facilities, networks, software, and technical expertise required to keep that compute running.   Crux is being capitalized and designed to address these constraints together. Its integrated model supports long-range capacity planning, efficient data center development, optimized TPU clusters, flexible service commitments, and dependable production operations. Crux will serve hyperscalers, frontier AI labs, sovereign and public-sector customers, global enterprises, and AI-native companies through compute-as-a-service and related infrastructure services.   This is the standard Crux is being built to set: every layer planned together, financed for the long term, and operated as one system—giving customers the freedom to focus on the AI they are here to create.

    View all Crux AI jobs
    Open roles
    29

    Interested in this role?

    Apply with HireFT

    Free to start — no card required.

    Your fit

    How well do you match?

    Sign in to see how your résumé lines up with this role.