HireFT
Browse JobsHow it worksPricingAboutSuccess Stories
    Back to jobs
    CR

    Crux AI

    Senior Site Reliability & Software Engineering Manager

    Palo Alto, United StatesHybridFull-time10+ yrs experiencePosted 1w ago
    All Crux AI jobs

    Job description

    Built to set the gold standard for integrated AI infrastructure

    Crux AI is a newly formed, U.S.-based integrated AI infrastructure company created to remove the physical and operational constraints on consequential AI ambitions. Crux brings together power, high-density data centers, TPU silicon, networking, orchestration software, and ongoing operations as one integrated system.

    Crux is being capitalized to plan every layer together, develop each one to demanding standards, and operate the whole system with efficiency and reliability. That gives hyperscalers, frontier AI labs, sovereign customers, enterprises, and AI-native companies greater freedom to pursue the AI they are here to create.

    Crux AI is led by CEO, Ben Treynor Sloss, who spent over two decades in executive technical leadership at Google and founded the Site Reliability Engineering (SRE) discipline. At Crux AI, we treat operations fundamentally as a software engineering problem.

    WHAT YOU'LL DO

    We are recruiting founding Senior Site Reliability and Software Engineering Managers to build and lead our initial fleet reliability engineering teams in Palo Alto, CA.

    In this organization, there is no separate software development team. Your team owns the software, control plane, telemetry, and automated remediation controllers that keep multi-gigawatt TPU clusters provisioned, resilient, and continuously executing customer AI workloads.

    This is a true hands-on, builder seat, not a supervisory position. In the early days, you will write the first remediation controllers, set reliability baselines, and take initial pages yourself to stay close to the work before expanding your team. You will lead an elite group of unusually senior software and reliability engineers: engineers with significantly greater technical and software depth than traditional operational SRE orgs. You must thrive on independence and revel in ambiguity, turning unknowns into concrete engineering priorities in a fast-paced, high-growth environment. While you will devise and participate in initial on-call rotations, your core mandate is to combine SRE disciplines with extensive AI/ML automation to drive operational pages down to zero.

    In this role, you will:

    • Build & Lead Senior Engineering Teams: Recruit, lead, and mentor an initial team of senior software and reliability engineers across Palo Alto and Europe as fleet capacity ramps rapidly.

    • Automate Pages to Zero: Own fleet availability end-to-end; establish on-call rotations while relentlessly developing self-healing systems and predictive remediation to eliminate manual pages.

    • Embed AI/ML into SRE Disciplines: Apply agentic techniques, machine learning models, and automated diagnostic workflows extensively to telemetry collection, root-cause analysis, and predictive cluster recovery.

    • Own Bare-Metal & Fleet Lifecycle Software: Drive software engineering for bare-metal node provisioning, firmware deployment, thermal/stress burn-in validation, host/TPU health monitoring, and decommissioning.

    • Control Plane & Fabric Reliability: Own software reliability for cluster orchestration, scheduling, capacity allocation APIs, and high-performance TPU host/interconnect networks.

    • Define Observability & SLOs: Establish customer-facing SLIs/SLOs (job goodput, time-to-detect, node availability) and build the telemetry pipelines serving operators, executives, and customers.

    SIGNALS OF SUCCESS

    After 60 days in this role:

    • Baseline reliability framework defined (v1 SLOs, severity structure, change management); first AI/ML-driven automated remediation controller shipped to production; recruiting active for senior engineering hires in Palo Alto.

    After 6 months:

    • First TPU cluster brought online under your team’s automated acceptance criteria; automated remediation pipeline running with pass rates tracked; observability v1 in daily production use; core senior team onboarded

    After 1 year:

    • TPU fleet operating against published customer SLOs; >90% of node/fabric faults automatically quarantined and remediated without human paging; team scaled ahead of rapid capacity ramps.

    EXPERIENCES, ATTRIBUTES AND MINDSET THAT INDICATE A GOOD MATCH

    Experiences

    • 10+ years of software or infrastructure engineering experience, with 3+ years managing engineering teams owning direct production SLAs and on-call.

    • Deep SRE Discipline: Grounded in foundational SRE principles (SLOs, error budgets, blameless postmortems) paired with a strict "code over heroics" mindset.

    • Hands-On Technical Depth (SRE + SWE): Track record shipping production code in Go, Python, or C++, with hands-on systems expertise across Linux OS kernels, bare-metal provisioning, firmware, and/or high-performance networking fabrics. Extensive experience with distributed systems and open source software.

    • Extensive AI/ML Adoption: Active utilization of AI agents and automated LLM/ML workflows in modern software engineering and diagnostic operations.

    • Senior Talent Magnet: Track record of attracting, evaluating, developing and leading unusually senior software engineers who thrive in fast-paced, high-stakes environments.

    Attributes

    • Possess a high tolerance for ambiguity. The first clusters will carry customer workloads while the SLOs are still being defined and the team is still being hired. You absorb that, translate unknowns into concrete near-term priorities, and never manufacture false certainty about reliability the data does not support.

    • Understands that the customer’s job is the unit of reliability. A node that is “up” while a training run stalls on a flapping link is down. You measure what customers experience — goodput, time-to-recover, lost progress — and hold the whole stack, and Google, to it.

    Mindset

    • Builder Mindset & Ambiguity: A true "builder, not supervisory" orientation; comfortable operating with high autonomy, navigating ambiguity, and establishing structure amidst rapid growth.

    • AI-Agentic First: Fluent with AI agents — or committed to becoming so quickly — and you embed them as first principles in how you and your team work, defaulting to agentic workflows before adding headcount or process.

    Nice to have (Preferred, not required):

    • Hyperscaler / Neocloud Scale: SRE or fleet leadership at a hyperscaler (Google, AWS, Meta, MSFT) or neocloud (CoreWeave, Lambda, Nebius, Nscale) during rapid fleet ramps.

    • Accelerated Compute: Direct TPU experience or large-scale GPU cluster ops (NCCL collective debugging, RDMA/GPU-Direct, Slurm/Kubernetes AI schedulers).

    • Custom Fleet Tooling: Hands-on experience building custom remediation controllers, event-driven fleet management software (Go, NetBox/DCIM), or OpenTelemetry/Prometheus pipelines.

    • Facility Telemetry & Thermal Signals: Familiarity with high-density, liquid-cooled environments and integrating facility telemetry (power, thermal, flow) into compute health signals.

    • Customer SLAs & Reporting: Proven experience constructing customer SLAs/SLOs, credit mechanics, and executive/customer-facing reliability reviews.

    Salary Range Information

    The annual salary range for this position has been estimated based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

    About Crux

    • We offer generous base, bonus and additional incentive based compensation

    • Health, dental, and vision coverage for you and your dependents

    • Company-paid life insurance and disability

    • Full suite of other optional benefits

    • 401(k) Plan with 4% company match

    • Hybrid schedule offering four days in-office collaboration paired with one remote workday for focused, individual work

    Job details are sourced from the employer's original posting.

    Open job posting
    CR

    About the company

    Crux AI

    Built to set the gold standard for integrated AI infrastructure   Crux AI is a newly formed, U.S.-based integrated AI infrastructure company created to remove the physical and operational constraints on the world’s most consequential AI ambitions. Crux will finance, develop, and operate the full chain required to deliver large-scale TPU compute—from power and high-density data centers to silicon, networking, orchestration software, and ongoing operations—as one integrated system.   Those ambitions extend far beyond frontier models. They include faster drug discovery, earlier diagnoses, personalized treatment, mathematical breakthroughs, adaptive education, robotics for disaster response, energy optimization, national security, and data sovereignty. Their progress depends on reliable access to large blocks of accelerated compute—and on the power, facilities, networks, software, and technical expertise required to keep that compute running.   Crux is being capitalized and designed to address these constraints together. Its integrated model supports long-range capacity planning, efficient data center development, optimized TPU clusters, flexible service commitments, and dependable production operations. Crux will serve hyperscalers, frontier AI labs, sovereign and public-sector customers, global enterprises, and AI-native companies through compute-as-a-service and related infrastructure services.   This is the standard Crux is being built to set: every layer planned together, financed for the long term, and operated as one system—giving customers the freedom to focus on the AI they are here to create.

    View all Crux AI jobs
    Open roles
    29

    Interested in this role?

    Apply with HireFT

    Free to start — no card required.

    Your fit

    How well do you match?

    Sign in to see how your résumé lines up with this role.