HireFT
Browse JobsHow it worksPricingAboutSuccess Stories
    Back to jobs
    SA

    Sail Research

    Member of Technical Staff - Inference

    San Francisco, United StatesOn-SiteFull-time$200k – $300k / yearPosted 1w ago
    All Sail Research jobs

    Job description

    Sail builds the world's most efficient software for inference (processing LLM tokens) and agent hosting (cloud VMs). Together, our technologies allow our customers to deploy AI agents at large scale to do the most challenging work.

    In this role, you'll own token processing down to the lowest layers of the stack. You'll do things like: develop a new request scheduling strategy, achieve better communication/computation overlap, investigate novel schemes for increasing cache hit rates, or identify a better way to benchmark inference performance.

    What you’ll do

    • Modify and extend state-of-the-art inference engines like vLLM and SGLang, and work on our own internal engine.

    • Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain every kernel launch on an nsys profile.

    • Design and implement exotic parallelism schemes to work with "interesting" hardware topologies.

    • Write and debug GPU kernels to excel in specific regimes, such as cascade attention

    What we’re looking for

    • Strong understanding of core LLM mechanics, like KV cache, mixture-of-experts, prefill vs. decode phases.

    • Interest in MLSys research - great ideas like speculative decoding and sparse attention come from research, that we need to follow closely.

    • Familiarity with modern, tile-based GPU programming, e.g. Triton, CUTLASS, ThunderKittens, etc. Or an interest in learning these!

    • Great interpersonal and technical communication. Please don't use LLMs to write prose. We desk-reject slopful cover letters and resumes.

    Interview process

    1. Meet the CTO, who will ask about your experience, and share as much technical detail about Sail as you want to hear. This is the first step because we respect your time.

    2. Share an online whiteboard with a team member and work through a technical problem. We spend a lot of time at whiteboards, building intuition about complex systems together. It's a great way for us to see how you communicate technically, and a even better way for you to see what working at Sail is like.

    3. Come in to Sail's SF office for an interview day. Meet the whole team, and work on a bunch of problems that closely simulates the work we do daily. We'll also ask you to give us a 20-30min 'chalk talk' about an interesting problem you've worked on before.

    4. Offer. Once the team decides we want to work with you, we make a strong offer quickly and will be quite persistent over email/text/calls :)

    Life at Sail

    We work out of a beautiful, sunny office in downtown San Francisco. All meals are on us (and actually great; SF is a food paradise!). Everyone gets a Studio Display (or two) at their desk. We are serious about investing in anything that saves us time or energy. There are six different ways to make coffee or tea in the office. A friendly (hypoallergenic) black cat named Coco visits occasionally.

    Job details are sourced from the employer's original posting.

    Open job posting
    SA

    About the company

    Sail Research

    View all Sail Research jobs
    Open roles
    5

    Interested in this role?

    Apply with HireFT

    Free to start — no card required.

    Your fit

    How well do you match?

    Sign in to see how your résumé lines up with this role.