HireFT
Browse JobsHow it worksPricingAboutSuccess Stories
    Back to jobs
    DE

    Deep Infra Inc.

    Artificial Intelligence

    Forward Deployed Engineer

    Palo Alto, United StatesOn-SiteFull-time$150k – $195k / yearPosted 1mo ago
    View all jobs

    Job description

    About DeepInfra

    DeepInfra is building the foundation for companies to use modern AI in production — simply, reliably, and at scale. Our team has deep experience building large systems that serve hundreds of millions of users, and we're bringing that same level of rigor to a rapidly evolving AI inference space. Our mission is to make advanced AI available to people and teams everywhere.
    We're an early, tight-knit team where you can influence product direction, try bold ideas, and drive meaningful work forward quickly. If you want to join a fast-growing company at a defining moment, we'd love to talk.
    DeepInfra is backed by leading investors including A.Capital, Felicis, 500 Global, Georges Harik, Samsung Next, Supermicro, Upper90, Peak6, SVAngel and Nvidia.

    Why this role matters

    As DeepInfra's enterprise pipeline grows, our customers need a technical partner who can run rigorous evals, defend benchmarks, and speak fluently to both engineering and procurement — someone who can own the technical win from first call through production.
    This is a pioneering role. You'll work closely with Sales, our co-founders, and the engineering team on the deals that matter most. You'll own the technical win end to end: running head-to-head bake-offs against leading AI providers, tuning deployments on the latest hardware, and turning what you learn into reusable assets that make every future deal faster to close. As our first FDE, you'll also define what the function looks like as GTM scales.

    What You'll Do

    • Own the technical win and the POC timeline, working closely with Sales and Engineering, from call one.
    • Design and run reproducible benchmark harnesses (TTFT, ITL, throughput/GPU, p95/p99) and quality-parity evals.
    • Run head-to-head bake-offs against leading AI providers — and win them.
    • Tune model-to-hardware deployments on B200/B300/GB300 NVL72.
    • Build cost-per-token models and write migration plans.
    • Handle enterprise security and compliance review, and get deployments to launch readiness.
    • Own account health post-signature, driving usage reviews and expansion.
    • Turn what you learn into reusable benchmark reports, reference architectures, and AE enablement material.

    What You Bring

    • Customer-facing engineering with an owned technical outcome at an infrastructure or ML platform company.
    • Strong Python skills.
    • Dual-audience presence with commercial instinct — credible with a skeptical staff engineer, clear with a CFO, and able to tell a technical objection from a procurement one.

    Bonus

    • Hands-on experience with inference internals: vLLM, SGLang, or TRT-LLM, batching, KV cache math, quantization.
    • Experience with agentic or coding-assistant workloads at scale.
    • Prefix-cache-heavy long context workloads.
    • Diffusion image/video, ASR/TTS, or multi-LoRA serving.
    • Open-source contributions to vLLM or SGLang.
    • Deep NVLink/InfiniBand topology knowledge.

    Why DeepInfra

    • Define DeepInfra's Forward Deployed Engineering function from day one and have a direct impact on its direction.
    • Work directly with co-founders and the inference team on the deals that matter most.
    • Join a small, high-performing team where your work ships quickly and reaches customers around the world.
    • Help shape how enterprises adopt some of the world's leading open-source AI models.

    How we work

    Three traits define the people who thrive here, and this role leans on all three.
    Initiative. We take ownership and step in where we can add value. Whether it’s starting something new, improving what exists, or helping move ideas forward, we aim to be proactive and thoughtful in how we contribute.
    Drive. We’re energized by hard problems. Building AI infrastructure is complex, and we lean into that. We care about doing things well, moving fast, and continuously improving — because solving meaningful challenges is what motivates us.
    Grit. Things don’t always work on the first try — and that’s expected. We stay persistent, adapt quickly, and learn as we go. We take setbacks seriously, but not personally, and use them to get better.

    Compensation

    The base pay range for this role is $150,000 – $195,000 per year.

    Job details are sourced from the employer's original posting.

    Open job posting
    DE

    About the company

    Deep Infra Inc.

    DeepInfra is a company focused on designing, developing, and deploying top open AI models at scale. They offer opportunities for software engineering interns to gain hands-on experience in building scalable and efficient software systems with cutting-edge AI.

    View all Deep Infra Inc. jobs
    Industry
    Artificial Intelligence
    Open roles
    17

    Interested in this role?

    Apply with HireFT

    Free to start — no card required.

    Your fit

    How well do you match?

    Sign in to see how your résumé lines up with this role.