HireFT
Browse JobsHow it worksPricingAboutSuccess Stories
    Back to jobs
    LU

    Luminal

    B2B Software and Services

    Cloud Inference Engineer

    United StatesOn-SiteFull-time$150k – $250k / yearPosted 1mo ago
    View all jobs

    Job description

    # Qualifications * CUDA + GPU inference optimization * vLLM, SGLang, or TensorRT-LLM experience * KV caching, paged attention, batching, token streaming, etc. * Distributed compute (with GPUs is a super plus) * No degree required # Company Luminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line. # Role Founding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud. Day to day responsibilities: * Deploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc. * Conducting model performance reviews * Improve scheduler, batcher, autoscaling; profile latency, cost, utilization * Sometimes write kernels and, yes, occasional tasteful shitposting

    Job details are sourced from the employer's original posting.

    Open job posting
    LU

    About the company

    Luminal

    Making AI run fast on any hardware.

    View all Luminal jobs
    Industry
    B2B Software and Services
    Open roles
    3

    Interested in this role?

    Apply with HireFT

    Free to start — no card required.

    Your fit

    How well do you match?

    Sign in to see how your résumé lines up with this role.