HireFT
Browse JobsHow it worksPricingAboutSuccess Stories
    Back to jobs
    OL

    Ollama

    Artificial Intelligence

    Software Engineer, Cloud

    Palo Alto, United StatesOn-SiteFull-timePosted 1mo ago
    View all jobs

    Job description

    Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.

    Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.

    About the role

    You'll build Ollama’s cloud, a scalable inference platform that lets developers run large, capable open models in their workflow. You'll work on high-throughput, low-latency distributed systems — inference serving, GPU fleet management, routing, metering, and the platform that Pro, Max, Team, and Enterprise customers rely on to process trillions of tokens.

    What you'll do

    • Build and scale the inference platform that serves every request from ollama.com.

    • Design the routing and capacity layer that places workloads across GPUs and regions for cost, latency, and availability.

    • Own multi-tenant infrastructure: isolation, quotas, usage metering, billing, and Pro/Max/team/enterprise tiering.

    • Build the reliability, observability, and cost controls for our team and customers

    You may be a fit if

    • You have deep experience with high-throughput, low-latency distributed systems — inference serving, traffic routing, real-time data pipelines, or large-scale APIs.

    • You're comfortable with cost/performance tradeoffs at scale and have owned a production service end-to-end.

    • You've worked with Kubernetes, GPU scheduling, or inference infrastructure.

    • You think in terms of reliability, SLOs, and honest capacity planning.

    • Bonus: experience building an inference platform, GPU fleet management, or billing/metering for an AI service.

    Job details are sourced from the employer's original posting.

    Open job posting
    OL

    About the company

    Ollama

    Ollama develops an open-source tool for running large language models locally on personal computers.

    View all Ollama jobs
    Industry
    Artificial Intelligence
    Open roles
    8

    Interested in this role?

    Apply with HireFT

    Free to start — no card required.

    Your fit

    How well do you match?

    Sign in to see how your résumé lines up with this role.