Most of our users never really touch our app. They text us. Restaurant owners and creators handle almost everything through conversation, and on our side that's a set of AI agents doing the booking, the follow-ups, the scheduling, and the problem solving. You own those agents. Not the prompts alone, the whole system underneath them: how they understand what someone actually wants, how they call tools without breaking, how they hold context across a conversation that's been going for 6 months, and how they take real actions in the real world without us having to check their work.
Any text that comes in, your systems handle it. ### **What you'll own** **The agents.** Full ownership of both of our agents end to end. How they're built, what they can do, and what happens when they get something wrong. **The harness.** The layer coordinating LLMs, tools, memory, and async workflows. This is the actual engineering problem here and it's most of your week. **Tooling** that doesn't fall over. Schema validation, retries, permissions, error handling.
An agent that calls a tool correctly 95% of the time is not good enough when it's booking real visits at real restaurants. **Evals**. Frameworks that measure task completion, accuracy, safety, latency, and failure modes. If we change a prompt on Tuesday, we should know by Tuesday whether it made things worse. **Observability.** Traces, logs, and failure analysis across agent workflows. When something goes wrong in a conversation, you should be able to see exactly where.
Turning vague into dependable. A restaurant owner texts something ambiguous at 11pm. Getting from that to a reliable agent behavior is the hard part, and it's the part we care most about. ### **What we're looking for** Required * 2+ years building software, with real experience building agent systems or harnesses. Not just calling an API in a side project * Strong conversational agent experience. This is the thing we weigh most heavily.
Our product is a conversation, and someone who has only built single-turn or task-runner agents will struggle here * Strong database and system design * You've shipped agents that real people used and dealt with the fallout when they broke * Comfortable with ambiguity. There's no established playbook for most of this * Based in Orange County / willing to relocate and able to work onsite * This isn't a 9-to-5 Nice to have * iMessage agent experience * You've built eval frameworks, not just run them * Observability and tracing work on LLM systems * Restaurant industry or creator economy experience
Job details are sourced from the employer's original posting.
Open job posting