AI assistants live or die by speed. An assistant has to get things done for you, but on iMessage it has very few ways to show you that it is working on your request: a tapback, a typing bubble, a short reply.
Partyline is a special kind of AI assistant in your iMessage. We're not focused on triaging your email: our opinion is that as a human you shouldn't have been on email in the first place. Our goal is to get you outside, with other people, having a good time. Our product metric is the amount of life our users have that's worth living.
For an assistant however, it means competing in a very, very, busy world that is doing it's best to tire out our users. Make them depressed and unhappy. But most of all, compete for their attention.
Because people text from chaotic environments we have to be fast. We only get a moment with our user before they put their phone away (as they should!) and we don't want them looking at the screen for a moment longer than they have to.
When we built Partyline, our replies were too slow. Finding a place or an event took 29 to 46 seconds. That is a long time when you are out with friends. If you are the one who pulls out your phone to find the next spot, you are putting your social capital on the line: you are the person who knows something cool is going on. At 46 seconds, someone else has already suggested a place and taken over the conversation.
So an iMessage assistant has to be very fast. An app gives you a screen you control while someone uses it. iMessage doesn't: people use it in busy, complicated moments, and the assistant competes with everything around them.
How we compare
We compared Partyline with two other personal assistants on iMessage, Tomo and Instinct. Instinct was the slowest: its answers to searches took 42 to 82 seconds, which is too slow to use while you are out with friends. Tomo is excellent at conversation. Its chat replies arrive within a few seconds, and its search answers take 28 to 37 seconds, with a tapback at about 2 seconds and a first text at about 5.

From 46 seconds to 3
Tomo set the baseline question for us: how could Partyline reply as quickly as Tomo does in conversation, but for bigger tasks like finding places and events? We started at 29 to 46 seconds. Today a search takes 2 to 3 seconds from send to reply.

Two changes made most of the difference.
First, we index events in advance for the places people ask about most. During a16z's Tech Week in San Francisco our inventory holds about 1,700 events for the week, and a search narrows them to a short list in under a tenth of a second. Where an area isn't indexed, a web search runs alongside and answers a few seconds later.
Second, we put Jev classifiers at several steps of our message pipeline. Jev is TypeSafe's classifier model. In one call of about 0.2 seconds it answers many questions about a message: what kind of request it is, whether it continues the last search, whether it is sensitive. Jev also ranks the candidate events and places, and checks each line of a reply against the data it describes. When a large model made those decisions, each one took about 4 seconds.

What's next
Partyline now finishes a search answer before Tomo sends even its first text, and it can hold many rounds of casual conversation before Instinct replies once. The downside is that some answers read less like a conversation and more like a fast response. We're working on it. Our next job is to tune the system toward a middle ground that still feels near-instant but reads like a conversation.
We're also building tooling to keep inference costs affordable as we grow, so we don't have to trade speed or quality for cost.
Today, Partyline is the fastest AI iMessage assistant in the world, especially for event search. If you are at a16z's Tech Week in San Francisco this week, text Partyline for the fastest event search you have seen.





