We stayed quiet and built the FASTEST search on Earth

@KZouAPT
अंग्रेज़ी5 घंटे पहले · 21 जुल॰ 2026
399K
351
82
146
138

TL;DR

Octen introduces an AI-native search API designed for high-concurrency agents, featuring 62ms latency and disruptive pricing at $1 per 1,000 calls. Built by search veterans, it aims to replace human-centric search stacks.

Search was built for humans. You type one query, scan the results, refine, and try again.

Thirty years of infrastructure is designed around that loop.

AI agents DON'T work that way. An agent wants to ask a question from many angles at once and get everything back together.

But the search APIs feeding agents today are the same human loop wrapped in an endpoint.

One query in, one response out, hundreds of milliseconds each.

We spent the last year rebuilding search for the thing actually doing the searching. We didn't post much about it.

Honestly, we weren't ready. Earlier this year I felt we were at maybe 70% of our own expectations, and I didn't want to make noise about 70%.

We're ready now.

The numbers

Octen's web search API returns results in 62ms (P50) with a P90 of 68ms, measured on SealQA Hard.

That is several times faster than Exa, Parallel, and similar services.

The fastest alternative measures around 244ms. The slowest, above 2.6 seconds.

Kuan Zou - inline image

One detail in this chart matters more than the headline: the gap between our P50 and P90 is 6ms. For some services it is over one second.

If your agent cannot predict how long a search takes, it cannot plan around it.

Pricing is $1 per 1,000 calls. Most services in this category charge $4 to $9.

On quality: Octen Embedding is #1 on RTEB, and our VL embedding model is #1 on MMEB-v2.

We open-sourced text embedding models and they have been downloaded more than 500,000 times in five months.

On DeepResearch Bench we score above OpenAI Deep Research, Gemini, and Grok by 10 to 17 points.

On FreshQA Strict and SimpleQA we lead every vendor we test against.

Kuan Zou - inline image

Note: All benchmarks are published and reproducible. Links at the end.

Why I started this

My name is Kuan Zou. Before Octen, I was the product lead of AI search at Alibaba Cloud for five years, running systems that served hundreds of millions of users.

Before that, I built Baidu’s enterprise search platform from the ground up.

Every year at Alibaba, my job was to survive Black Friday. Billions of queries in a day, and no tolerance for failure.

So when I looked at the search APIs being given to AI agents, I was genuinely surprised. Most of them start to break when queries per second reach the tens.

An agent doing real work will trigger 20, 50, 100 searches together. The infrastructure agents depend on most is the part that fails first.

We raised a $10M seed led by Square Peg, brought together a team from Alibaba, Baidu, Meta, Google, TikTok, DeepSeek, and Xiaohongshu, and got to work.

From linear to concurrent

Give Octen a question and it decomposes it into dozens of sub-queries, fires them all simultaneously, and assembles everything that comes back into one answer.

Not one query at a time. All of them at once.

Kuan Zou - inline image
Kuan Zou - inline image
Kuan Zou - inline image

This is what it does to deep research: a full source-backed report in under 3 minutes. Systems that take over an hour on the same task score below us on DeepResearch Bench.

Kuan Zou - inline image

At 62ms, search stops behaving like an external tool and starts behaving like memory.

Your agent reasons over the live web at close to the speed it reasons over its own context.

That opens applications that were not practical before. Real-time trading agents. Live sports and market data. Voice agents that answer without the awkward pause.

Kuan Zou - inline image

Built for reliability

A single Octen account supports more than 1,000,000 queries per second. New content is indexed within 5 minutes of publication.

We also offer guaranteed QPS with an SLA.

As far as I know, we are the only ones in this category who can promise this, because it is only possible when you own the index end to end. We do.

Kuan Zou - inline image

Beyond text, we run image search and video search, and we ground image and video generation with real references from the live web.

This layer is in early access now.

Kuan Zou - inline image

How it is also 5x cheaper

There is no trick.

Most "search for AI" products are built on open-source search stacks that were designed for humans. They took the engine behind a query box and moved it behind an API.

The technology did not change. Only the interface did. Calling that "search for AI" makes no sense.

We rebuilt everything. The search engine, the ranking, the serving layer, all written from scratch for machine consumption. Nothing in our stack was designed for a person scrolling results.

That is why the price is possible. A fully AI-native engine doesn't inherit thirty years of human-search overhead. The business is healthy at $1 per 1,000 calls. This is not a subsidy waiting to expire.

Price is the most honest proof of architecture. Anyone can claim their search is built for AI. The cost structure shows whether it is true.

Some companies in this category already run our API against theirs every day.

I take it as a compliment.

Why I'm sharing everything

Knowing what we do and being able to build it are different problems.

Our moat is a team that has spent a decade handling some of the hardest search traffic in the world.

The first movers educated the market, and I'm grateful for that.

But developers always check the numbers, and they route to whatever performs best and costs least.

I think that is the most honest market in the world.

The incumbents spent the last two years on marketing. We spent them on engineering.

Now the engineering is public.

The physical era

The next generation of AI does not live on a screen.

Glasses, robotics, world models. In that world, search becomes the eyes: real-time perception of the live web.

A robot cannot wait 3 seconds for an answer.

When search returns in tens of milliseconds, real-time interaction for physical AI finally breaks through the bottleneck that has held it back.

In that era, I don't believe anything above 100ms of latency survives. Today we are the only ones under that bar.

The agent stack is settling into three pieces: foundation models, GPUs, and live search.

The first two are solved problems with clear winners. The third is still being decided. That is the race we intend to win.

Try it yourself

The benchmarks are public and the API is open.

$1 per 1,000 calls, with $5 in free credit. Available as an API, as Skills, over MCP, and from the CLI. It works in Claude Code, Cursor, VS Code, and any MCP client.

Run us against whatever you are using today.

Same queries, side by side. Post what you find.

Companies in this category already benchmark us daily, and you should too.

Docs: docs.octen.ai

Note: Latency and accuracy figures measured on SealQA Hard, FreshQA Strict, and SimpleQA. Test date and region are stamped on each chart. Reproduce the benchmark: https://github.com/Octen-Team/search-eval

YouMind में रीमिक्स करें

Turn one viral article into a full content workflow

Collect the source, decode the pattern, create assets, draft the story, and distribute from one AI workspace.

Explore YouMind
क्रिएटर्स के लिए

अपने Markdown को एक साफ़-सुथरे 𝕏 आर्टिकल में बदलें

जब आप अपना लंबा कंटेंट पब्लिश करते हैं, तो इमेज, टेबल और कोड ब्लॉक को 𝕏 के लिए फ़ॉर्मेट करना मुश्किल होता है। YouMind पूरे Markdown ड्राफ़्ट को एक साफ़-सुथरे, पोस्ट के लिए तैयार 𝕏 आर्टिकल में बदल देता है।

Markdown से 𝕏 आज़माएँ

समझने के लिए और पैटर्न

हाल के वायरल लेख

और वायरल लेख देखें