Developing Real-Time Voice AI for a Multilingual World

@stevejang
ENGLISH3 weeks ago · Jun 30, 2026
1.3M
108
24
5
39

TL;DR

Kindred Ventures led a $10M seed round for Kotoba, a research lab developing real-time voice AI and translation models purpose-built for Japanese, Korean, and Chinese markets.

For many of us in Silicon Valley and similar global hubs, we are well aware that voice AI is fast becoming the new core modality of how people work, communicate, and interact with agents and each other. The shift becomes strikingly apparent as agent systems increasingly move beyond coding into new sectors of knowledge work like Perplexity Computer and Claude Cowork, consumer-facing applications like Wispr Flow, Sierra, and Granola, and into agent embodiments in myriad cars, robots, and wearables. And yet outside of our regional chambers, many of the world's most important languages have been treated as an afterthought and little progress has been made on the interconnection of these languages and their speakers.

By current count, Asia is now home to nearly 5 billion people. East Asia alone represents 1.6B – 20% of the global population. Roughly half of the world's knowledge workers speak an Asian language. A new set of speech AI models, trained specifically for Asian languages, will enable us to truly achieve multimodal intelligence within reach of this global majority.

With hundreds of distinct languages, each carrying its own linguistic nuances and data characteristics, building for East Asia requires far more than building off of an English-first model: Building the future of a global-first knowledge work demands a ground-up approach to model training and market expertise.

Taking a step back, we’ve all been watching as much of frontier research work in Asia centering in China, particularly in open-weights large language models and generative media. In the past year in Japan and Korea, we are now seeing a new wave of research labs emerging. These research teams focused on not only variations of homegrown large language models like Upstage and Sakana, but also new labs developing multimodality with speech models and video understanding, and on physical AI with robotic intelligence and world models.

Today, we’re excited to announce that @KindredVentures led a $10 million seed round in Kotoba (@kotoba_tech), alongside Salesforce @SalesforceVC and Sony Ventures (@Sony_Innov_Fund). In our very first conversations with the founders about training data and model architecture, we were super impressed by their highest quality ASR and TTS models which are perfect for various agent pipelines, but also their research progress on smaller edge models for on-device inference, and their frontier speech-to-speech realtime translation models which outperform translation models from Google, Microsoft, and OpenAI.

Founded by @noriyuki_kojima (PhD, @Cornell and @jungokasai (PhD, @UW), @kotoba_tech is building speech AI for East Asian languages. In their prior work, they were the co-founders of an early Japanese government and university research project called the LLM-Fugaku project — Japan’s large‑scale language model initiative built on the Fugaku CPU-only supercomputer. They were able to train a Japanese LLM successfully using a transformer architecture without any GPUs, only CPUs. Today at Kotoba, the Koto proprietary model family delivers industry-leading performance across Japanese, Korean, and Chinese, powering AI voice agents, devices, wearables, robotics, and real-time speech translation and reasoning with the accuracy and latency these markets demand.

What continues to stand out about this team was the rare combination of world-class research, deep cultural fluency across East Asia, and a product already demonstrating meaningful traction. Kotoba's models aren't adaptations of English-first systems—they're purpose-built for the linguistic realities of the markets they serve with a unique training approach. Just 6 months after release of their first model, their models consistently perform at lower latencies and higher quality on prosody than other models from Western companies. In the first six months are releasing their models privately to customers, Kotoba now counts several Fortune 100 enterprises, global hardware companies, and high-growth AI-native startups as their initial customers.

We're thrilled to partner with @noriyuki_kojima, @jungokasai , and the entire @kotoba_tech team as they build a new frontier research lab for Japan and a Voice AI platform for broader Asia and RoW.

You can read more about our investment below:

https://kindredventures.com/announcement/kotoba-developing-voice-ai-for-a-multilingual-world/

Remix in YouMind

Turn one viral article into a full content workflow

Collect the source, decode the pattern, create assets, draft the story, and distribute from one AI workspace.

Explore YouMind
For creators

Turn your Markdown into a clean 𝕏 article

When you publish your own long-form writing, images, tables, and code blocks make 𝕏 formatting painful. YouMind turns a full Markdown draft into a clean, ready-to-post 𝕏 article.

Try Markdown to 𝕏

More patterns to decode

Recent viral articles

Explore more viral articles