Safety vs utility - how startups win

351K
118
6
10
27

TL;DR

Sivesh Sukumar analyzes how startups like xAI and Anthropic use the safety-utility trade-off to disrupt incumbents, building data flywheels by shipping more aggressive AI agents.

There has always been a tradeoff between safety and utility in AI. AI is “probabilistic software”, there will always be unknown edge cases and these can be scary - especially for incumbents with a lot to lose (brand risk, fines, churn etc…). There’s a reason why research teams are split into “capabilities” and “safety”, they are often seen to be orthogonal areas of research.

This tradeoff has consistently kept the door open for startups to displace “incumbents” and win over consumers. Google was so well placed to launch something like ChatGPT - they had the distribution, research, compute and even the idea. They didn't launch due to the risk of hallucination, brand risk and the negative press around LaMDA - given the unclear commercial opportunity, they didn't take this risk.

OpenAI were less worried about this, as a startup, so were happy to launch. This launch gave a wedge into a data flywheel enabling the capture of market share in one of the strongest monopolies we’ve ever seen (consumer search) and drove the success of one of the most valuable private companies in the world.

There are many of these case studies so I wrote a post to explore how startups can navigate the trade-off to win in consumer AI / the explosive personal agent market forming.

Sivesh Sukumar - inline image

Grok Bot

The recent inflection of personal agents e.g. Grok Bot, “managed open claw” services (or what I prefer to call Proxy / Convergence fast followers, iykyk) is a great case study for this. Anthropic and OpenAI both have similar “personal agent” products but having used Grok Bot for the past few days I've noticed it’s incredibly “aggressive” - there are plenty of examples of Grok doing things Claude / ChatGPT would never do automatically (e.g. resetting a password to login to a service).

Claude and ChatGPT’s respective computer-use agents are capable of doing these things but Anthropic and OpenAI both have a lot to lose now so there is more red-teaming / alignment to prevent the model from doing bad things, this limits how useful their agents can be. xAI / Cursor are happy to take that risk given a lack of market share in personal agents and it seems to be working!

Anthropic / Claude Code

Anthropic is possibly the best case study for a company who has won by navigating this trade-off not just to find a wedge but also finding the right trade-offs over time to expand market share. This would make sense given their roots as an “AI safety” company.

There’s no denying Anthropic has had an edge in coding capabilities over the last 12 months but Claude Code as a product launch felt kind of insane to most CISOs and risk officers. It was a powerful agent, prone to making mistakes, deployed locally with full access to your filesystem / OS - it could leak anything on your machine or even delete your entire computer. In reality, there aren’t many examples of it doing this and with appropriate use the product is pretty safe. Given the utility of the product, millions of developers (+ consumers) were happy to take this risk and this gave rise to a data flywheel for Anthropic - it’s safe to say Anthropic wouldn’t be where they are today if they didn’t take this risk with Claude Code.

What’s more interesting is how Anthropic evolved Claude Code to expand their universe of users / tap into the user-base that was too “scared” to use Claude Code. Cowork is essentially a wrapper of Claude Code but has guardrails in place to prevent bad things going wrong, if anyone has used both it’s quickly evident that Cowork is less powerful despite being built on the same models and harness.

Sivesh Sukumar - inline image

The recent launch of “auto mode” (Anthropic made it default for all users a couple weeks back) is another example - a good proportion of developers have been using “YOLO mode” i.e. “—dangerously-skip-permissions” to let agents run for hours or overnight solving a task without ever asking for permission. Some would think it's crazy to even launch this feature given the risks but I’m sure a huge amount of useful work has been done this way.

Anthropic built Auto Mode (a classifier to escalate potentially harmful actions during autonomous runs) to find a balance and enable long running agents whilst still having “some” guardrails for when the agent should ask for permission. This enables more confidence in long running agents but means a run can stall, inherently making the agent less useful. However, this trade-off makes sense as it expands the universe of people using long running agents.

Black Forest Labs / Mistral

Some less obvious examples of this trade-off can be found at the model layer as well - early BFL and Mistral models were able to top many AI leaderboards, beating out the biggest labs in the world - and they made a great brand for themselves in doing so. Whilst this might’ve been due to better research, it definitely wasn’t due to better data or compute.

My very cynical view is that they were able to do this by launching “less safe” models - the incumbents have to put a lot of effort into alignment and red teaming for obvious reasons, this is even more important for image and video gen. It was very clear that when BFL / Mistral were at the top of the leaderboards their models were less safe vs competitors and didn’t have the same red-teaming (but were no doubt more capable!).

Conclusion / Opportunities

“Ship an unsafe product” is not the takeaway of this post - there are plenty of examples of companies taking risks and it backfiring (e.g. Grok’s imagery mess). I’m encouraging startups to understand (and ideally quantify) the actual risks of an AI experience and make the leaps which incumbents will not.

It’s also important to understand that the first adopters of a “risky” product will likely be a small segment of your customer universe and may be hard to find - Harvey and Legora did a great job of finding the law firms who were willing to make the leap into AI first and built their business around these early partnerships. Most of their product differentiation vs Cowork comes from building “safe” features (e.g. governance, audit trails etc…). Whilst most of these lessons are more applicable to B2C I’m sure there’ll be some in B2B land as well.

I can keep going with countless examples of this trade-off and we are going to see plenty more given the fundamental nature of AI being “probabilistic software”. I’m now thinking about the future opportunities where startups can take advantage of this trade-off - financial services / robotics seem great.

AI only became capable enough to make financial services products viable at the start of this year. Financial services incumbents are some of the most risk averse out there, most of their business models are built on trust so the opportunity is very exciting. A good example is in AI agents for wealth management / investing - there’s an obvious opportunity here but the best example I can see is from Public.com and it’s sub-par. All offerings are very narrow, with many rules to be defined and they’ve restricted what the agent is actually capable of doing - they are very worried about an agent doing something wrong and losing their customers' money and rightly so. This opens up an opportunity for a startup to make an autonomous agent and find users who are willing to take that risk!

Consumer robotics is going to be a very interesting market to watch and learn how companies navigate this. Especially when trying to understand the risks consumers are actually willing to take - will consumers be okay with someone tele-operating a robot in their house? I probably am but I’d guess most aren’t so building your business model around this might not make sense!

If you’re a founder, navigating this safety vs utility trade-off, reach out to me on ssukumar@balderton.com

Guardar com um clique

Faça leitura aprofundada de artigos virais com IA no YouMind

Guarde a fonte, faça perguntas específicas, resuma o argumento e transforme um artigo viral em notas reutilizáveis num único espaço de trabalho com IA.

Explorar o YouMind
Para criadores

Transforme o seu Markdown num artigo 𝕏 impecável

Quando publica os seus próprios textos longos, formatar imagens, tabelas e blocos de código para o 𝕏 é uma dor de cabeça. O YouMind transforma um rascunho completo em Markdown num artigo 𝕏 impecável e pronto a publicar.

Experimente Markdown para 𝕏

Mais padrões para decifrar

Artigos virais recentes

Explorar mais artigos virais