My interview experience with @Cloudflare for AI Engineer role.

@aditya_ghai07
ENGLISHAug 09, 2026
112K
1.0K
27
15
1.7K

TL;DR

Aditya Ghai details his interview process for Cloudflare's AI Workers team, highlighting the importance of cold emailing and deep technical knowledge in AI inference optimization.

CTC: 44-58 LPA

I cold emailed an engineer at @Cloudflare, we had a bit of discussion around the role and my background, and a few days later I got an interview invite with a link to pick a slot.

The first round was with a Lead Staff Engineer with 10+ YOE and was focused on the AI Workers team.

Cloudflare’s AI stack is pretty interesting: running models on serverless GPUs across their global network, with inference closer to users, while increasingly moving into larger and frontier scale models.

I wasn't very familiar with this side of Cloudflare beforehand, so we spent some time discussing what the team has been building and how they are approaching fast, efficient AI inference at scale.

Then we got into my own experience.

During my internship, I worked on model deployment, rooflining and inference optimisation of encoder based models, mainly rerankers and embedding models, running in-house on GPUs.

These were encoder-only models, mostly stacks of MLPs, so the usual LLM inference stack around attention, KV cache, vLLM or SGLang wasn't really relevant in these cases.

We went fairly deep into:

how I approach rooflining and profiling choosing batch sizes and measuring their impact scaling inference workloads finding bottlenecks across the inference pipeline optimising p50 and p99 latency TTFT and overall response latency

We also discussed how I would approach further optimisation once the obvious bottlenecks had been removed.

One question I particularly liked:

What does frontier AI mean to you?

The round went well and we even discussed the next rounds, which would have involved hands-on PyTorch.

Didn't make it through in the end. I think they were looking for someone with more hands-on LLM inference experience. I had been actively studying LLM inference and systems, but at the time I didn't have much relevant production experience in that area. Most of my hands-on work had been around distributed training, ML systems and infra.

Still, a really good interview experience. The discussion gave me a much better understanding of the problems involved in serving AI models at Cloudflare's scale.

One-click save

Use YouMind for AI deep reading of viral articles

Save the source, ask focused questions, summarize the argument, and turn a viral article into reusable notes in one AI workspace.

Explore YouMind
For creators

Turn your Markdown into a clean 𝕏 article

When you publish your own long-form writing, images, tables, and code blocks make 𝕏 formatting painful. YouMind turns a full Markdown draft into a clean, ready-to-post 𝕏 article.

Try Markdown to 𝕏

More patterns to decode

Recent viral articles

Explore more viral articles