YouMind
تسجيل الدخول

We reverse engineered ChatGPT Intelligent UI. Here's how it actually works

@rabi_guha
الإنجليزية08 أكتوبر 2026
206K
2.0K
146
66
3.5K

ليرة تركية؛ د

This article details the technical architecture behind ChatGPT's newly launched Intelligent UI, explaining the use of a custom inference language (DIL), server-side compilation into JavaScript and JSON, and secure client-side execution via sandboxed workers.

ChatGPT launched Generative UI (finally) and renamed the field over night to Intelligent UI. We had to see how they implemented it.

https://x.com/OpenAI/status/2107894997538525580

This post is an in depth study of how OpenAI implemented their flagship Intelligent UI, the layers, formats and rendering natively across Web & Mobile.

Building blocks

ChatGPT's implementation divides the work between the model, the backend server, and the client:

  • Inference format: the model writes the interface in DIL, which combines Markdown with JSX-like tags and JavaScript.
  • Server-side compilation: the server converts each partial response into a JavaScript program and a JSON document of text and data.
  • Client runtime: a sandboxed runtime executes the program and produces UI operations.
  • Rendering: ChatGPT applies the operations to its own native components.
  • Design system and catalog: the components, properties, and design tokens available to the model.
Rabi Shanker Guha - inline image

ChatGPT Intelligent UI Architecture Diagram

Inference format

This is what the model writes. In ChatGPT, it is a language OpenAI calls DIL: Markdown for prose, JSX-like tags for components, and JavaScript for state and logic. We will follow one small response through every layer:

text
1## Team plan estimate
2Drag the slider to see the **monthly price** for your team.
3{@body const [seats,setSeats] = DIL.useState(8)}
4{@body const price = seats*29}
5<box border padding={3} gap={2}>
6 <slider min={1} max={50} value={seats} onChange={setSeats}/>
7 <title size="xl">${price}/mo</title>
8</box>

The heading and the paragraph are ordinary Markdown. The tags are components from ChatGPT's catalog. The two {@body …} lines are JavaScript: the first declares a piece of state, seats, and the second derives price from it. The slider is bound to seats, so moving it updates the price.

A dedicated format is needed because the model writes the interface token by token.

  • It has to be easy to write reliably, so it is built from notation the model already knows well.
  • It has to stay usable while half-written. Statements sit on their own lines, and any open element can be closed automatically. That lets the server cut a partial response at its last complete construct and still compile it.

Server-side compilation

The client never executes the model's output as written. OpenAI's server compiles it into a JavaScript program and a JSON document, which are stored with the message (as model_dil_v2). The response compiles to this (formatted for readability):

javascript
1function __dilSafe(evaluate, failureValue) {
2 try { return evaluate(); } catch { return failureValue; }
3}
4
5DIL.render(__dil.jsx(() => {
6 const __dilConstants = DIL.useConstants();
7 const __dilModelDataBindings = DIL.useAppData((appData) => appData.opGenui?.modelDataBindings ?? {});
8 const [seats, setSeats] = DIL.useState(8, { key: "seats" });
9 const price = __dilSafe(() => seats * 29, undefined);
10 return __dil.jsx(__dil.Fragment, null,
11 __dil.jsx("title", { size: "lg" }, __dilConstants["0"]),
12 __dil.jsx("text", null, __dilConstants["1"], __dil.jsx("bold", null, __dilConstants["2"]), __dilConstants["3"]),
13 __dil.jsx("box", { border: true, padding: 3, gap: 2 },
14 __dilSafe(() => __dil.jsx("slider", { min: 1, max: 50, value: seats, onChange: setSeats }), null),
15 __dil.jsx("title", { size: "xl" }, __dilConstants["4"], __dilSafe(() => price, null), __dilConstants["5"])));
16}, { key: "body:2" }));
json
1{
2 "constants": {
3 "0": "Team plan estimate",
4 "1": "Drag the slider to see the ",
5 "2": "monthly price",
6 "3": " for your team.",
7 "4": "$",
8 "5": "/mo"
9 },
10 "appData": { "opGenui": { "componentResults": {}, "modelDataBindings": {} } }
11}

The Markdown is compiled into the same tree as the components. The heading becomes a title, the paragraph a text with a bold inside it, and their words move into the constants table.

Compilation does work that every client would otherwise have to repeat:

  • Plain function calls. Markup becomes calls to __dil.jsx, so a JavaScript runtime can evaluate the program without a parser for DIL.
  • Error isolation. Expressions are wrapped in __dilSafe, so an expression that throws removes one element instead of aborting the whole render.
  • Text in a separate table. Static text moves into the constants table, so as a response streams, growing text changes the data rather than the program.
  • Stable state keys. Each piece of state receives a key ({ key: "seats" }), so its value survives every recompilation.
  • Repair and validation. Incomplete statements and tags are dropped, unclosed elements are closed, and properties that fail validation against the catalog are removed and recorded as diagnostics.

The JSON document holds the text constants and any data the server resolves for the response, such as image search results (see Data).

Client runtime

The client receives the compiled program and the JSON document. Its work is split between a runtime, which executes the program, and a renderer, which draws the result.

The program is model-written code, so it does not run in the ChatGPT page. ChatGPT loads a hidden iframe (runner.html), sandboxed with allow-scripts and a content security policy of default-src 'none', which starts a Web Worker.

  • Lockdown. Before evaluating a program, the worker removes network access, timers, messaging, and dynamic code evaluation from its global scope, and freezes the remaining globals.
  • Evaluation. It then evaluates the program with new Function. The runtime objects (DIL, __dil, GenUI) and the catalog's composite components are passed in as parameters.
  • Watchdog. A program that does not respond within a timeout is quarantined, and the worker is restarted.

The runtime is a small reconciler in the style of React. It renders the component and keeps hook state in keyed slots. It then compares the resulting tree with the previous one and encodes the differences as a list of operations. It does not draw anything.

The following illustrative example shows the operations from a first render, with one line per node. Entries that list each element's property names are omitted:

text
1CREATE #1 title SET size = "lg" PLACE under root at 0
2CREATE #2 text "Team plan estimate" PLACE under #1 at 0
3CREATE #3 text PLACE under root at 1
4CREATE #4 text "Drag the slider to see the " PLACE under #3 at 0
5CREATE #5 bold PLACE under #3 at 1
6CREATE #6 text "monthly price" PLACE under #5 at 0
7CREATE #7 text " for your team." PLACE under #3 at 2
8CREATE #8 box SET border = true, padding = 3, gap = 2 PLACE under root at 2
9CREATE #9 slider SET min = 1, max = 50, value = 8, onChange = fn#1 PLACE under #8 at 0
10CREATE #10 title SET size = "xl" PLACE under #8 at 1
11CREATE #11 text "$" PLACE under #10 at 0
12CREATE #12 text "232" PLACE under #10 at 1
13CREATE #13 text "/mo" PLACE under #10 at 2

Functions never leave the worker; the slider's handler is sent only as an identifier (fn#1). On the wire, the operations are encoded as a binary sequence of integers, with strings held in a separate table.

Rendering

The ChatGPT page applies the operations to its own component tree. Each CREATE instantiates a native component from ChatGPT's design system, and the page animates changes as they arrive. The page accepts operations only for known component types, so model output cannot introduce arbitrary markup or styles. The exceptions are the raw CSS values some properties accept (see Design system and catalog) and AppBlock apps, which run in an iframe (see Escape hatch).

Interaction runs in the opposite direction. When the user drags the slider to 9, the page sends the handler's identifier and arguments to the worker. The worker calls setSeats(9), re-renders, and returns update operations. No model call is involved.

Design system and catalog

The catalog defines what the model can request. It is needed because the model does not build an interface from raw layout and styling rules. It chooses from components ChatGPT already knows how to draw, and styles them with design tokens such as padding={3}. Raw CSS values, such as pixel widths and hex colours, are accepted for some properties, but design tokens are preferred. As a result:

  • Generated interfaces look like the rest of ChatGPT on every platform.
  • The compiler has a schema to check output against. A property that does not exist on a component, or a literal of the wrong type, is removed during compilation and recorded as a diagnostic.

In one response we captured, the compiler removed two properties: fill on an icon (an unknown_prop diagnostic) and gap="1" on a box (an invalid_literal diagnostic).

The catalog has three parts:

  • Native components. Around 70 components are defined in the component registry in ChatGPT's client code; 39 of them appear in the responses we captured.
  • Design tokens for spacing, radius, colour, and size.
  • Composite components written by OpenAI in DIL and sent to the sandbox prebuilt, such as the image and product components. In our captures, the model used these components but never defined its own.

Streaming

Streaming text is simple: each new token is appended to what is already on screen. Streaming an interface is harder, for three reasons:

  • The output is usually not runnable yet. At most moments it is an incomplete program, with a tag or expression still open, and it cannot be executed as written.
  • The interface has to keep working while it grows. Components the user has already touched must keep their state.
  • Some content arrives separately. Data such as images comes from the server, not from the text.

Server Side Streaming

A plain stream of appended tokens cannot express this. ChatGPT instead streams patches to a structured message that holds the raw text, the compiled program, and its data side by side.

The response reaches the browser over a server-sent event stream (POST /backend-api/f/conversation). Each event is a JSON-Patch-style update to the message being built. A single event usually updates the raw DIL text and its compiled form together. This is one update from a captured response, shortened:

json
1{"o": "patch", "v": [
2 {"p": "/message/content/parts/0", "o": "append", "v": " Sunday lamb roast with friends — generous food, …"},
3 {"p": "/message/metadata/model_dil_v2/code", "o": "replace", "v": "DIL.render(__dil.jsx(()=>{…"},
4 {"p": "/message/metadata/model_dil_v2/constants", "o": "append", "v": {"0": "Here's a plan for a proper Sunday lamb roast with friends — …"}},
5 {"p": "/message/metadata/model_dil_v2/constants", "o": "append", "v": {"1": "Since you're"}},
6 {"p": "/message/metadata/model_dil_v2/fallbackMarkdown", "o": "append", "v": " Sunday lamb roast with friends — …"}
7]}

The server does not compile incrementally. Every few hundred milliseconds, most likely with each new chunk of model output, it recompiles everything the model has written so far and sends the result. Compilation starts with the first token, before any tag has appeared.

The compiler has to take care of:

  • Compiling a half written response
  • Updating text
  • Updating UI

The timeline looks something like this:

Rabi Shanker Guha - inline image

GIF

Client Side Streaming

The page passes each new update to the sandboxed worker. The worker evaluates it, re-renders with the existing state, and sends update operations to the page. State keeps its values across recompiles because of the keys added during compilation. If a new program fails to evaluate or render, the worker keeps the last one that worked.

The page then animates each change:

  • text fades in over 0.7 s;
  • new rows and grid items slide in over 0.42 s;
  • charts draw over 1.8 s;
  • container heights transition instead of jumping.

Escape hatch: AppBlock, an app in an iframe

Rabi Shanker Guha - inline image

ChatGPT generated inline app

Some requests call for things the native components are not designed for, such as a drum machine that synthesizes sound with Web Audio. For these, the model can write an AppBlock: a self-contained web app in HTML, CSS, and JavaScript, embedded in the response. This is the start of one, shortened:

text
1<AppBlock title="Drum Lab" icon="app-chatgpt" variant="inline" app_block_id="drum-lab-01">
2<div id="dl" class="w-full min-w-0 space-y-4 text-base">
3 <style>
4 #dl{color:var(--viz-text)}#dl button{touch-action:manipulation}#dl .panel{background:var(--viz-panel);border:1px solid var(--viz-border);border-radius:15px}…
5 </style>
6 …
7 <button id="dl-play" class="btn" style="background:var(--viz-text);color:var(--viz-card);min-width:100px">▶ Play</button>
8 …
9</div>
10<script>
11(function(){
12const root=document.getElementById('dl');if(root.dataset.init)return;root.dataset.init="yes";
13…
14function audioInit(){if(!audio){const C=window.AudioContext||window.webkitAudioContext; if(!C)return false;audio=new C();…
15…
16})();
17</script>
18</AppBlock>

AppBlocks are rendered differently than Intelligent UI components.

Bringing it all together

You type a prompt. The model starts writing the interface, the server turns it into something ChatGPT can run, and the page builds it piece by piece as the response streams in. Once it’s there, moving a slider or ticking a checkbox updates the interface locally, without asking the model again.

A model specific language, a server compilation step, native renderers and a grounded design system bring it all together. Each part serves a critical part and together they serve the next generation of AI native interfaces to their billions of users worldwide. What a time to be alive!

Rabi Shanker Guha - inline image

Screenshots of ChatGPT Intelligent UI

Methodology

All observations come from our own ChatGPT accounts, from the traffic the ChatGPT web app generates, and from the JavaScript that chatgpt.com serves publicly. They were made in October 2026 with GPT-6 and GPT-6 Thinking.

Analysis done with the help of Codex & Claude. Written with Codex, Visualisations by Claude

(A more in-depth version is at https://www.openui.com/blog/how-chatgpt-intelligent-ui-works

بنقرة واحدة حفظ

استخدم YouMind للقراءة العميقة للمقالات سريعة الانتشار بتقنية الذكاء الاصطناعي

احفظ المصدر، واطرح أسئلة مركزة، ولخص الحجة، وحوّل المقالة واسعة الانتشار إلى ملاحظات قابلة لإعادة الاستخدام في مساحة عمل واحدة تعمل بالذكاء الاصطناعي.

اكتشف YouMind
للمبدعين

حول Markdown إلى مقالة 𝕏 نظيفة

عندما تنشر كتاباتك الطويلة، فإن الصور والجداول وكتل التعليمات البرمجية تجعل تنسيق 𝕏 مؤلمًا. YouMind يحول مسودة Markdown كاملة إلى مقالة نظيفة وجاهزة للنشر 𝕏.

حاول Markdown إلى 𝕏

المزيد من الأنماط لفك التشفير

المقالات الفيروسية الأخيرة

استكشاف المزيد من المقالات الفيروسية