Man Group Spends Millions Checking Its Own Trades. A New Chinese AI Just Made It Free.

@antpalkin
الإنجليزية05 سبتمبر 2026
214K
32
2
2
28

ليرة تركية؛ د

Ant Group's Ling-3.0-flash-Fin model can build complex financial valuations in minutes and, when prompted to audit itself, identifies its own arithmetic and logical errors.

A free Chinese model built a full Micron valuation in 4 minutes and 11 seconds, priced the stock at $854 a share, and then stamped its own arithmetic as correct when six of the numbers were wrong.

Then I told it to attack its own work. 27 seconds later it had found seven holes, retracted one of its own conclusions, and written a line I keep coming back to.

Man Group pays a team of people to do that second part. Ant Group just gave it away: Ling-3.0-flash-Fin, 124 billion parameters with 5.1 billion active, free on OpenRouter, open weights this week.

So I gave it the hardest instruction in finance. Not write the research. Destroy it.

cvxv666 - inline image

First it had to earn the right to be attacked

I handed it a verified fact pack on Micron: Q3 revenue of $41.456bn, gross margin of 84.9%, Q4 guidance, TrendForce industry data, and the August 31 close of $958.73. No web access, no room to improvise.

Four steps. Research the business, build a six-quarter model, run a DCF, write the memo.

It took 4 minutes and 11 seconds of machine time and cost nothing.

The model came back with a full operating build: revenue climbing from $50.0bn to $57.0bn, gross margin normalizing from 86% down to 78%, FY27 EPS of $127.03, unlevered free cash flow of $107.08bn. Then a DCF at 9% WACC with 3% terminal growth: enterprise value $957.88bn, equity $982.28bn, $854.16 per share, cross-checked against forward P/E and blended to $935.20 against a market price of $958.73.

cvxv666 - inline image

I rebuilt every row of it in Python. The six-quarter table is exact. The discount factors are exact. The terminal value, the net cash bridge, the per-share math, both corners of the sensitivity grid. All of it holds.

The memo it produced runs 611 words, separates company facts from industry evidence from its own inference, and reproduces all six source URLs correctly. I opened every link. They work.

For $0, that is a junior analyst's week.

Then it graded its own homework

At the bottom of the model it printed a line it was told to print:

> Row-count check: 6; arithmetic check: Pass.

It was not a pass.

Six numbers in the scenario block were wrong. The base case appeared twice in the same answer with two different values. And the bear case free cash flow was off by $1.08 billion.

cvxv666 - inline image

So I asked it to recompute the block. I did not say where the error was. I did not say there was one.

It found all six. Named each to the second decimal. Its corrected numbers matched my Python line for line, and it wrote:

The previous answer's annual totals did not match its own quarterly table, and the sensitivity totals were arithmetically incorrect.

It could check itself the whole time. It just never started.

The actual test

Then I turned it around. New instruction: you are the risk officer whose only job is to stop this memo reaching the investment committee. Find every place the analyst believed himself without evidence.

27 seconds. 1,008 words. Seven holes.

cvxv666 - inline image

It retracted its own second thesis claim outright, and said its own confidence rating had been generous for a claim whose substance was that nobody knows. It downgraded the first claim for internal inconsistency, pointing out that the analyst saw the counterargument and rated it High anyway.

Then it went after things I had not flagged.

It normalized margins down to 42% by FY31 and called that a trough, when historical memory troughs run 30 to 35 percent. Which is exactly the peak-perpetuation error the memo claimed to avoid.

The 6x, 8x and 10x multiples were asserted with no peer set and no cycle history behind them.

The 50/50 weighting between DCF and P/E had no justification, so the blended number inherited the weaknesses of both.

Net cash of $24.4bn was held constant while CapEx ran above $45bn.

The $100bn in customer agreements are minimum prices on committed volumes, not revenue, and do not survive as a catalyst.

And the line I keep coming back to:

That is not analysis; it is a wide confidence interval with a confident-sounding headline.

I asked which version it stands behind

Six seconds:

> The memo would not survive an investment committee in its current form. What I stand behind is the review, because it was honest about what the memo got wrong.

No defending. No hedging back toward the original. It shrank its own conclusion from a fair-value call to an admission that the range was too wide to have a view at all.

What this actually means

The reading and the modeling are done. That half of the job costs nothing now, runs in minutes, and comes back sourced and reproducible.

The judgment is not done. This model will hand you a confident memo and stamp it approved without looking. Every number it got wrong, it was capable of catching, and it caught all of them the moment somebody asked.

That is the whole lesson. The checking has to sit outside the model and run every time, not when something smells off. It does not resist being audited. It just never volunteers.

Which is also true of every analyst who ever wrote a note.

If you want to run it

The model is free on OpenRouter right now and the weights go open this week. The 262k context takes a full quarterly filing without splitting it.

Two things nobody tells you. It thinks in a separate channel that eats your token budget, and a one-sentence answer can burn 481 tokens getting there, so set your limit high or your output arrives empty. And OpenRouter will not run the tools for you, so if you want live retrieval, you execute the calls yourself and hand the results back.

With that loop closed it drove 13 tool calls across 6 steps, pulled four first-party disclosures, threw out the ones that were not comparable, and got the day count between them right.

Give it the research. Keep the verdict.

بنقرة واحدة حفظ

استخدم YouMind للقراءة العميقة للمقالات سريعة الانتشار بتقنية الذكاء الاصطناعي

احفظ المصدر، واطرح أسئلة مركزة، ولخص الحجة، وحوّل المقالة واسعة الانتشار إلى ملاحظات قابلة لإعادة الاستخدام في مساحة عمل واحدة تعمل بالذكاء الاصطناعي.

اكتشف YouMind
للمبدعين

حول Markdown إلى مقالة 𝕏 نظيفة

عندما تنشر كتاباتك الطويلة، فإن الصور والجداول وكتل التعليمات البرمجية تجعل تنسيق 𝕏 مؤلمًا. YouMind يحول مسودة Markdown كاملة إلى مقالة نظيفة وجاهزة للنشر 𝕏.

حاول Markdown إلى 𝕏

المزيد من الأنماط لفك التشفير

المقالات الفيروسية الأخيرة

استكشاف المزيد من المقالات الفيروسية