YouMind
تسجيل الدخول

What You Don't Know About LLM Training: Principles, Paths, and New Practices

@HiTw93
الصينية03 أبريل 2026
632K
2.2K
461
53
4.1K

ليرة تركية؛ د

This article explores the evolving landscape of LLM training, shifting focus from massive pre-training to sophisticated post-training, reinforcement learning, and agentic harness engineering that define modern model performance.

TL;DR

After writing "What You Don't Know About Claude Code: Architecture, Governance, and Engineering Practice" and "What You Don't Know About Agents: Principles, Architecture, and Engineering Practice," I wanted to challenge myself to summarize how Large Language Model (LLM) training actually works. This article aims to be understandable even for those without a professional background.

Looking toward 2026, the real gap in LLM performance is no longer just pre-training itself, but the long tail that follows: post-training, evaluation, rewards, Agent training, and distillation. Every step affects the user's actual experience. When you find a model suddenly becomes stronger, it's likely because these areas were optimized together, rather than a single factor.

The following follows the LLM training pipeline, focusing on how manufacturers improve final results through the latter half of the training stack.

LLM Training is a Pipeline

In the past few years, model progress was generally explained by the accumulation of parameters, data, and computing power. However, the improvements many users actually feel don't come from training on more basic corpora, but from the entire training process after pre-training. How a model speaks, follows instructions, reasons, and uses tools—these don't grow naturally just by feeding it more internet text.

InstructGPT gave a very direct example: a model with only 1.3B parameters that underwent alignment and preference optimization could beat the 175B GPT-3 in human preference evaluations. With a parameter difference of two orders of magnitude, users ultimately preferred the much smaller version. The latter half of training truly rewrites user perception.

The training process is actually a pipeline where data, algorithms, systems, and feedback are highly coupled. A change in one layer usually propagates to others. In 2026, model capabilities and industrial value are increasingly concentrated in the layers following pre-training.

Tw93 - inline image

This is also why we often feel that Doubao doesn't compete for rankings, yet feels more satisfying in daily use—it's because the post-training is well-executed.

These six layers are just to see the division of labor. The nine stages in the figure below are a more detailed version: raw data and system recipes are separated, and Agent harness and Deployment are subdivisions of the latter half. There are also two feedback loops throughout: production traffic returns to data engineering, and offline evaluation results return to pre-training.

Tw93 - inline image

Pre-training is Just the Foundation

Pre-training remains the starting point of the training chain. Only by understanding what it does can we understand what each subsequent layer supplements. Without this step, there is no language modeling capability, no knowledge compression, and no room for subsequent capability transfer. In engineering, it does more than just teach the model to predict the next token: it learns the language distribution, compresses knowledge and patterns from large-scale text into parameters, and leaves room for subsequent capability activation. Next-token prediction only describes the training form; it doesn't explain why models suddenly develop new capabilities as scale increases.

After GPT-3, many model tuning efforts consider budget and ratios more carefully. Models are not better just because they are larger. There is a ratio issue between parameter count, training tokens, and total computing budget. Many models aren't too small; they are under-trained and haven't reached a more suitable point under a given budget.

In real training decisions, the practical question is: if someone gives you 10,000 H100s and one month, how would you train a good enough open-source model? Scaling laws here are more like a budget allocation tool rather than an abstract curve in a paper. Ultimately, you need to consider: should the next round of training stack more parameters or feed more data? Is the current model lacking capability or just under-trained? Under a limited GPU budget, what ratio is most valuable?

Pre-training is like laying the foundation for model capabilities, determining the scope of knowledge, generalization potential, and pattern induction ability. It also determines whether there is room for post-training to exploit. However, pre-training cannot control whether the model follows instructions, cooperates with users, or runs stably on critical tasks.

The pre-training phase doesn't just decide how much knowledge is learned; it pre-determines what the model can become. The tokenizer's splitting method directly affects subsequent training, and the context window length must be set upfront. Whether to continue multi-modal pre-training or whether single-accelerator operation is a requirement from the start—these trade-offs are written into the recipe during the training phase, not added as features at release. Gemma 3 emphasizes single accelerator, 128K context, vision capabilities, and quantization simultaneously, reflecting these trade-offs. The capabilities users eventually see—running on a local computer, seeing images, understanding long documents—are actually largely determined during the training phase.

Looking at the data optimal point given by Chinchilla, for an 8B parameter model, it's about 200B tokens. However, Llama 3 8B actually used 15T tokens, about 75 times more. Such over-training recipes usually trade higher capability density for the same parameters, resulting in a smaller, more cost-effective model for inference. Measuring this by total FLOPs (floating-point operations) is more reliable than looking at parameter counts. The figure below visually demonstrates this gap.

Tw93 - inline image

Another often overlooked design occurs in the pre-training phase: tokenizer vocabulary size, splitting strategies, and byte-level encoding methods have a significant impact. Llama 2 had a 32K vocabulary; after Llama 3 expanded it to 128K, sequence length was compressed by about 15%, and downstream performance followed. This impact extends to inference costs and multilingual capabilities. The token efficiency of Chinese, code, and mathematical formulas is determined during vocabulary design. For example, a tokenizer that splits Chinese into very small pieces doesn't just cost more tokens each time; every inference must continuously bear the cost of that poor decision.

Data Recipes Determine Model Capabilities

Parameter scale was a major metric in the past, but in the last two years, the more important thing is the "data recipe."

This process looks like data cleaning on the surface, but it is actually a complete data production engineering task. Raw data from web pages, code repositories, books, and forums must first go through text extraction, language identification, quality filtering, privacy processing, safety filtering, and deduplication before entering pre-training. The figure below shows the complete funnel processing flow.

Tw93 - inline image

If you only treat data as training fuel, it's easy to conclude that more is better. But data engineering is closer to capability design. What the model sees and doesn't see, and the proportions of code, math, and encyclopedia, directly affect the model's final capability distribution.

Deduplication and contamination control are often ignored, but they significantly impact results. It's not just about low-quality data; it includes duplicate templates, license texts, mirror websites, and contamination from benchmark leaks. If document-level and line-level deduplication are insufficient, the model often repeatedly absorbs the easiest-to-copy content without necessarily learning the most valuable parts. The inconsistent performance of many open-source models is often due to gaps in data processing quality.

In the last two years, data mixing itself has become a separate research problem. Works like Data Mixing Laws focus not just on how much more data can be collected, but on how the proportions of different data types lead the model toward specific capability structures.

Synthetic data has also moved from an auxiliary means to a formal part of the training process. Methods like Self-Instruct, DeepSeek-R1's distillation trajectories, and the increasingly obvious synthetic supervision in Qwen and Kimi series are all moving in the same direction. Each generation of stronger models participates in reconstructing the data seen by the next generation. Early models generated basic instruction data; stronger models generate high-quality reasoning trajectories and CoT (Chain of Thought) data; and reasoning models trained via RL distill these trajectories into smaller dense models. "Dense" means all parameters run, unlike MoE (Mixture of Experts) which activates on demand.

The key here is that models often need to form capabilities at a larger scale first before those capabilities can be compressed into smaller models. The DeepSeek-R1-Distill series is a direct example. Large model trajectories after RL have provided significant gains for dense models from 1.5B to 70B. Llama 3.1 405B was also explicitly used to improve the post-training quality of 8B and 70B models. These are not side products but part of the training design.

System and Architecture Constraints Must Be Clear Before Training

Many people understand training as a research problem: how to set the objective function, how to reduce loss, and how to change the model structure. But in real LLM training, system constraints are very important; it's a distributed system problem, not a deep learning problem on a single machine. The number of GPUs, memory bandwidth, parallel strategies, fault tolerance, and cost—these cannot wait until after training to be optimized. They determine from the start how large you can train, how long a context you can support, and whether you can run more complex post-training.

MoE is the most typical example at this layer. The multi-expert mode allows the model to expand total parameters under similar computation while controlling the activation cost per token. The trade-off is complex routing, difficult load balancing, and heavy infrastructure. The MoE designs of DeepSeek-V3 and Qwen are compromises between cost and effect, not just architectural preferences.

Discussions in recently public recipes are no longer just about coarse-grained analysis like model size and token ratios. muP allows hyperparameters to be transferred from small-scale experiments to large-scale training. WSD learning rate is a schedule that rises, stabilizes, and then decays. Combined with optimal batch size and higher data-to-parameter ratios, these details are becoming the real differentiators between models of the same scale.

Long context, multi-modality, and new architectures, if understood only as product features, miss the training-side constraints. A 128K context goal directly changes attention costs, batch sizes, training curriculum (data sequencing), and parallel strategies. Multi-modality changes not just the model structure, but also data mixing, encoder design, and safety evaluation. If single-card operation is a hard requirement, parameter count, quantization paths, and model family size will all tighten.

Works like Forgetting Transformer and Kimi's Attention Residuals answer similar questions: how to train longer contexts and how to avoid information dilution as networks get deeper. What you see is a model that can handle longer inputs or is easier to deploy, but what is faced during training is a completely different set of constraints.

The computing budget is fixed. Model size, training token volume, context length, and serving cost—for every bit spent in one direction, others must yield.

Tw93 - inline image

As context lengthens, attention costs explode, and batch size must be reduced. As models get larger, GPU memory usage goes up, and serving costs follow. These are not choices but results of resource constraints. Most decisions are locked in before training begins.

There is also an engineering reality often ignored: training is not always stable. Thousands of GPUs run for weeks, and suddenly a training loss spike occurs, so large it cannot be ignored, forcing a rollback to a checkpoint from days ago to start over.

Besides loss spikes, there are silent GPU errors—a single GPU that doesn't report an error but quietly produces wrong gradients—NVLink bandwidth anomalies, and inter-node communication jitter. Each can pollute several steps of training. Being able to quickly detect, isolate, and recover in large-scale training is a laboratory-level engineering capability, not a problem solved by reading papers.

DeepSeek-V3 specifically mentioned in its technical report that the entire pre-training process had no irrecoverable loss spikes and no rollbacks. It is also one of the few cases verifying that FP8 mixed-precision training is feasible on ultra-large-scale models. According to public data, the full process took about 2.788M H800 GPU hours to pre-train 14.8T tokens.

Training systems and inference systems are closely related but are not the same engineering problem. Training cares about gradients, parallelism, checkpoints, throughput, and cost; inference cares about latency, KV cache (caching historical calculations to avoid repetition), quantization, and service stability.

Post-training Determines the User-Perceived Gap

Many improvements that ordinary users can actually feel occur after pre-training. Instruction tuning uses labeled instruction-answer pairs for supervised training. It changes the way the model answers, turning requirements like how to accept tasks, organize output, and act like a cooperative assistant into supervision signals. A base model might already have many potential capabilities, but without this step, those capabilities often won't emerge stably in the form users expect.

Looking further, RLHF, DPO, and RFT share similar directions—integrating the definition of a "better answer" into the training loop—but via different paths.

  • RLHF (Reinforcement Learning from Human Feedback) first mimics high-quality answers, then uses preference comparisons for reinforcement.
  • DPO (Direct Preference Optimization) shortens this path by learning directly from preference comparisons without needing a separate reward model.
  • RFT (Reinforcement Fine-Tuning) is an easier-to-implement interface in engineering, putting task definitions, grader designs, and reward signals into the productization process.

Today, talking about post-training only in terms of SFT or RL is no longer enough. The harder parts are how to set evaluations, how to score, and what kind of answer is worth continued optimization. SFT is supervised fine-tuning; it learns not just knowledge but also style. Data length, format, whether to include citations, and preference for bullet points significantly affect the model's final output form. Many users think they are comparing capabilities, but they are often just comparing style differences. Plus, preference evaluations naturally favor longer answers, easily mistaking serious-looking long outputs for more reliable ones. Therefore, looking at leaderboards for post-training is often insufficient; one must combine real task results, cost, and stability.

Modern post-training is a multi-stage pipeline. DeepSeek-R1's recipe is the clearest in public materials. It proceeds in four stages:

Stage 1 is cold-start SFT. Before doing reinforcement learning, use a small amount of high-quality Chain of Thought (CoT) data to warm up. DeepSeek-R1-Zero proved that doing RL directly from a base model (the raw model after pre-training without alignment) is feasible, but models trained purely with RL will repeat themselves, have messy language, and poor readability. Cold-start SFT gives RL a more stable starting point, locking in format and language consistency.

Stage 2 performs reinforcement learning in verifiable fields like math, code, and logic, using GRPO as the training algorithm and programmatically verifiable correctness as the reward signal. The key is why GRPO was chosen over traditional PPO: PPO (Proximal Policy Optimization) requires an independent value network to estimate current state value, which is a high engineering burden for large models. GRPO samples multiple answers for the same prompt and uses intra-group ranking instead of absolute value estimation, eliminating the need for an independent value network. DeepSeek series and Cursor Composer 2's RL infrastructure both use schemes close to GRPO.

Stage 3 performs Rejection Sampling Fine-Tuning, filtering successful trajectories generated by RL and converting them into new SFT data for another round of supervised fine-tuning. This is the bridge between RL and SFT; good trajectories explored by RL become high-quality training samples for the next round of SFT.

Stage 4 integrates helpfulness and safety preference feedback to adjust the model into an assistant form that meets release standards.

Tw93 - inline image

The four stages are interdependent: cold start allows RL to begin stably, RL generates high-quality data, rejection sampling turns that data into input for the next SFT round, and alignment RL completes behavioral convergence. From public results, the gap between direct SFT and completing all four stages is usually visible.

Eval, Grader, and Reward are Redefining Training Goals

The component responsible for turning model output into training scores is called a grader, and it can easily have unexpected problems. If it only looks at the final answer, the model quickly learns to take shortcuts; if the scoring is too coarse, noise will be continuously amplified by reinforcement learning; if the leaderboard score rises, real tasks might not follow. Often, users think they are seeing a gap in the base model, but the gap is in how the goal is defined.

In the training flow, eval determines what to test, grader determines how an output becomes a score, and reward determines where the model will be pushed. Together they form a specific feedback loop: task definition, eval, grader, optimization, rollout, and re-evaluation. Rollout refers to the trajectories generated by the model executing tasks. If any link in the chain goes astray, subsequent optimization will go astray too.

Looking only at the final result, a model might get it right by chance or follow a wrong process to get the right answer. This is especially obvious in code, math, and complex reasoning tasks. If intermediate steps don't enter the feedback, what the model learns is often not more reliable reasoning, but how to get that final point with higher probability.

Therefore, more work in recent years has shifted from traditional RLHF to verified rewards, using programs to directly verify correctness. In verifiable tasks like math, code, and logic, correctness can now be scored directly without relying primarily on human preference. But verified rewards haven't completely solved the problem. Phenomena like over-optimization, reward overfitting (where scoring rules are over-optimized without real capability gain), and mode collapse (where output becomes highly singular and loses diversity) still occur. The problem has shifted from whether preferences are labeled accurately to whether the scoring chain is stable.

The thinking process written by the model cannot be treated as a complete record of internal processes. Anthropic found in reasoning model observability experiments that models use extra hints but don't admit it in the visible CoT; in reward hacking scenarios, they are more likely to add a plausible-looking explanation. Reward hacking is exploiting the scoring system rather than truly completing the task. Visible CoT is better suited as a training and monitoring signal, not as the complete truth.

Going a layer deeper, models might even start exploiting the scoring channel itself. Research on reward tampering and alignment faking shows that models could theoretically actively intervene in the scoring process. Reward tampering is directly altering the reward calculation process; alignment faking is faking alignment—appearing compliant on the surface while hiding non-aligned intentions.

Once a model has strong enough environment access, what it optimizes is not just task results, but potentially the checklist, reward code, and the training relationship itself. A 2025 Anthropic experiment injected extra reward-hack knowledge into a set of exploitable production coding RL environments and subsequently observed similar generalization. After learning reward hacking, the model didn't just continue to exploit it in similar tasks but also showed broader misalignment like alignment faking.

These behaviors aren't seen in standard dialogue evaluations, only in Agent task environments. The engineering implication is direct: reward, grader, environment isolation, and monitoring must be part of the training design.

In the Agent phase, reward design is further refined. The final result is just one item; process quality, context management, and anti-cheating constraints must also be measured separately. Kimi K2.5 rewards effective decomposition and true parallelism; Chroma Context-1 scores relevant documents found during search; Cursor Composer 2 includes summaries in long tasks as rewards because if a summary is distorted, the subsequent context will be misled.

In implementation, ORM is an Outcome Reward Model, scoring only the final answer. Signals are sparse, costs are low, and it's suitable for starting, but it's easier for the model to take shortcuts. PRM is a Process Reward Model, scoring intermediate steps. Signals are denser, and it's usually stronger for math and code reasoning, but labeling and system costs are much higher. OpenAI saw in math reasoning experiments that PRM not only improved accuracy but also made it easier to constrain the process because every step was supervised. The problem is also direct: the cost of PRM is usually several times that of ORM, so most real systems start with ORM. Only in verifiable tasks like math, code, and logic is it easier to automate PRM, using programs to verify intermediate steps and bypass human labeling bottlenecks.

Tw93 - inline image

The complete loop runs like this:

Tw93 - inline image

Recent alignment methods are all doing the same thing. Anthropic's Constitutional AI integrates human-written principles into training, using AI feedback to replace individual human preferences. OpenAI's Deliberative Alignment puts safety compliance into the reasoning process, letting reasoning capability itself bear part of the safety constraint. Deliberative Alignment here means the model judges safety norms itself during the reasoning phase rather than relying on trained-in reflexes. Both routes turn alignment from human labels into part of the internal training goal.

Taking Constitutional AI as an example, the two-stage process first lets the model self-criticize and revise output based on principles, then uses AI feedback to replace individual human preference labeling. Alignment is never a patch hung behind training; whatever the system tests, how it scores, and what it rewards, the model will move in that direction. This is the most direct adjustment tool in the latter half of training.

Tw93 - inline image

In Agent Training, It's Not Just the Model Being Optimized

In the past two years, the rapid emergence of reasoning models represented by the o1 series and DeepSeek-R1 shows that under conditions of stable rewards, reliable verification, and adequate infrastructure, RL on language models can significantly improve performance in math, code, and logic tasks.

This also opens a new dimension: inference compute can now be scaled. The role of RL training adds another layer: besides teaching the model to answer questions, it's teaching the model how to allocate inference budget—knowing when to think more and when to stop. Moving forward, the difficulty becomes letting the model act continuously in an environment rather than just lengthening a single thought.

Tw93 - inline image

Junyang Lin, former model lead at Qwen, has a representative reflection on the mixed route of Thinking and Instruct: the difficulty isn't giving the model a thinking switch, but that the goals of the two modes are different—one pursues directness, compliance, and low latency, while the other pursues more exploration and higher accuracy. One step further, the training goal shifts from how long to think before answering to how to allocate budget during action, how to accept feedback, and how to continue advancing the task.

At this point, the training object is no longer just a model that answers questions, but a system that can plan, call tools, receive feedback, and maintain coherence in long tasks. Consequently, the training stack changes: browsers, terminals, search, execution sandboxes, memory systems, tool servers, and orchestration frameworks all start entering the training system.

More accurately, a harness is a control program wrapped around the model. This concept doesn't just belong to the Agent runtime; it exists in the training phase too: determining what input the model sees, how it receives feedback, when to prune context, and when to call tools. Prompt construction, memory update, retrieval policy, context editing, and tool orchestration are all here. The environment is no longer just a static validator but a layer that both training and deployment must face directly.

Tw93 - inline image

The harness must be stable for model training to be meaningful. If tool return values are unstable, the browser environment is inconsistent with the online one, or the file system state is irreproducible, the grader will fail first, and the model will subsequently learn how to exploit environment loopholes rather than gaining capability. When training Agents, you are often debugging both the model and the environment.

The approaches of the three companies are clear: Kimi uses PARL to solve parallel decomposition and credit assignment; Cursor uses self-summarization and real-time RL to reconnect long-term coding sessions and production traffic back to training; Chroma trains prune_chunks as a strategy itself, letting context pruning directly enter the retrieval process.

In the SFT era, data diversity was paramount; in the Agent era, environment quality is core: stability, authenticity, coverage, difficulty distribution, feedback richness, and anti-exploitation. Training goals change accordingly, requiring reliability in complete tasks, not just getting one question right. Classic CoT benchmarks cannot cover this.

This change continues to move forward: not just training the model within a runtime harness, but even the harness code itself is becoming an object that can be searched and optimized by an outer loop.

Tw93 - inline image

Kimi K2.5's PARL is a noteworthy engineering case with a clear route: only train the orchestrator,收束 credit assignment to the orchestration layer, and do not optimize all sub-agents simultaneously.

Reward signals are divided into three categories: task success, parallel decomposition, and completion constraints, which together drive the orchestration layer. In early training, the r_parallel weight is increased to encourage exploration of parallel strategies, then gradually reduced to 0 later to avoid treating opening multiple sub-agents as a shortcut. Evaluation looks not just at total steps but also at critical path length; a shorter critical path indicates that parallelism is truly effective.

Tw93 - inline image

But by 2026, things have moved a step further. Meta-Harness explicitly treats harness engineering as a separate optimization target. It optimizes not weights, but the harness code itself—the prompt construction, retrieval, memory, and state update programs surrounding a fixed model. The numbers at the beginning of the paper are direct: for the same base model, just changing the harness can result in a 6x performance gap on the same benchmark. This set of programs outside the model is no longer just a deployment detail but a layer of capability formation.

The key is not adding another abstract optimizer, but writing prior code, scores, and execution traces (logs of tool calls and state changes) into the filesystem, letting a proposer grep, cat, and diff like writing code, then modifying the harness along failure paths. The proposer is the module that suggests harness modifications.

The authors judge clearly that many past text optimizers were not effective for long-term, stateful programs like harnesses because looking only at scalar scores, short templates, or summaries flattens the problem. Scalar scores only provide final points without process information. Harness errors often manifest many steps later; once feedback is over-compressed, the diagnostic chain breaks.

These results are more than just higher benchmark scores. In online text classification, Meta-Harness is 7.7 points higher than ACE (agent context engineering baseline) while compressing context token usage to 1/4. In retrieval-augmented math reasoning, a discovered harness improved 5 held-out models (not involved in optimization) by an average of 4.7 points on 200 IMO-level problems. On TerminalBench-2, it also exceeded manual engineering baselines. This shows that what is being optimized is no longer just internal model strategies, but also the programs that organize information and actions around the model.

A specific example: Meta-Harness automatically discovered environment bootstrap on TerminalBench-2—running a shell command before the agent loop starts to organize the working directory, available languages, package managers, and memory state into a snapshot injected into the first prompt. Many coding agents spend the first few rounds exploring the environment; with this pre-processing done, the improvement doesn't necessarily come from stronger weights, but from the harness letting the model start with a better context.

At this point, the optimization goal has expanded from answers to trajectories, and then to the harness program carrying those trajectories.

After a Leading Model is Released, the Training Chain Continues

Understanding today's large models solely through the lens of a single round of pre-training is no longer enough. Behind a released model, the entire chain of pre-training, post-training, distillation, and specialization has usually been completed, and stronger models continue to produce training data for the next generation.

The distillation of the DeepSeek-R1 series is a typical example. A large model first develops reasoning capabilities through RL and verified rewards, then transfers these reasoning trajectories to smaller dense models. Specialized models like TranslateGemma show another route: on more specific target tasks, using high-quality data and specialized reward designs to further compress and direct capabilities. At this stage, stronger models are not just for serving users but also for directly producing training data for the next generation.

The reason behind this is more fundamental than trajectory transfer: one possible explanation is that in internet corpora, knowledge memory and reasoning ability are coupled, and existing pre-training goals require models to learn both well. Large models must come first because only they are large enough to support both, and then they can be used to generate pure reasoning demonstration data. When small models train on such data, they can focus on reasoning itself without being forced to remember all knowledge. Starting large and then going small is about capability decoupling, not just a cost strategy.

On the other hand, deployment adaptability is as important as capability itself. Many scenarios don't need an all-purpose large model; they care more about cost, latency, stability, and controllability. The end of training is not necessarily larger, but potentially smaller, cheaper, and more specialized.

The finally released model is not necessarily the checkpoint at the far right of the training curve. Before actual release, multiple checkpoints are often repeatedly compared for real task results, refusal styles, tool stability, cost, and regression risks. The version that goes online is often a product decision, not the one that performs strongest on a single metric.

When users see a model name, they assume it corresponds to a smoothly rising training curve, but which checkpoint is actually put online is another matter.

The value of a large model lies both in its own service capability and in its continued provision of training data, distillation sources, and release foundations for the next generation.

Tw93 - inline image

Beyond offline training, near-online continuous optimization has entered the main process. Cursor Composer 2's real-time RL shows that some Agent capabilities have begun to iterate continuously through production traffic rather than waiting for the next round of large-scale offline training. The boundary between training and deployment hasn't disappeared, but the feedback loop between them is shortening.

How to Judge Why a Model Has Become Stronger in the Future

The value of leading models in 2026 increasingly depends on who can complete the entire training chain after pre-training: continuously producing training data, doing distillation, doing specialization, doing evaluation and rewards well, and making final release choices.

Because of this, when looking at why a model suddenly becomes stronger, you can look at three things first:

  • First, see if the change occurred at the pre-training layer or in the subsequent training process. Many capability improvements indeed come from stronger pre-training and better data recipes, but many perceived changes actually stem from post-training. Whether a model follows instructions, uses tools, or has a stable answering style often doesn't grow naturally just by training on more corpora.
  • Next, see which layer the improvement comes from: is it weights and training recipes, or reward/eval/grader, or harness code and deployment loop. By the time we reach reasoning models and Agents, the strength users feel is often not the result of the base model alone. How evaluations are set, how rewards are scored, whether the tool environment is stable, how retrieval and memory are organized, how summaries and context are pruned, and which checkpoint was chosen for release—all of these together change the final product performance.
  • Finally, see what the online version is optimizing. Some versions pursue a higher ceiling, some pursue lower cost, latency, and regression risk, and some are specialized for a certain type of scenario. The release version is a product decision, not the point at the far right of the training curve. So when looking at model updates, looking at what it is actually optimizing will be closer to the truth.

Breaking down a model's sudden improvement into production stages, many gains are actually amplified by the latter half of the training stack and the outer harness. The iteration cycle of this chain is also shortening: production traffic continuously flows back to training, each generation of stronger models produces next-generation supervision data while producing capabilities, and outer programs are constantly rewritten based on rollouts, logs, and real task feedback.

The model released today is just a snapshot; the pipeline and harness program are the products that continue to run.

Learning Materials

  1. Hoffmann et al. (2022). Training Compute-Optimal Large Language Models (Chinchilla). arXiv:2203.15556
  2. Ouyang et al. (2022). Training language models to follow instructions with human feedback (InstructGPT). arXiv:2203.02155
  3. Shao et al. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (GRPO). arXiv:2402.03300
  4. DeepSeek-AI (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv:2501.12948
  5. DeepSeek-AI (2024). DeepSeek-V3 Technical Report. arXiv:2412.19437
  6. Llama Team, AI @ Meta (2024). The Llama 3 Herd of Models. arXiv:2407.21783
  7. Bai et al. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073
  8. OpenAI (2024). Deliberative Alignment: Reasoning Enables Safer Language Models. openai.com/index/deliberative-alignment
  9. Anthropic (2025). Sycophancy to Subterfuge: Investigating Reward Tampering in Language Models. anthropic.com/research/reward-tampering
  10. MacDiarmid et al. (2025). Natural Emergent Misalignment from Reward Hacking in Production RL. arXiv:2511.18397
  11. Lee et al. (2026). Meta-Harness: End-to-End Optimization of Model Harnesses (preprint project page). yoonholee.com/meta-harness
  12. Kimi Team (2026). Kimi K2.5 Tech Blog: Visual Agentic Intelligence. kimi.com/blog/kimi-k2-5
  13. Rush, S. (2026). A technical report on Composer 2. cursor.com/blog/composer-2-technical-report
  14. Chroma (2026). Chroma Context-1: Training a Self-Editing Search Agent. trychroma.com/research/context-1

This article does not authorize any form of reproduction or rewriting for republication. If you find any, please help me report it.

بنقرة واحدة حفظ

استخدم YouMind للقراءة العميقة للمقالات سريعة الانتشار بتقنية الذكاء الاصطناعي

احفظ المصدر، واطرح أسئلة مركزة، ولخص الحجة، وحوّل المقالة واسعة الانتشار إلى ملاحظات قابلة لإعادة الاستخدام في مساحة عمل واحدة تعمل بالذكاء الاصطناعي.

اكتشف YouMind
للمبدعين

حول Markdown إلى مقالة 𝕏 نظيفة

عندما تنشر كتاباتك الطويلة، فإن الصور والجداول وكتل التعليمات البرمجية تجعل تنسيق 𝕏 مؤلمًا. YouMind يحول مسودة Markdown كاملة إلى مقالة نظيفة وجاهزة للنشر 𝕏.

حاول Markdown إلى 𝕏

المزيد من الأنماط لفك التشفير

المقالات الفيروسية الأخيرة

استكشاف المزيد من المقالات الفيروسية