Updated: August 2026.
The AI model race will take on a completely different character starting in 2025 and will accelerate throughout the first half of 2026. If 2023–2024 was a race to see “who has the bigger model and knows more,” then 2025 is a race to see “who can actually get the job done—and do it more cheaply.” Three things have shifted simultaneously: how models generate capabilities, how they generate revenue, and who is leading the pack.
Significant Leaps Since the Start of the Year
- Run-time inference (test-time compute). OpenAI’s o1 → o3 sequence pioneered the “think before you answer” approach. In early 2025, DeepSeek R1—an open-source model—achieved the same feat at a much lower cost: the “DeepSeek moment” shook the market and proved that reasoning capabilities are no longer exclusive. Google, Anthropic, and xAI all have “thinking” modes. This represents a new frontier in capability expansion: applying computational resources during inference, not just during training.
- Executing agents. Models are shifting from “answering” to “doing”—calling tools, browsing the web, performing computer tasks, and, most importantly, writing code. Agent-based programming assistants have made a quantum leap, handling the bulk of the writing, editing, and testing in many real-world projects.
- Multimodal & long context. The model is inherently multimodal (text, images, audio, video), with a context window of millions of tokens, and can understand video and screen content.
- Efficiency and open-source development are catching up. MoE, model distillation, and small models are reaching the performance levels of last year’s large models; open-source families (DeepSeek, Qwen, Llama, Mistral) are narrowing the gap with the top performers, turning what was once cutting-edge into mainstream technology.

Economic breakthroughs are what make the race so intense
The intense competition isn’t due to impressive benchmarks, but because, for the first time, capability is being directly converted into measurable economic value:
- Programming. AI can write real code—boosting the productivity of engineers, the most expensive resource in the software industry—while also generating clear revenue streams (subscriptions, APIs).
- Automation of knowledge work. Agents conduct research, provide customer support, handle operations, and analyze documents. The market it targets is the labor market—a market worth trillions of dollars.
- Computational costs are plummeting. The price per million tokens has dropped sharply year over year, making AI affordable enough to embed in every product. The Jevons Paradox: cheaper ⇒ more usage ⇒ total computational demand increases.
- Platform rewards. Market leaders dominate the economy’s “AI layer”: enterprise contracts, developer ecosystems, and data loops. Therefore, massive investment in chips and data centers makes sense—the expected value generated far exceeds hardware costs.

It is this very flywheel that is the driving force behind the global memory chip crisis: each rotation demands more HBM and GPUs, putting pressure on a supply constrained by physical limits and time.
First Half of 2026: Where Is the Race Headed?
If 2025 is the year that proves the viability of “thinking models” and “doing models,” then the first half of 2026 is the year those models enter actual production—and immediately reveal their limitations:
- Programming agents have become the backbone of work. They are no longer just assistants that suggest command lines; they now handle multi-step tasks: reading code repositories, editing multiple files, running tests, and opening pull requests. The metric has shifted from “solving puzzles” to the rate at which they complete lengthy tasks without human intervention.
- Reliability has become the main battleground. The biggest problem with agents isn’t that they’re “not smart enough,” but that they drift off course over long sequences: veering off target, getting stuck in loops, or failing on the twentieth step. A host of techniques have emerged to address this—self-verification, running in a sandbox, checkpoints for rollback, and supervisors at critical nodes.
- Long-term memory and massive context have become standard features: the assistant remembers projects across multiple sessions, as well as organizational habits and constraints. This is also what consumes the most memory.
- The cost per unit of computing power continues to fall thanks to sparse models (MoE), distillation, quantization, and new hardware—but total spending still rises, just as the Jevons paradox predicts.
- Open-source is closing in. The gap between the best open-source models and the best closed-source models is now measured in months rather than years, driving the continued democratization of the infrastructure layer.
- Constraints are shifting to hardware and power. What will determine who moves faster in 2026 is no longer ideas, but whether there is enough HBM, GPUs, and power—which ties directly into the memory chip crisis.
Map of the Parties (mid-2026)
- OpenAI — focusing on consumer products, reasoning, and agents; making massive bets on infrastructure.
- Google (Gemini) — strong in multimodal and long-context processing, with the advantage of its proprietary TPU chips.
- Anthropic (Claude) — strong in programming and enterprise agents, with a focus on safety.
- xAI (Grok) — Rapid growth, integrated with the X platform, and more lenient moderation.
- Meta (Llama) — open strategy, broad adoption.
- China (DeepSeek, Qwen…) — highly efficient, open-source, closing the gap rapidly, and driving up prices across the entire industry.
Commonality: The gap at the top is narrowing. The difference shifts from a "smarter model" to reliability, cost, and the degree of integration into actual work.
Trends for the Second Half of 2026 and 2027
- A reliable agent for long-running tasks. Moving from “getting it done” to “doing it right and consistently”: reliability, self-verification, and multi-step on-chain error handling.
- Deeper reasoning combined with tool invocation—there’s still plenty of room for test-time compute.
- The underlying layer is rapidly commoditizing as open-source catches up → value (and profit margins) are shifting to the agent, product, data, and infrastructure layers, rather than remaining in the “weighting model.”
- Memory & context. Models have persistent “memory” and extremely long context—extremely memory-hungry, directly tied to the chip shortage.
- Multimodal expansion into the real world—world models, robots, and embodied agents.
- On-device & sovereignty. Small models run directly on the device; “sovereign AI” on a country-by-country basis.
Prediction
- The race is shifting from “raw IQ” to “reliability + cost + integration.” The top models will be similar; success or failure will hinge on the agent and the products built around it.
- Second half of 2026: Reliable agents will prevail. The side that achieves the highest long-term task completion rate—not the highest benchmark score—will secure enterprise contracts.
- Profit margins are shifting away from “weighting.” The open-source layer is catching up quickly; money is flowing into applications, infrastructure, and proprietary data.
- Demand for chips will remain strong through 2026 and into 2027 because memory, long-term context, and runtime reasoning all consume significant hardware resources—just as discussed in the article on the memory chip crisis.
- The risk that expectations will outpace reality could trigger a correction; but the underlying capabilities have now created enough real value that this is not a pure bubble.
In summary: the race has moved beyond the “demonstration” phase and entered the “implementation” phase. The side that transforms its capabilities into a reliable, cost-effective agent that is tightly integrated into the process—rather than simply flaunting benchmark scores—will reap the rewards as the flywheel keeps turning.
Thảo luận