12 MONTHS, A ROLLERCOASTER RIDE IN THE AI MARKET

How the race for the best model turned into a race for computing power, capital, talent and value creation

02# Value Insider

INTRO

August 2025 to August 2026

Back in August 2025, the AI market was still easy to summarise: OpenAI, Google, Anthropic and others were releasing new models in quick succession, and observers were primarily asking: Which model is currently in the lead? Twelve months on, this perspective has become too narrow.

The model remains important, but an industry has emerged around it whose momentum is just as crucial: hyperscalers are investing hundreds of billions, chip suppliers are becoming investors in their largest customers, and agents are gaining increasing access to tools and corporate systems, raising new security concerns in the process. The scale of this development is hard to overlook: according to the Stanford AI Index, global corporate investment in AI reached around US$581.7 billion in 2025, 130 per cent more than in the previous year. 88 per cent of organisations were already using AI, whereas autonomous agents were in use in only single-digit percentages almost everywhere.

It is precisely this gap that aptly describes the market in the summer of 2026: technology is developing faster than its organisational integration. Competition has long since moved beyond simply vying for the best model; it is now a race for capital, computing power, talent and data. DeepSeek-R1 had laid the groundwork for this in early 2025: cutting-edge performance no longer had to come from the largest US laboratories, nor did it have to be tied to a proprietary model. Efficiency suddenly became just as important as absolute performance.

This is the starting point for our twelve-month window.

THE MODEL DISAPPEARS BEHIND THE SYSTEM

August 2025

5 and 7 August 2025 mark a good starting point: first, OpenAI released gpt-oss-120b and gpt-oss-20b, the first open-weight models since GPT-2; two days later, GPT-5 followed with a contrasting promise: simplicity.

GPT-5 combined a fast model for simple tasks with a model capable of deep reasoning for complex problems. A real-time router decided which model to use depending on the task.

This marks the start of a development that is more important for businesses than any individual model name: the focus of decision-making is shifting from the user to the architecture, with the crucial difference being that the system itself now decides how much computing power to allocate to a task.

The strategic question is therefore less and less often ‘Which model is the best?’, but rather ‘Which model is assigned which task, and at what cost?’

Model selection becomes orchestration.

INFOBOX: FOUR TERMS IN 90 SECONDS

Open Weight: The trained model weights are available and can be run on your own infrastructure – which is not necessarily the same as fully open-source development.

Routing: A system automatically distributes requests across different models to balance performance, speed and cost.

Reasoning: Additional computational steps that enable a model to process complex tasks step by step. This can improve quality, but may also increase latency and costs.

Inference: The day-to-day operation of a trained model. As usage increases, it is not the training but the day-to-day operation that becomes a cost factor.

CAPITAL BECOMES PART OF THE ARCHITECTURE

Autumn 2025 to spring 2026

On 22 September 2025, OpenAI and NVIDIA announced a partnership of historic proportions: OpenAI plans to build AI data centres with a capacity of at least ten gigawatts using NVIDIA systems, whilst NVIDIA, in return, has pledged to invest up to 100 billion US dollars.

Sam Altman summed up the logic behind the agreement in four words: ‘Everything starts with compute.’ This statement captures a key aspect of the past twelve months: progress in modelling has long been a question of industrial capacity as well – data centres, electricity and chips all play a part in determining how quickly a provider can grow.

Even more interesting is how capital and computing power are linked: in November, Anthropic committed to purchasing US$30 billion worth of Azure computing power, whilst Microsoft and NVIDIA, in return, announced investments of up to US$5 billion and US$10 billion respectively in Anthropic.

In February 2026, a similar circular arrangement was struck between Amazon and OpenAI: Amazon invested US$50 billion, whilst OpenAI committed in return to two gigawatts of Trainium capacity on AWS. Anthropic extended its AWS commitment to over US$100 billion in April.

In the essay ‘The Trillion Bet’ in this issue, we explore whether this is already generating sustainable value creation or whether it is primarily the expectation of future demand that is being priced in. The International Energy Agency estimates that investment by just five major technology companies will total more than 400 billion US dollars in 2025, and expects an increase of around 75 per cent in 2026. 

Google expects to invest between 180 and 190 billion US dollars in 2026, with a large proportion of this going towards its own chips. Amazon spoke of around 200 billion US dollars; CEO Andy Jassy raised this forecast to around 220 billion US dollars during the Q2 earnings call.

The economic gamble is therefore clear: providers are building the infrastructure of the AI economy today in the hope that tomorrow’s demand will be great enough.

Circular Vendor financing

Hidden logic: Capital is flowing into model providers – part of this is offset by long-term cloud, chip or computing deals. The demand is real, but not independent: investors, suppliers and customers are often the same players. This makes growth more predictable in the short term, but largely eliminates the market’s normal checks and balances – if one link fails, it drags investors, suppliers and customers down with it in equal measure.

THE CHATBOT IS BECOMING A WORKING SYSTEM

Autumn 2025 to summer 2026

Alongside this wave of investment, the purposes for which computing power is used are changing: in September 2025, Anthropic’s Claude Sonnet 4.5 was explicitly positioned for coding and long-running agents. Other providers followed suit.

A traditional chatbot responds to enquiries; an agent pursues a goal and carries out several steps independently to achieve it. The request ‘Summarise these invoices’ can thus become a process that scans invoices, flags discrepancies and only refers cases requiring a decision to a human. The difference may sound minor, but it is fundamental for businesses: AI moves beyond the text window and into the process itself.

In spring 2026, this shift becomes a product promise: at its I/O conference, Google speaks of the ‘agent-driven era’ and reports more than 3.2 quadrillion tokens per month – seven times as many as in the previous year. OpenAI tailors GPT-5.5 for multi-stage knowledge work, followed in July by GPT-5.6 with ‘more intelligence from every token’.

At I/O 2026, Sundar Pichai expressed a realistic expectation: ‘We’re now at the stage of the AI cycle where people want to see the value in the products they use every day.’ Users increasingly want to see the real added value that the products they use on a daily basis actually bring.

This is precisely where the next stage of development lies: in 2024, the question was whether generative AI would make an impression at all; in 2025, which model would come out on top; and in 2026, the decisive factor will be whether this develops into a reliable, cost-effective workflow.

After all, a model can be brilliant and still fail within a company – due to poor data, a lack of authorisation or excessive costs.

INFOBOX: WHAT IS AN AI AGENT?

An AI agent links a model to a goal, data and tools. It does not necessarily stop once it has produced an answer, but carries out several steps in succession. The more tools and permissions an agent is granted, the greater its potential benefit and its ‘blast radius’ – the damage caused by a security issue. Agents are therefore always a governance issue as well.

‘MORE GPU’ IS NO LONGER THE ONLY ANSWER

Winter 2025 to summer 2026

These billion-euro investments are entering a hardware market that is currently undergoing a period of reorganisation. NVIDIA GPUs remain the backbone of many Frontier systems: flexible for both training and inference, with a mature software ecosystem. With the Rubin platform, unveiled in early 2026, NVIDIA is further extending this lead.

However, universality comes at a price, because the larger the AI operation becomes, the more attractive specialised chips become.

In December 2025, Amazon launched Trainium3, its first 3-nanometre AI chip, which offers around four times the performance per watt of Trainium2. Microsoft followed in January with Maia 200, an accelerator optimised for inference that offers 30 per cent better performance per dollar.

In 2026, Google will split training and inference across two TPU types: the 8t for training and the 8i for cost-effective inference, which boast up to twice the performance per watt of the previous generation.

The reasoning behind this is simple: training is the infrequent process of building a model, whilst inference is its day-to-day operation. If millions of people use a model every day, the total cost of inference becomes more significant than the cost of a single training run.

Amazon is being unusually transparent about these figures: the company puts the annualised turnover of its chip business, comprising Graviton and Trainium, at more than 25 billion US dollars and expects savings running into the tens of billions.

This explains why the hyperscalers are expanding vertically downwards: Google is developing models, cloud services and TPUs; Amazon is developing AWS and Trainium; and Microsoft is developing Azure and Maia – all with the aim of gaining control over the most expensive cost item.

Alongside GPUs and specialised ASICs, another approach is emerging: Cerebras is focusing on wafer-scale systems, in which a large portion of a silicon wafer serves as a computing unit. In early 2026, OpenAI secured 750 megawatts of such inference capacity.

The most likely end state is therefore not a single ‘winner’ chip, but a heterogeneous computing system: GPUs for flexible workloads, ASICs for mass inference.

It is not just the AI that is being orchestrated; its hardware is too.

Who controls which layer of the AI stack?

TALENT IS THE MOST SCARCE RESOURCE

2025 to August 2026

Despite all the billions being spent on chips and data centres, one bottleneck remains surprisingly ‘analogue’: people. Hardly any other sector is as vulnerable to the departure of individual staff members as Frontier AI: crucial knowledge is held by small teams and individual researchers, who could take it to the competition.

Ruoming Pang, former head of Apple’s Foundation Model team, moved to Meta in 2025. According to reports, he was offered a multi-year package worth more than 200 million US dollars. Just seven months later, he moved on to OpenAI.

In June 2026, Noam Shazeer, co-author of the Transformer paper and co-lead on the Gemini models, followed suit: he left Google for OpenAI less than two years after Google had brought him back via the Character.AI deal for a reported sum of around 2.7 billion US dollars.

Just two days later, John Jumper announced his move from Google DeepMind to Anthropic. Jumper is one of the minds behind AlphaFold; he shared the 2024 Nobel Prize in Chemistry with David Baker and Demis Hassabis.

In early August, this became a leadership issue: Google reorganised its AI leadership team, with Demis Hassabis moving into a more research-focused role, whilst Jeff Dean, Sanjay Ghemawat, Oriol Vinyals and Quoc Le left the company to found Discovery Loop. Alphabet’s share price reacted with a fall. This is because computing power can be procured as long as the capital is available. Frontier expertise, on the other hand, is harder to scale. This explains why personnel news in this sector becomes market news.

WHEN BENCHMARKS OBSOLETE FASTER THAN BUDGET CYCLES

February to summer 2026

Der Februar 2026 zeigt besonders deutlich, wie schnell klassische Vergleichslogiken an Grenzen stoßen. On 23 February, OpenAI announced that it would no longer use SWE-bench Verified to evaluate new Frontier models, as an investigation had uncovered problematic test cases and evidence of contamination from training data.

This shows just how difficult it is to measure a market whose systems are changing so rapidly.

Stanford describes the same trend: in March 2026, the leading providers – from Anthropic, xAI and Google to OpenAI and Chinese providers such as Alibaba and DeepSeek – were unusually close together in a preference ranking. Some benchmarks are achieving, within a matter of months, levels that were actually intended to take years to reach.

For companies, this means that a simple league table is of less value: a lead enjoyed today may have vanished by the time the group-wide roll-out is complete.

Evaluation must be brought closer to the reality on the ground: a CFO does not need to know which model wins a coding test, but rather what a correctly completed task actually costs. The price per million tokens is easy to compare, but is often misleading from a business perspective: input, output, reasoning and human rework significantly alter the actual cost of a task.

The more meaningful metric is therefore: cost per task, i.e. the cost per task completed fully and correctly.

INFOBOX: WHY ‘COST PER TASK’ IS BECOMING MORE IMPORTANT

The price per million tokens is only part of the equation: model usage, tool calls, human verification and error costs all factor in. For management, therefore, it is not the cheapest token that is of interest, but the most cost-effective reliable process. 

WHO IS PROFITING FROM THE BOOM AND WHO IS BEARING THE RISK?

February to summer 2026

The economic structure of the AI market can be broadly divided into two groups. On the one hand, there are the frontier labs, which invest enormous sums: At the end of March 2026, OpenAI completed a funding round worth US$122 billion at a valuation of US$852 billion. Anthropic followed at the end of May with US$65 billion, a valuation of US$965 billion and an annualised revenue of more than US$47 billion.

On the other hand, there are companies that sell infrastructure such as cloud services, chips, data centres, storage and energy: they continue to make a profit even when the leading position amongst model providers shifts.

Amazon illustrates this logic well: in 2025, AWS generated revenue of around 128.7 billion US dollars, but an operating profit of 45.6 billion – that is, just 18 per cent of consolidated revenue, but 57 per cent of operating profit.

At the same time, this expansion is putting pressure on cash flow in the short term: Amazon’s free cash flow fell from around 38 to 11 billion US dollars in 2025, and in the second quarter of 2026 it stood at minus 7.6 billion on a 12-month basis – primarily due to investments in AI.

For Frontier Labs, the situation is more challenging: competition is driving down prices per unit of intelligence, whilst the demand for computing power continues to grow. The more interchangeable models become, the harder it is to maintain consistently high margins on the model alone.

This may seem paradoxical: whilst AI is becoming increasingly affordable and useful for users, it is precisely this that makes it harder for individual providers to differentiate themselves commercially. Langfristig könnte deshalb nicht derjenige den größten Wert abschöpfen, der einmal das beste Modell hatte, sondern wer mehrere Ebenen kontrolliert: Distribution, Cloud, Chips, Daten und Kundenbeziehung.

WHO IS GOING TO PAY FOR ALL THIS?

The B2C Paradox

At this point, the AI debate extends beyond the technology sector: almost every investment narrative assumes that AI boosts productivity. However, if automation replaces human labour or reduces the need for it, a second question arises: what happens to income, purchasing power and demand?

Amazon generated revenue of around US$716.9 billion in 2025. Approximately US$588.2 billion – around 82 per cent – came from the consumer-focused segments North America and International. 

Meta is even more dependent on demand from the advertising industry: in 2025, around 196.2 of the total US$201 billion in revenue came from advertising – approximately 97.6 per cent. 

At Alphabet, Google’s share of advertising revenue in 2025 stood at around 294.7 billion US dollars out of a total of 402.8 billion US dollars – a good 73 per cent. This creates a feedback loop that has so far received little attention in the AI debate: the very same companies that are investing billions to make work automatable are dependent on an economy in which people have an income, consume and remain attractive to advertisers.

The distribution of productivity gains will therefore be crucial. If these gains mainly accrue to companies and shareholders, whilst earned income comes under pressure, the productivity surge could turn into a demand problem in the long term. 

Vulnerable business models

Open feedback, no forecast: If automation displaces jobs faster than new ones are created, falling purchasing power may hit precisely those companies that are driving the expansion of AI.

AI AUTONOMY. THREATS INCLUDED

How to limit the potential harm caused by powerful AI agents

The economic boom coincided with a security issue that became very real in July 2026. During an internal cyber assessment, OpenAI models with reduced security restrictions were running in an isolated test environment. They discovered a zero-day vulnerability, gained unrestricted internet access and compromised Hugging Face’s infrastructure in order to access test solutions. Both companies halted the activity.

It is important to put this into context: this was not an ‘escape attempt’ by an AI with a mind of its own, but a narrowly defined test objective for which the models found a technical solution that the environment was actually designed to prevent.

A capable agent does not need to have malicious intentions to cause harm. It is enough for it to pursue a goal relentlessly and interpret a safety limit differently from what is expected.

Anthropic describes this same challenge as the ‘blast radius’: the more access an agent is granted, the greater the theoretical damage if protective mechanisms fail. The company is therefore increasingly relying on technical containment measures: sandboxes, segregated permissions and controlled network access.

A productive agent should only have access to the data and systems it needs to carry out its task. Access credentials should be short-lived, and a failure in the first layer of protection must not immediately expose the entire corporate network.

At the same time, regulation is becoming more specific: the GPAI obligations under the AI Act have been in force since August 2025, and from August 2026 the European Commission will be able to enforce them by imposing fines. Deadlines for high-risk systems have been postponed to December 2027 and August 2028 respectively in the ‘Digital Omnibus on AI’ (not yet published in the Official Journal of the EU at the time of going to press).

The relevant management question is therefore not ‘Is our model secure?’, but ‘How much damage will be caused if our first line of defence fails?’

Productive autonomy with a limited blast radius

QUESTIONS THAT THE BOOM HAS NOT YET ANSWERED

Greater autonomy requires clear boundaries

The past twelve months have shown just how quickly technological boundaries are shifting and how many economic and social questions remain unanswered: Is this still infrastructure development, or is it already a bubble?

The parallels with previous infrastructure booms are obvious, but unlike purely speculative projects, there is already genuine demand for hyperscalers. The next twelve months will show whether it will be large and profitable enough to justify the pace of capacity expansion in the medium and long term. 

At the same time, competition forces providers to deliver more for less, which, whilst good for the customer, poses a problem for profit margins. The market could therefore follow a trajectory familiar from the traditional computer industry: high benefits for users, but value creation concentrated on a few bottlenecks such as computing power, chips and energy. Will agents take over our jobs or break them down?

The sober answer: both are conceivable; the timing and scale remain to be seen. Rather than entire professions disappearing, a reorganisation of tasks is more likely to occur initially: research and analysis will become more automatable, whilst responsibility will remain with humans for longer. And if a great many jobs do disappear: who will remain the customer?

We have already seen this feedback loop in the B2C paradox: the more companies automate their work at the same time, the more it depends on whether the productivity gains translate into new demand. 

The security incident in July 2026 shows that the limitation lies not solely in the model, but in what an agent is technically capable of achieving and how it is stopped if it behaves differently to expected.

And finally: Who controls the infrastructure of intelligence?

Competition is taking place across a wide range of models and applications. However, the key resources underpinning this are concentrated amongst just a few providers: cloud infrastructure, chips and the expertise required for frontier models. Google, Amazon, Microsoft and Meta are therefore increasingly seeking to control several levels of this value chain themselves.

The next twelve months could therefore be shaped less by a single ‘best model’ than by the question of which architectural approach is economically viable. The boom is not over yet, but its most challenging phase may only be just beginning.

For the technological question is increasingly being overtaken by an economic and social one: as we scale up AI, can we also scale up value creation, security and purchasing power in equal measure?

Contents