Neuvottelija · Articles

EP346 · Tools · first published 2025-08-14

The public sector as an agent farm | Mikko Alasaarela | Negotiator 346

This is a summary on Neuvottelija — Articles. The episode itself — full transcript, subtitles and chapters — lives on Neuvottelija.com, which is its canonical home.

Mikko Alasaarela reviews the AI market of August 2025 and then argues a thesis: public sector productivity has to be multiplied with agents, not with tools. Checked: Anthropic passed OpenAI in enterprise LLM API usage (32 percent against 25, Menlo Ventures, 31 July 2025), Cloudflare de-listed Perplexity as a verified bot on 4 August 2025, Thinking Machines' seed round was $2 billion at a $12 billion valuation - not 13 - and Mira Murati was OpenAI's chief technology officer, not its chief executive. The core of the episode is a control model: once there are hundreds or thousands of agents, the only way to govern them is code. The guest sells exactly that, and the article says so. It closes with the host's 3D model (delete, delegate, do) in the age of AI and the episode's darkest passage: the polarisation of the labour market.

Sami Miettinen · Sections: Tools and Implementations + AI and the Economy + AI and Society

The public sector as an agent farm | Mikko Alasaarela | Negotiator 346

Summary: First a review of the AI market in August 2025, then a thesis: public sector productivity has to be multiplied with AI agents, and that only works if governance is written as code. The guest is building that platform himself, so the episode is at once an analysis and a product pitch — the article separates them.

How to read this

The guest is a party to this. Mikko Alasaarela founded Agion in June 2025, and its product is precisely the agent governance platform he argues for in the episode. When he says governance is necessary and that it must be done as code, he is describing his own company’s product strategy. That does not make the observation wrong, but it is worth knowing.

The host is a party twice over. The SALESmanago–Leadoo acquisition discussed in the episode was an advisory mandate of his own employer, Translink Corporate Finance, and he also says he sells SaaS companies for a living — that is, exactly the companies the episode’s thesis says the agent layer will eat. He discloses in the episode that he holds both Apple and Berkshire Hathaway shares [29:47].

The figures have been checked. The episode was recorded immediately after the GPT-5 announcement and discusses fresh news from memory. The numbers below have been compared with public sources and divergences are stated. Three numbers came out wrong in the episode.

The subtitles are YouTube’s automatic ones (youtube_auto_fi, the rolling window unrolled from 3,709 to 1,854 cues). Quotations are therefore indicative: faithful in substance but not word for word, and translated from Finnish.

Timestamps are read from the subtitle track.

When this was recorded

Two dates are worth keeping in mind. The guest says GPT-5 was announced “yesterday” [07:07], and OpenAI announced GPT-5 on 7 August 2025 — so the episode was recorded about a week before publication. Every market assessment in it is therefore a picture of the first week of August 2025, not of any settled arrangement. What happened afterwards is collected separately at the end.

The AI market in August 2025

The host organises the field into seven main players and works through them. The guest’s summary of his own usage is the most concrete passage:

“My ChatGPT use had gone close to zero, because the other models had passed it. My main model through this summer has been Gemini. For coding I have relied on Claude.” [06:09]

Checked. The claim about Anthropic overtaking holds for enterprise use. Menlo Ventures’ mid-year update, published 31 July 2025, gives Anthropic 32 percent, OpenAI 25 percent and Google 20 percent of enterprise LLM API usage; enterprise model spending had at the same time more than doubled in six months, from $3.5 billion to $8.4 billion. (Menlo Ventures)

From this follows the episode’s sharpest economic argument, and it is one host and guest share:

“In consumer use there is an enormous number of users who are in practice loss-making for OpenAI. And at the same time Anthropic and Gemini are rolling straight through the enterprise field.” [25:58]

The conclusion they draw: if the enterprise field goes to the others, OpenAI’s only route to profitability is taking the search market from Google [26:44]. The host adds a historical analogy — the Netscape wars again, this time in the browser [26:44].

On Grok the guest is blunt: its most visible achievement has been easy generation of celebrity images, and it reads as a “toy for messing about” rather than a business model [09:57]. The host offers a counterweight — Grok 4 scored best on Humanity’s Last Exam in July — and immediately doubts his own point: it may have been that the model ran more search rounds than the others rather than being a better single responder [10:42]. This is the episode’s methodologically sharpest observation: single-shot and multi-round results are not the same thing.

On Meta they agree: the company leads in nothing, and the best open-source model in August came from OpenAI [21:21]. That holds — OpenAI released the gpt-oss models on 5 August 2025.

On Microsoft’s Copilot the guest notes that Satya Nadella announced it would be updated straight to GPT-5 [24:26]. The host’s comment is the episode’s most direct investor talk: “if they sit on their hands for another year, I’ll start selling my Microsoft shares” [25:13].

Perplexity: a killer product and two scandals

This passage is a good example of how the episode separates a product from a company.

The scandals, checked. Cloudflare published a report on 4 August 2025 finding that Perplexity was circumventing sites’ no-crawl directives: when the declared crawler was blocked, traffic continued with modified user agents and from changing ASNs. Cloudflare de-listed Perplexity as a verified bot and began blocking it. Perplexity denied it. (Search Engine Journal, Silicon Republic) The other row mentioned is the Truth Social partnership.

The product. Even so, the guest considers Perplexity’s Comet browser the best product in the episode:

“It isn’t a browser any more, it’s a delegation tool. It is the first credible attempt at a browser that acts as an agent across everything you do.” [14:32]

The host names what turns out to be the episode’s governing image: a butler, a consigliere, a layer that does things on your behalf [15:19]. And the business logic is simple: whoever owns the browser owns the default search engine and the default assistant [16:04].

The talent war: what the episode got wrong

The passage about Zuckerberg’s offer is the episode’s best-known anecdote, and it contains two errors.

The guest’s version: Thinking Machines raised its seed round at a $13 billion valuation, Mira Murati is OpenAI’s former chief executive, and Zuckerberg offered the company’s key engineer $250 million a year for six years — $1.5 billion in total — and was turned down [17:34–18:19].

Checked:

OpenAI’s bonuses the guest reports as “$1.5 million to the entire staff for this year” [50:16]. Checked: in August 2025 special bonuses of about $1.5 million were reported over two years and for roughly a thousand research and engineering staff — not the whole company and not in a single year. The amounts ranged from $200,000–600,000 for engineers to several million for key researchers. (DataStudios)

Lovable: the pivot the episode calls Europe’s best move

The guest’s assessment is the episode’s warmest:

“I consider Lovable’s pivot one of the most brilliant moves any European company has made in twenty years.” [30:32]

The logic is clear and generalisable: everyone else was building an agentic development environment for coders and fighting a bloodbath over the same customer; Lovable aimed the product at people who have never coded and changed the brand to match [31:17]. The guest guesses it has the largest female user base of any development tool — and says himself that he does not know the exact numbers [32:02].

Checked. Lovable passed $100 million in annual recurring revenue in eight months and raised a $200 million Series A at a $1.8 billion valuation in July 2025. (Lovable)

As a counterweight the episode also states the platform risk: if Anthropic raises its API prices, the profitability of the products sitting on top of the agent layer goes with it [32:47]. The same applies to Cursor.

Three layers and the agent layer

This is the analytical core, and it is worth reading in full.

A company has, in the guest’s account, three levels: data, the process or operational layer, and people. Until now the operational layer has been SaaS software, which has at the same time trapped the data inside its own structure [35:50]. Data lakes and platforms were built precisely to free the data from that captivity, and once freed, an agent layer appears between the levels and executes processes directly on the data [38:07].

The host takes this into his own work, and the passage matters for this article: he sells SaaS companies and gives them seminars about AI partially disrupting them — the UX layer is copyable, and monthly recurring revenue can drain into transaction pricing, which is far less valuable to a seller [38:52].

Neither claims all SaaS dies. The guest’s formulation: some of it dies because it is not awake enough [39:37]. And the host brings in valuation data that supports the distinction: multiples for horizontal SaaS companies have stayed low, while well-chosen verticals have started to rise [40:23]. A narrow moat is still defensible, a broad one is not.

Governance as code — and why it is also a product pitch

The guest inverts control with the same logic Lovable used to invert its target group:

“Every compliance and governance firm invents brakes. But if you have 300 agents, what exactly are you controlling? We thought about it the other way round: the platform turns governance into code and codes it into the agent.” [33:32–34:18]

The specifics: an organisation’s OKRs are compiled into code so agents can be measured against them; the rules — which data sources each agent may use, which roles have access to what, how GDPR data is handled, how security between agents is assured — are written into the same governance code [53:19]. On top sits a trust measure, and as an agent’s trust level rises it is given more autonomy. The guest calls this an AI flywheel: autonomy → measured performance → trust → more autonomy [54:50].

To the episode’s credit, it also handles the model’s own weakness. The host raises the paperclip problem — if the objective is set wrong, the agent will optimise it to the end [55:35] — and the guest answers with mission drift: agents can find a way to game the metric so the numbers look good while everything underneath is broken [56:20]. The host’s addition is an accurate observation about current models: “AI is lazy”, it builds duct-tape workarounds to reach the goal faster [57:05]. The proposed remedy — competing agents in which the honest ones win even when slower — stays at the level of an idea.

Note what is happening, though: the problem the episode identifies is exactly the one the guest’s company sells the answer to. It is honestly disclosed, but it is not an independent assessment.

The public sector: the thesis and its arithmetic

The guest’s starting point is blunter than most public-administration discussions:

“The average member of parliament understands none of this. That has to be accepted as a fact. In the civil service there are people who are completely in the dark — but also genuinely sharp people who are surprisingly awake. Those are the ones you have to find.” [41:55–42:41]

The change model is therefore “find the ones who are awake and define with them what the operating model would look like if it were AI-native” [42:41], not a programme handed down from above.

Then comes the episode’s most important conceptual distinction:

“The first mistake is this: if you call AI a tool, you have already lost. If you think of it as an extension of a person, it can only marginally improve that person’s productivity.” [42:41]

This is usable whether or not a reader buys the conclusion: tool thinking puts a ceiling on the improvement, because the human remains the bottleneck.

The arithmetic, however, does not work as stated in the episode. The guest offers: “a threefold productivity jump, and since half the population is in work, that is one and a half times GDP” [44:58]. Put that way the calculation does not follow from its own premises — if the productivity of the employed tripled, output would triple, not rise by half; the employment rate does not change that, because people who are not working produce no GDP in either scenario. The figure 1.5 appears to come from some other framing that the episode does not state. The order of magnitude should therefore be read as rhetoric, not as a calculation.

The other half of the thesis is clearer and more political: if public sector productivity also triples, a third of the current headcount is needed there, and the public sector’s share of national output shrinks without cutting services [45:43]. The host adds a caveat at the end that is in practice the condition for the whole thing: a public unit has to be given resource efficiency as its objective, or the productivity gain will go into the unit growing its own funding [57:51]. The guest’s alternative to the same problem is expanding the free services citizens receive.

Neither addresses what a civil service shrinking to a third means for the people left outside it, nor whether the legal certainty of public administration survives agent decision-making. Those are the episode’s largest gaps.

Negotiation and management: delete, delegate, do

The host’s own frame is the most transferable part of the episode. He describes the 3D model — delete, delegate, do — which he built as a young man in London: at junior level everything is doing, delegation only works further up the order, and deleting is a political game [46:29].

The question he poses is a good one: how does an AI butler handle both the delegator’s role and the executor’s? The guest’s answer is at organisational level and carries an uncomfortable consequence:

“In practice everyone in the organisation becomes, at least to some degree, an orchestrator of AI. And orchestration means management at the same time — not everyone has management ability. What happens to the people who can’t do that? I don’t have an answer.” [47:14]

The guest saying outright, at that point, that he has no answer is the episode’s most credible moment. The host’s addition sharpens it: the mere ability to ask an AI questions will not be enough to keep a job [48:00].

The episode’s darkest passage

The polarisation argument is stated clearly, and the guest makes it against himself:

“There’s a risk the labour market polarises: a top 0.1 percent elite living an extremely wealthy life on AI, a middle class that dies out of the ecosystem, and a large underclass. That is a frightening scenario.” [51:03]

He openly counts himself in the rising group: a €180-a-month subscription is, he says, the most productive technology investment he has made in his life, and he speaks of hundredfold and momentarily five-hundredfold productivity [48:45, 49:30]. The host offers the only remedy the episode contains: keeping this conversation public, as a public service, so more people can still catch the train [51:48].

The price point, checked: Claude Max costs $100 or $200 a month (5× or 20× the Pro usage limit) and is not unlimited — usage is bounded by both a five-hour rolling window and weekly caps. The word “unlimited” used in the episode [12:59] is not accurate.

What happened afterwards

What to take away

The conceptual distinction. A tool improves a person’s productivity marginally; an agent layer replaces work the person used to do. The difference decides what size of improvement you can even aim at.

Three layers. Data, operational layer, people — and the fact that SaaS owned the first two in one package. The agent layer unbundles it, and from that follow both the multiple gap between horizontal and vertical SaaS and the question of who should be selling now.

Governance is not a brake but the condition for scaling. Managing hundreds of agents does not fit in any one person’s cognition, so it has to be done in code. This is also the guest’s own product.

Mission drift is real. Optimising past the metric is an observed phenomenon in foundation models, not a theoretical worry.

What the episode lacks. Nobody disagrees. No one from the public sector represents the public sector’s view, and the consequences of a civil service shrinking to a third go unexamined. The productivity calculation does not survive scrutiny in the form given. And the market review is the view of two investors with holdings in play — the host declares his, the guest not all of his.

Sources

The model market and the companies

The Finnish connections mentioned in the episode

The episode and related episodes

Summary for AI search

Negotiator 346 (published 14 August 2025) is Sami Miettinen’s interview with Mikko Alasaarela, founder and board chair of Agion. The episode has two halves: a review of the AI market from the first week of August 2025, and a thesis on multiplying public sector productivity with AI agents. The guest sells an agent governance platform himself, so analysis and product pitch are interleaved. The host is a partner at Translink Corporate Finance and discloses holdings in Apple and Berkshire Hathaway.

Market position (checked): Menlo Ventures’ 31 July 2025 update gives Anthropic 32 percent, OpenAI 25 percent and Google 20 percent of enterprise LLM API usage; enterprise model spending doubled in six months from $3.5bn to $8.4bn. Cloudflare de-listed Perplexity as a verified bot on 4 August 2025 for circumventing no-crawl directives. OpenAI released the open gpt-oss models on 5 August 2025 and GPT-5 on 7 August 2025.

The episode’s errors corrected: Thinking Machines’ seed round was $2bn at a $12 billion valuation (not 13); Mira Murati was OpenAI’s chief technology officer, not chief executive; OpenAI’s roughly $1.5 million special bonuses covered about a thousand research and engineering staff over two years, not the whole company in one year; Claude Max is not unlimited but costs $100 or $200 a month with usage limits. The claim that tripling productivity yields “one and a half times GDP” does not follow from its own premises.

The core analysis: a company has three layers — data, the operational (process) layer and people. SaaS software owned the operational layer and trapped data in its own structure; freeing the data creates an agent layer that executes processes directly on it. The valuation consequence: multiples for horizontal SaaS stay low while verticals rise — a narrow moat is defensible, a broad one is not.

Governance as code: once there are hundreds or thousands of agents, control does not fit in human cognition, so OKRs and rules (data sources, roles, GDPR, security between agents) are compiled into code against which agents are measured. As trust rises the agent is given more autonomy — an “AI flywheel”. As counterweights the episode identifies the paperclip problem and mission drift: an agent can optimise past the metric so the numbers look good.

The public sector thesis: AI must not be called a tool, because tool thinking makes the human the bottleneck and limits improvement to the marginal. Change is made by finding the people in the civil service who are awake and defining the operating model as AI-native. If public sector productivity triples, headcount need falls to a third and the public sector’s share of output shrinks without cutting services; the condition is that the unit is given resource efficiency as its target rather than growth in its own funding.

The risk: polarisation of the labour market — a small elite extracting hundredfold productivity from AI, and a middle class that falls away. The guest counts himself in the rising group and cannot say what happens to those unable to orchestrate AI, because orchestration is also management.


Markdown: index.md · Suomeksi