Field notes · The map for the uncharted · 2 of 6

How many agents will each person have? And what kind?

The bet right now is two. One personal, one for work. You already have one of each. An openclaw in your messages. A Claude Code in your terminal. By next year you'll probably have a third for something you didn't think of yet.

The fight isn't for which model is smartest. It's for which one earns the slot in front of you.

I set up openclaw for several teams last year. The pattern was always the same. Anything front-line got handed to openclaw. Anything that needed to actually do the work got handed to Claude. Two agents. Two slots. Same person.

That's the middle layer. The place where the relationship between AI and humans actually gets decided.

That is why so many agents are trying to enter this category. openclaw, Claude Code, Codex, Character AI. All front-line agents. All of them are making money. They will make a lot more.

Most startups building front-line agents reach users through OpenRouter. Any model can plug in. Every model competes for the same thing. Tokens used per month.

Think about it like the stock market. The stock market trades company shares for dollars. OpenRouter trades model tokens for dollars. Pay dollars, get tokens, use the model. Any model, from any company.

The first stock exchange in the world opened in Amsterdam in 1602. The joint-stock companies needed somewhere to clear the price of risk in uncharted territory. The exchange wasn't an accident. It was the structural piece that made the bet legible.

The agent economy needs the same thing. OpenRouter is the closest thing we have to it today.


There is a contradiction in how the exchange is pricing what it trades. Look at it for a week and you will see it. Token prices are falling. Token prices are also rising. Both are true at the same time.

The cheap tier of models — Haiku 4.5, Gemini Flash-Lite, GPT-5-mini, DeepSeek V4 — keeps getting cheaper. About ten times cheaper per year for the last three years. That is the part of the story everyone repeats.

The frontier tier is doing the opposite. GPT-5.5 is four times the per-token price of GPT-5. Opus 4.7 is more expensive than Opus 4.6. The frontier is getting more expensive, not less. OpenAI is selling three-year forward contracts on this capacity. Anthropic is selling Adobe forty-eight-thousand-dollar unlimited plans because per-token pricing no longer captures the value.

Both things are happening because there is not one model market. There are three.

Call them what the labs themselves call them when you ask. Ferrari. Mercedes. Toyota.

Ferrari is the frontier. The most expensive training run. The highest reasoning. The model the press writes about. There are a handful of these and the price is going up because the cost to train them is rising faster than the cost to serve them is falling. The labs sell forward contracts on Ferrari capacity the way utilities sell forward contracts on electricity. The buyer is not the consumer. The buyer is the application company that needs guaranteed delivery so its product can run.

Mercedes is the workhorse. Reliable, fast, good enough for ninety percent of production traffic. Sonnet 4.6. GPT-5. Gemini 2.5. Priced where most of the actual revenue lives. The lane that funds the labs. The lane the developer integrates with first.

Toyota is the volume tier. Cheap, fast, dumber. The model you call ten thousand times a day in a batch job. Haiku. Flash-Lite. Mini. DeepSeek. Priced near marginal cost because the lab knows the customer's alternative is open weights running on rented GPUs.

The labs do not lead with this framing in keynotes. They lead with the Ferrari because it makes headlines. But the actual product strategy at every American lab — OpenAI, Anthropic, Google — is to sell into all three lanes from one company. Same brand on the Ferrari and the Toyota. Different prices, different margins, different customers.

three lanes, three prices ferrari rising, toyota falling, both at the same time $0.10 $1 $10 $100 $1k $ per million tokens (log) 2023 2024 2025 2026 2027 ferrari opus, gpt-5.5 mercedes sonnet, gpt-5 toyota haiku, deepseek one company sells all three lanes. the chinese labs add a fourth one underneath.

The Chinese labs play the same game with a twist. They play all three lanes and they add a fourth one underneath.

Zhipu sells GLM-4.7 at the Ferrari tier and tells founders, face to face, that the goal is parity with the Western frontier. GLM-4.6 Air at Mercedes. Smaller GLM variants at Toyota. Three lanes, one company, sold to enterprises across China and Asia.

Alibaba does the same thing with Qwen. Qwen3-Max at Ferrari. Qwen3 at Mercedes. Qwen3-Coder and Qwen-VL as specialized vehicles. And underneath all of that, the Qwen open weights — 7B, 14B, 32B, 72B — released free for anyone to download and run. That is the fourth lane. The lab is selling closed at the top and giving away open at the bottom on purpose.

DeepSeek is more or less only the fourth lane. Open weights. Near-marginal-cost hosted inference. They are the floor of the market. Every Western lab's Toyota pricing has a DeepSeek-shaped constraint on it.

Moonshot does Ferrari in a specialized lane — Kimi K2, the long-context machine. Different car, same lane.

Ant Group, MiniMax, StepFun, ByteDance Seedance. Multi-tier, multi-vehicle. Distribution is the play. Whatever the customer can use, they ship.

The Chinese labs are not catching up. They are reshaping the floor on purpose so the ceiling has to defend itself. That is the second contradiction in the price data. The cheap tier is cheap because someone wants it to be cheap.


But all of this is one vehicle. The car. The general-purpose conversational model.

The agent economy is not one market. It is a fleet.

Suno generates music. ElevenLabs synthesizes voice. Cartesia ships sub-hundred-millisecond text-to-speech. Vapi runs real-time voice agents. Different competitive structure. Different pricing. Different scarcity. The Ferrari-Mercedes-Toyota math applies inside the audio market separately. ElevenLabs is the Ferrari of voice. XTTS is the Toyota.

Sora generates video. Runway, Veo, Seedance, Kling. Few suppliers. Very high compute cost. Currently the most Ferrari-shaped market in AI because video is still capacity-constrained at every layer of the stack from training to inference. There is no Toyota of video yet because the science is not done.

Claude Code, Codex, Copilot, Cursor, Devin. Code agents. Integrated into developer workflows. Mixed model dependency. This is the first vehicle where the integration layer captures more value than the model layer underneath it. Cursor is worth more than the model running inside it.

OpenAI ada-3, Cohere, Voyage, BGE, Nomic. Embedding models. Heavily commoditized. All Toyota-tier dynamics across the whole category. The brand premium evaporated in 2024 and nobody has been able to rebuild it since.

Kimi K2, Gemini Pro long-context, Llama 4 with ten-million-token windows. Long-context retrieval vehicles. Different pricing curve. Different unit of work.

GPT-5.5 with vision, Gemini, Claude with attachments. Multimodal. Blurring with the car category but still a separate competitive layer because vision-language data is a separate scaling constraint.

one engine, many vehicles each row is its own market. each market has its own three lanes. vehicle ferrari mercedes toyota open cars opus 4.7 gpt-5.5 gemini pro sonnet 4.6 gpt-5 qwen3-max haiku 4.5 flash-lite deepseek v4 qwen open llama 4 audio elevenlabs suno cartesia hume xtts bark coqui video sora veo, seedance runway kling emerging code claude code codex cursor windsurf copilot aider continue embeddings cohere voyage ada-3 bge nomic long-context kimi k2 gemini pro llama 4 the transformer is the engine. these are the vehicles built on top.

The transformer is the engine. More and more of these vehicles share it. Suno and Claude run on the same fundamental science. The engine block is the same. The vehicle on top is different because the use case demands a different shape.

This is the auto industry pattern at the level of the whole stack. Toyota and Lexus share platforms but sell to different customers. Stellantis makes both Ferrari and Fiat from related engineering organizations. Honda makes cars and motorcycles and lawnmowers and jets. Same engineering discipline. Completely different markets. Completely different competitors.

The AI labs are doing the same thing. OpenAI ships GPT, Sora, Whisper, DALL-E, Codex. Anthropic mostly ships cars but is moving into specialized variants. Google ships everything because Google always ships everything. The Chinese labs ship every vehicle category because the alternative is too small: being only a car company in the country they are operating in.

The next phase of competition is not at the engine layer. The engine is becoming fungible. The next phase is at the vehicle layer, where each category has its own structure, its own tiering, its own scarcity. And the phase after that is at the integration layer above the vehicles. That is where the harness picks which vehicle to use for each step of a workload.

OpenRouter is the routing exchange for cars. There is no OpenRouter yet for audio. There is no OpenRouter yet for video. There will be. Whoever builds those routing layers captures the value migration the same way OpenRouter is capturing it for text.

The exchange we have today is one shore. There is a different shore for every vehicle category. The treasure sits above all of them, where the agent harness decides which exchange to clear through for which step.