The model price war is finally learning to count the thing customers actually buy: completed work. Anthropic released Fable 5.1 with pricing of $10 per million input tokens and $50 per million output tokens, while cutting cache-read pricing to $0.25 per million tokens. The company estimates that typical workloads cost about 25 percent less than Fable 5, with savings approaching 45 percent for some highly agentic workloads.

Those are Anthropic's estimates, and real costs will depend on prompt design, tool use, retry behavior, context size, and how much work can be cached. Still, the direction is right. An agent does not produce value because its token price looks attractive on a slide. It produces value when it completes a process reliably for less than the process is worth.

Long-running agents expose every hidden tax in a model stack. They reread instructions. They call tools. They recover from failures. They pass context between steps. They sometimes burn a small novel of tokens discovering that an API key expired. A modest improvement in model intelligence can lower total cost if it eliminates retries, while a cheaper model can be expensive if humans keep cleaning up after it.

Fable 5.1 is generally available across Anthropic's paid plans, API, and cloud marketplaces. The company is also introducing Mythos 5.1 to a limited group of vetted organizations. Anthropic says sensitive cyber and biology prompts can be routed to less capable systems as a safety fallback. That separation shows how pricing, capability, and access policy are becoming different control planes for the same product family.

For builders, the useful metric is cost per finished job. That means measuring the full loop: model calls, tool calls, infrastructure, latency, human review, and failure recovery. A system that costs eight dollars but completes a valuable task correctly beats a forty-cent workflow that fails often enough to need supervision.

For model providers, agent economics creates a stronger moat than a temporary benchmark lead. Once a company tunes prompts, tools, caching, permissions, and evaluation around a model, switching costs become operational. Lowering the cost of that established loop can retain customers even when another model wins a leaderboard.

The market is growing up. Intelligence still matters, but intelligence without workload economics is a science project with an invoice. The winner will not necessarily offer the cheapest token or the flashiest benchmark. It will make the largest class of useful work dependable enough, fast enough, and cheap enough to run every day.

LaunchPad positionThe model market is moving from benchmark theater toward workload economics. Builders should measure the cost, latency, and failure rate of a completed process, because cheap tokens can still produce expensive work.
Reporting standard

This report draws on the linked primary sources and reputable reporting. Company statements are treated as claims until independently demonstrated.