Blog

The Harvey Lesson Isn't About Margins: It's About Architecture

Harvey's reported margin swing reveals a broader lesson: models and their economics change. Your organizational context should outlast both.

Guy DerryGuy Derry

Harvey's reported margin swing exposed a constraint that will matter far beyond legal AI: as agents consume more inference, organizations need the freedom to change the models underneath their work without rebuilding the context that makes those models useful.

Harvey, the legal AI company valued at more than $15 billion, reportedly began 2026 with gross margins near 50%. By June, according to Bloomberg's reporting, that figure had fallen to roughly negative 50% after a March update to Harvey's AI agents caused customer usage to surge. Harvey did not confirm the specific margin figures to Bloomberg.

The dramatic number drew attention well beyond legal AI. The architecture behind Harvey's response is more instructive.

Agents Change the Economics of Software

A conventional software action can be cheap enough that another click barely changes the cost of serving a customer. An AI agent behaves differently. Give it a substantive task and it may search documents, call tools, reason over results, retry failed steps, and continue working without another instruction from the user.

Harvey says a single agent run can involve hundreds of model and tool calls over a large corpus. The company describes cost as one of the fastest-growing constraints on deploying agents at scale and says that routing every task to the best frontier model is not economically sustainable.

The mechanism is not specific to Harvey. As organizations move from occasional prompts toward agents doing persistent work, more useful agents can also consume more inference.

For an AI vendor, that pressure appears in gross margin when usage-linked costs rise faster than contract revenue. For an organization using agents, it appears in the cost required to produce an accepted business outcome.

Model Choice Is an Architectural Decision

Harvey's response is instructive. Rather than assume every workload should run through the most capable model available, it built infrastructure that routes different work to different models.

Its own testing helps explain why. Harvey says that across many legal tasks, smaller or open-source models now reach the required quality threshold at substantially lower cost than a top frontier model. Its goal has shifted from finding the best model in absolute terms to finding the model that is capable enough, economical enough, and fast enough for the work at hand. Harvey says this multi-model approach has produced cost reductions of roughly three to five times compared with routing everything through frontier models.

This makes model choice an architectural concern rather than a procurement decision made once. The best option will change as capabilities, prices, and the work itself change.

Harvey also argues that once an agent workforce is deeply tied to one provider, changing models becomes operationally expensive. The work, instructions, accumulated context, and surrounding processes may all have to move with it.

Most organizations will never need infrastructure as sophisticated as Harvey's. They still face the same question: how difficult will it be to change the model underneath a workflow when its capabilities or economics change?

Context Creates Another Dependency

Changing models is only useful if the context that makes them effective can move with the work.

An AI system needs more than general intelligence to operate inside a company. It needs the company's policies, decisions, processes, research, terminology, constraints, and accumulated reasoning. Increasingly, AI products retain some of this through persistent memory, project instructions, and retrieval systems.

Persistence solves one problem while raising another. If organizational knowledge accumulates separately across assistants, vertical AI products, and internally developed agents, the company can gradually create multiple machine-readable versions of itself. Each may be useful, but each must also remain accurate as the organization changes.

The architectural question is therefore not whether AI systems can remember. It is whether the organization's important context exists independently of whichever system currently remembers it. A pricing decision should not become authoritative because one agent retained it. A process should not have to be reconstructed because the team that documented it has moved to a different model. The durable asset is the underlying organizational context. The technology using it should remain replaceable.

Your context should outlive the models that use it.

AI Puts the Context Tax on a Meter

Organizations already spend money reconstructing information they possess. An employee searches several systems for the latest policy. A new hire reconstructs why a decision was made. A team repeats research because the earlier reasoning cannot be found. Kernel calls this accumulated cost the Context Tax.

Agents incur the same tax computationally. When an agent searches too broadly, reconciles conflicting sources, or retries work because critical context was missing, the organization pays for that ambiguity in model processing and human review.

This does not mean large context windows or expensive reasoning are inherently wasteful. Some work genuinely requires them. The useful distinction is between computation required by the work itself and computation required to reconstruct organizational context that could already have been resolved. The relevant measure is not the price of a single model call. It is the total retrieval, processing, retrying, and review required to produce an accepted outcome.

As agent activity grows, that distinction starts to matter financially. Organizational ambiguity that once appeared as employee friction now appears in the cost of machine work. AI doesn't eliminate the Context Tax. Increasingly, it puts part of that tax on a meter.

Own the Layer That Compounds

Harvey's experience is not an argument against frontier models. Harvey continues to use them, and its architecture is designed to route work to them when they are the right choice. The lesson is that the organization should retain the ability to make that choice as economics and capabilities change.

Organizations have a related choice about context. They can repeatedly encode how the company works inside each AI system they adopt, or they can make that context durable and portable enough that different systems can consume it. This is where the economic lesson from Harvey meets the work Kernel describes as Operational Architecture: separating the asset that should compound from the technology that will keep changing. When context belongs to the organization, it can adopt better or more economical tools without rebuilding the foundation beneath the work.

Own what compounds. Stay flexible on what changes.