The Jevons Paradox of AI: Why Cheaper Models Create More Work, Not Less

The price of running a model is falling by roughly 50-fold a year. The work around the model is not falling at all.
Table of Contents
Meet your new assistant.

One assistant, every channel, the work handled. Create an account in seconds.

The Jevons paradox says that when a resource gets cheaper to use, people do not use less of it. They find more jobs for it. AI is proving the point right now, and most buyers have not priced what that means.

The cost of AI is dropping fast. Epoch AI tracked six benchmark trends at fixed performance. The price fell between 9-fold and 900-fold a year, with a middle estimate near 50-fold. That figure covers the price of matching a benchmark through an API. It is not the price of work you would ship. The trend is still real, and it is changing how much AI people use.

It is not changing the work that sits around the model. That work is growing.

Most buyers are watching the wrong line on the invoice. The model got cheap. What the model hands back did not.

Cheap tools do not get used less. They get used more.

In 1865, William Stanley Jevons wrote about coal. Britain was building better steam engines, and most people read that as proof the country would burn less fuel. Jevons said the opposite would happen. A cheaper engine widens the field of use, and the wider field burns more coal than the old one ever did. He put it plainly in chapter VII of The Coal Question.

He was writing about energy, not AI, and the borrowing deserves care. Reviews from the UK Energy Research Centre and from Steve Sorrell both treat rebound across a whole economy as hard to test and far from settled. Nobody has proven a Jevons paradox for AI. We are not claiming one.

We do not need the strong version of the claim. We need the weak one, and the weak one is already visible in how people work.

Watch what happens to a single request. Someone asks an assistant for a morning brief. It is a favor at first. Then it runs every day. Then it grows to cover more threads, more files, more meetings, and more replies drafted on their behalf. Nobody sat down and chose to widen it. The price fell, so the ask spread out to fill the room.

Every team we talk to has some version of this story. None of them planned it. That is the point.

Three things scale with every extra attempt, and none of them are the model

Every added attempt pulls three costs along with it. They hide well, because none of them show up on the AI bill.

  • More private data goes in. A better answer needs richer context. Richer context means more of your work, your clients, and your calendar moving through a system you do not own.
  • More candidates land in front of a person. Every draft, summary, and suggested reply has to be read, fixed, kept, or thrown out. That job grows with the pile.
  • More load lands on a device that did not get bigger. Battery, memory, heat, and storage stay as finite as they were last year.

The privacy toggle is not a fix. It is a transfer of work.

The standard answer to the first cost is a toggle and a longer policy page. That is not a fix. It hands the work back to you.

The survey data shows what people do when a vendor hands them that job. In a Cisco survey of more than 2,600 adults across 12 countries, 84% of regular AI users said they worry the data they type in could be shared. Forty-five percent said they hold back private or sensitive facts. Cisco sells security products, so read the framing with that in mind. The behavior is familiar anyway, because most of us have done it.

Holding back works, and it costs you the thing you wanted. Take the private parts out of a brief and you get a brief that is safer because it knows less, and worse for the same reason. An assistant that only knows what you were comfortable typing is a mirror.

That is the trade most people are making without naming it. It is a bad trade, and the vendor built it, not the user.

Generated output is not accepted work

This is the difference the market keeps skipping. It is also the one that decides whether AI pays for itself.

A study of 758 consultants makes the case better than any vendor claim. Inside the frontier the study tested, the people using AI finished 12.2% more tasks, worked 25.1% faster, and produced work rated roughly 30 to 34% higher in quality. On one task placed outside that frontier on purpose, the same people were 19 percentage points less likely to reach the right answer.

Same tool. Same room. Same afternoon. The work was timeboxed and simulated, and the failure came from a single task, so the numbers do not stretch. The shape of the result still holds.

The tool helped and hurt in one session, and the people using it could not always tell which was happening. That is a sorting problem, and sorting is work.

So counting output tells you nothing. Generated output is proof of activity. Accepted work is a different state, and a draft only reaches it when five things are true:

  • It meets a bar somebody stated out loud, before the work started.
  • It guards the context it touched.
  • It respects who was allowed to act.
  • It stays open to inspection.
  • It has an owner willing to answer for it.

That is a working definition, not a metric anybody has tested. It says nothing about how many people adopt a tool or how often a model is right. It does give you something to ask for, which is more than most demos hand you.

Provenance records the trail. It does not check the work.

Three ideas get mixed up here, and keeping them apart is worth the effort. Making a draft is one thing. Tracking where that draft came from and what changed is a second thing. Signing off on it is a third.

The middle one gets oversold. C2PA content credentials attach a tamper-evident record of origin and edits to a file, which is useful. The spec itself then says the credential cannot tell you whether the content is true, accurate, or worth trusting. Knowing where a claim came from is not the same as knowing it is right.

Signing off is the step nobody sells, because it is the step that takes a person. It is also the only one that turns output into work.

The Jevons paradox moves scarcity. It does not delete it.

Turn the 50-fold figure around and it looks different. Benchmark prices move through a data center. A phone does not. Nothing in that number tells you what a model costs in battery drain, heat, memory, or storage on the device somebody is holding.

That is the half of the Jevons paradox people skip. AI gets cheap while hardware, attention, and review time stay as scarce as they ever were. Abundant intelligence still has to fit through a finite machine.

The consumer test is blunt. Does the model drain the battery? Does it heat the laptop? Does it eat the memory, or ask its owner to download, pick, and pin model packs? A product that fails those is not ready for normal people. A product that dodges them by sending everything to a server has not solved the problem either. It moved the problem somewhere the user cannot see it. Local-first hybrid is the honest description of what works. Local-only is a slogan.

Nobody asking for a brief should have to choose a model, budget memory, watch a temperature, or weigh a trip to the cloud against the exposure it costs. Those are real calls with real trade-offs. They belong to the system.

What Pax does about all three

A model is capacity. Capacity is not the same as control, and the gap between them is where a product either exists or does not.

Paciva is AI-native, and Pax is an executive assistant built around these three pressures rather than around a chat box. Here is what that means in the product today, stated at the level the engineering supports and no higher.

Recognized text PII is redacted and sealed through shared provider funnels before a request goes out. Speech-to-text audio and file uploads are the remaining raw exceptions, and we name them rather than fold them in. Suitable work can run through a local inference adapter that targets local servers. Local-first hybrid means what it says. It does not stand in for a promise that nothing ever leaves.

Actions that carry weight stop for human approval. The policy is checked again at the moment they are applied, and the action itself is checked once more right before anything is sent. What Pax remembers about your work sits in a timeline you can read and correct. Past sessions replay the same way every time, so a result can be inspected instead of trusted on faith.

The resource side belongs to the system, not to you. Battery, thermal, memory, VRAM, and model-pack choices are governed by Pax. You never have to pick a model or budget memory to get your work done.

One honest note on all of the above. This work is merged and running in our main branch. No outside party has validated it in production, and we would rather say that here than let you find out later.

The test to run on any AI product, including ours

The price of AI will keep falling. More uses will keep showing up. Jevons is not an argument for slowing that down. It is a design brief, and the brief is simple: safeguards have to scale with use, because use is not going to ask permission.

So stop grading the model and start grading the circuit around it. Three questions do most of the work.

Ask what leaves your control, and what gets checked before it goes. Ask what makes an answer finished, and whose name ends up on it. Then ask what the hundredth run costs, in battery, in bill, and in somebody’s afternoon.

A good vendor will answer all three without flinching. A vendor who cannot has handed you its homework, and you will find that out on the hundredth run instead of the first.

Generated output is proof of activity. Accepted work is the thing you were paying for.

See how Pax works →

Put Pax to work

An executive assistant that carries the working memory of your professional life and finishes the work across the tools you already use.

Transform how your organization operates with Paciva

Transform how your organization operates with Paciva

Product updates, how-tos, community spotlights, and more. Delivered monthly to your inbox.