The AI Slop Report 2026
The recipient pays. AI slop is an unpriced transfer of work, and authorship is the investment that prevents it.
Generation got cheap. Review, context, and repair did not. This report follows one handoff from the sender who saved minutes to the recipient who got the invoice, and proposes the measure that puts the cost back where it belongs.
The sender saved minutes. The recipient got the invoice.
A decision brief lands in your queue. It is fluent, formatted, and long. It does not contain the decision, the evidence, or the next action. Someone saved twenty minutes producing it. You are about to spend two hours finding out.
Start with the one measured signal this field has. In a commercial survey of 1,150 full time US desk workers by BetterUp Labs and the Stanford Social Media Lab, 40 percent reported receiving workslop in the prior month, and respondents estimated nearly two hours to resolve each incident. The same study models the cost at $186 per month per affected employee. Those are self reports and a vendor cost model, not audited time studies, and this report treats them that way. What they establish is direction: recipients are already paying, and nobody is sending them the bill.
Follow the illustrative brief above through the recipient's actual work. First you detect that something is off. Then you interpret what the sender probably meant, reconstruct the context they did not supply, verify the claims they did not check, repair what you can, respond, and quietly update your trust in the sender. Every step is real work. None of it appears in any productivity metric, because generation cost is private and visible while recipient cost is distributed and usually unmeasured. Economics has seen this shape before: a 2012 analysis of email spam estimated nearly $20 billion in annual US costs against roughly $200 million in spammer revenue, a hundred to one externality. That is a historical analogy about cost asymmetry, not an AI estimate, and the mechanism is the point: when producing costs almost nothing, the cost does not disappear. It moves.
The unit that matters is not output produced, or even output accepted. It is realized useful outcome, net of producer, reviewer, defense, recipient, defect, and delay costs. Everything in this report works back from that measure.
Assumptions, stated: burden equals incidents times hours times hourly cost; the exposure share scales individual burden to the team; failed and unresolved handoffs stay in the total. This is a model of self reported time, not an audit, and it excludes relational and downstream costs entirely, so it understates as often as it overstates.
At these settings a 100 person team pays about $7,440 a month in recipient time. The sender's saved minutes never appear on this invoice.
The flood is measurable. Slop is not one number.
Two claims survive the evidence. Synthetic production rose fast on every surface anyone has measured. And nothing in that evidence supports a universal slop percentage.
The production surge is documented across very different methods. Detector estimates put AI attributed posts at 37 percent of Medium and 39 percent of Quora by late 2024, up from around 2 percent in early 2022. A detector ensemble found that the share of primarily AI generated new articles reached roughly half by 2025, then plateaued. NewsGuard's tally of AI content farms grew from 49 sites in April 2023 to 3,749 in its June 23, 2026 update. A manual audit found 5.3 percent of sampled biomedical education videos were both AI generated and low quality. None of these numbers measures the same thing, and that is the trap this section exists to spring.
The texture behind those numbers is worth one paragraph. When Clarkesworld, a science fiction magazine, received roughly 500 machine written submissions in a single February against about 700 legitimate ones, it closed submissions entirely; the queue itself became the casualty, and every honest writer paid the price. When researchers sampled 125 Facebook pages posting AI images at volume, one image drew 40 million views and 1.9 million interactions; the audience was real even where the intent was not. Neither event measures slop prevalence. Both show what a channel looks like when production cost stops gating release, and both explain why recipients now greet fluent work with suspicion by default.
Every study in this set sits somewhere on four separate axes: provenance, quality, deceptive intent, and recipient burden. A detector estimates provenance and nothing else. A discovery tally counts sites that met inclusion criteria, not a quality score. One audit measured quality directly, on ten queries in one domain, as a deliberate lower bound. No study in the set measures all four axes, so no honest reading of it produces a sentence like "half the internet is slop." AI provenance is never proof of slop. The specimen wall on this report's cover is built from recreated genres for exactly that reason: the failure is in the handoff, not the tool.
Slop begins at the handoff.
Private drafts and rejected experiments tax nobody. Failure starts when unfinished judgment enters someone's attention. Here is the definition this report runs on: slop is work whose producer has not invested in making it useful.
Investment means judgment, context, evidence, testing, revision, acceptance rules, monitoring, or accountability. It is not measured in keystrokes or elapsed time, and the producer is the accountable source, a person or an institution, never the language model. The operational version is stricter: slop is a failed handoff that imposes avoidable interpretation, verification, repair, delay, or trust costs because the work was not made useful enough for its recipient, purpose, and level of risk. AI slop is that failure with AI supplying the volume. AI provenance alone is never sufficient; the handoff must also fail the recipient burden test.
A handoff carries an implicit claim: someone considered you, supplied judgment, and will answer for the result. AI makes that signal cheap to imitate, which is why recipients now inherit sincerity checks along with verification work. Bounded studies back the relational half: extensive AI mediation reduced perceived apology authenticity, and an AI label independently reduced how heard people felt, even when the AI written response was rated better. Systems inherit different work: schema, queue, verification, and repair.
Discipline matters when counting. Record the gross burden first, then count only what a proportionate, producer owned control could likely have reduced at lower total system cost. Legitimate disagreement, joint uncertainty, and inherent complexity are ordinary coordination, not slop. Bad output alone does not even prove a lazy producer; recent philosophical work insists on asking who controlled the deadline, the workload, and the release decision before assigning blame. A diagnostic may find burden without finding slop, and that finding is still worth having.
Context slop
Epistemic slop
Relationship slop
Judgment slop
Volume slop
Execution slop
Ecosystem slop
Human slop
A person sends the decision brief without the decision, the evidence, the context, or the next action.
AI slop
AI expands the same missing judgment into fluent volume, and the producer releases it anyway.
Authored AI assistance
The producer defines the outcome, recipient, evidence, acceptance test, and repair owner, then uses AI and assesses the result before sending.
Why rational systems produce irrational amounts of slop.
Cheap generation, volume incentives, and superficial fluency can reward release before work is useful. Weak intent and missing acceptance gates make avoidable recipient burden more likely, not inevitable.
Quantity metrics reward drafts, submissions, messages, tickets, and apparent activity, while cleanup lands elsewhere. This is an analysis, not a finding from any single study, but each component is individually evidenced. One scholarly journal documented a 42 percent submission volume increase after ChatGPT, with abstracts in the highest detector band desk rejected more often than the lowest, a bounded picture of reviewer strain under cheap production. A preregistered experiment with 758 consultants showed the fluency trap directly: AI users finished 12.2 percent more tasks and worked 25.1 percent faster inside the tested capability frontier, and on one task outside it they were 19 percentage points less likely to reach the correct answer. Fluency reads the same on both sides of that line.
The cleanest producer versus recipient separation comes from a randomized study of 680 participants in simulated workplace discussions: AI interventions increased comment length and some participation signals on the producer side, while none improved recipient perceptions. Some reduced perceived quality, and all of them increased dislikes. More was produced; nothing was better received. Power decides where the residue lands, because the person who can not reject a transferred cost absorbs it. That is why this is an incentive design problem before it is a writing problem, and why the fix in Sections 6 and 7 targets accountability rather than word choice.
The recipient tax compounds.
The invoice from Section 1 was one person's. The tax does not stay that small. It moves through four levels, and each level makes the next one more expensive.
At the individual level, it is interruption, interpretation, verification, repair, and delayed decisions. At the relational level, it becomes annoyance, reduced trust, uncertainty about intent, and lower confidence in the sender's judgment; the bounded relational studies in Section 3 show how directly mediation and labeling touch trust. At the institutional level, it shows up as review queues, duplicated work, weaker knowledge bases, and escaped defects, with a plausible but unproven adverse selection risk: if careful contributors withdraw from an overwhelmed channel, low cost producers occupy more of it. At the ecosystem level, synthetic material enters search and future training data. Controlled experiments show recursive training on generated data degrades information unless it retains original human data, and follow up work shows external verification can change that outcome. At the same time, an imperfect verifier imposes its own limits. Neither collapse nor cure is guaranteed; the mechanism is real, but the present tense claim is not available.
Measure the total and its distribution, not the average. An average hides one powerful sender exporting small costs to many recipients. And every recipient who stops trusting a channel builds a private defense system, which taxes legitimate senders too. That is the compounding nobody prices: the defense you were forced to build charges everyone who reaches you honestly.
Individual
Interruption, verification, repair, delay
Relational
Trust discounts on every future handoff from the same sender
Institutional
Queues, duplicated work, weaker knowledge bases, contributor fatigue
Ecosystem
Search pollution and recursive degradation risk in future training data
Authorship is the control system, not keystrokes or detection.
Authorship is not proof of quality. It is the control system that assigns intent, judgment, final approval, and responsibility for repair, so quality can be specified, assessed, and improved.
Adapt the standard medical publishing test, which has carried this weight for decades: substantive contribution to intent and evidence, critical review and revision, final approval for the actual audience, and accountability for accuracy and repair. Three operating states follow. Individual authorship: a person understands, approves, and can repair the artifact. Institutional authorship: routine low risk automation releases work without per item review, because the institution collectively owns the workflow's intent, acceptance rules, monitoring, exceptions, and repair path. Pseudo authorship: approval with no competent role able to comprehend the decision, assess failure, or own repair. That last one is a rubber stamp, and the experimental record warns exactly here: with a simulated 75 percent accurate AI, participants caught wrong answers 8 percent of the time under ordinary explanations, 27 percent with cognitive forcing, and 49 percent with no AI at all. Passive review fails quietly.
Authorship is a gradient, not a purity test. Translation, editing, and language help preserve it; a randomized writing study found AI primary drafting reduced ownership, pride, and perceived accountability more than editing assistance did, which is why the two should never be treated as equivalent. Accessibility is part of authorship, not an exception to it: surveyed disabled students described AI helping them process material and express authentic intent, and authorship must never be confused with linguistic polish or unaided composition. Detection is not authorship either. Current evaluations of sixteen English detectors found inconsistent biases, including against English language learners, while a Czech evaluation found no systematic bias; test your own population and instrument instead of generalizing either result. AI can plainly improve work, with a preregistered experiment showing 40 percent faster task completion and higher quality scores on short writing tasks. The villain is abandoned responsibility, not saved time.
| State | Who holds intent and judgment | Who can repair it | Verdict |
|---|---|---|---|
| Individual authorship | The named producer, end to end | The producer | Authorship kept |
| Institutional authorship | The workflow's owners, through tested rules and monitoring | A named, competent role | Authorship kept |
| Assisted work | The producer, with AI on drafting or language | The producer | Authorship kept |
| Delegated judgment | Nobody engaged with the decision itself | Unclear | Slipping |
| Rubber stamp | An approver who cannot assess failure | Nobody competent | Pseudo authorship |
| Abandonment | Nobody; volume ships itself | The recipient | Slop |
Four phases. Ten gates.
The method is four words: intend, challenge, prove, own. Each phase has gates that hold work back until specific evidence exists. State it plainly: this protocol is a proposed synthesis extending adjacent evidence. It has not been shown to reduce recipient tax, which the diagnostic in Section 9 is designed to test.
The components have support. Cognitive forcing measurably improved independent judgment where passive explanations failed. A structured deliberation system that ran human views through AI synthesis, human critique, AI revision, and human selection produced consensus statements that 5,734 participants preferred to human mediated ones in 56 percent of head to head cases, a genuine worked example of AI inside an authored loop. Early adopting firms have documented staged, risk sensitive oversight tactics. What this method adds on top of that adjacent work is recipient modeling, empirical acceptance criteria, proportional effort, and abstention as a first class outcome.
Two gates keep the method from becoming process slop. Match effort scales review depth with stakes, uncertainty, reversibility, and audience size, because a calendar confirmation and a consequential recommendation should never traverse the same path. Abstain ends the pipeline: if the work is not worth the recipient's attention, do not send it. If the process costs more than the burden it prevents, the proportionality test says reduce the process.
PHASE 1Intend
PHASE 2Challenge
PHASE 3Prove
PHASE 4Own
Two-sided protection.
Sender side authorship cannot govern every handoff that reaches you, because no organization controls every sender. The handoff itself is the shared unit: prevent low investment work before it ships, and protect people from imported work across their connected surfaces.
Widen the word recipient first, because a person is only one kind. A submissions queue is a recipient. A review board is a recipient. A knowledge base, a search index, and a training corpus are recipients, and each inherits a different bill: people inherit sincerity and trust checks; systems inherit schema, verification, and repair work. Section 3's rule holds across all of them, and so does the economics: every unguarded channel eventually builds its own defense, paid for by whoever legitimately depends on it.
Sections 6 and 7 cover the sender side: fewer failed handoffs leave the building. The recipient side is the mirror obligation, and it is where most organizations have nothing but individual willpower. Incoming work should prove relevance, actionability, and acceptable risk before it consumes scarce attention, regardless of the tool or person that produced it. This is an outcome description, not an architecture; this report makes no product performance claim, and the only honest test is the one Section 9 defines: protection earns its place when the burden it removes exceeds the burden it adds for the producer, reviewer, and defender, with defense errors counted against it.
Sender side: prevention
- Authorship gates before release
- Acceptance criteria and assessment
- Proportional review depth
- Abstention when the work is not worth attention
Recipient side: defense
- Incoming work proves relevance before it takes attention
- Actionability and risk checked at the boundary
- Burden routed, batched, or bounced with a reason
- Defense errors measured, not assumed away
Pax is an executive assistant for individuals and a chief of staff to organizations. It exists so the work that reaches people is worth their attention, and the work that leaves them is work they can stand behind.
Measure useful outcomes, not generated output.
The governing metric is total system burden per realized useful outcome, reported with its distribution across senders and recipients. Everything else is a supporting measure.
The denominator is where scorecards go to lie, so guard it first. Predefine the eligible handoffs and the observation window. Count eligible handoffs, realized useful outcomes, unsuccessful cases, and unresolved cases separately, and report exclusions with reasons beside them. Failed and unresolved work stays in total burden; dropping difficult cases never improves the ratio. With zero useful outcomes, report the counts and the burden rather than a finite ratio. Acceptance is not the bar, because exhaustion, hierarchy, and weak alternatives produce compliance without value; test the intended outcome after the handoff. An uncontrolled before and after comparison does not establish causality, so match work where practical and describe the design limits. No universal ROI or target number is attached to any of this, on purpose.
The sequence is small enough to start this month: pick one high volume or high consequence workflow, have the accountable owner name the purpose, the recipient's next action, an observable success condition, and any authority constraints before production, measure a baseline sample with every denominator count, apply proportional authorship and recipient controls to comparable work, compare with the same eligibility rules and window, then expand, revise, or stop, with all three outcomes equally available. The remedy succeeds only if burden removed exceeds producer, reviewer, and defense burden added.
| Measure | The question it answers |
|---|---|
| Time to useful outcome | How long until the intended recipient can use the work for its stated purpose? |
| Recipient rework | How many minutes went to reconstructing, checking, or rewriting? |
| Escaped defects and repair | What failed after release, and what did correction require? |
| Producer and review overhead | How much work did the authorship control itself add? |
| Defense errors | How often did recipient protection delay or block useful work? |
Workflow shape
Observation window
Stakes
Abundance changes what is scarce.
Go back to the brief that opened this report. It was never hard to produce. It was hard to receive.
AI makes production abundant. Judgment, context, and responsibility remain scarce, and abundance makes them more valuable, not less. The competitive advantage on the other side of this shift is not producing more plausible work; it is producing realized useful outcomes without exporting avoidable burden, and being able to show the numbers. Acceptance alone settles neither question. Authorship is how organizations keep the benefits of AI without quietly converting every recipient into an unpaid editor, and the recipient tax is how they find out whether it is working.
One action, then this report is done with you: nominate one costly handoff, the recurring report nobody reads, the brief that always comes back with questions, the queue that eats a team's mornings, and run the bounded baseline and comparison from Section 9 on it. Expand, revise, or stop. All three are wins over not knowing.
Bring one high volume or high consequence workflow. The bounded diagnostic in Section 9 is exactly what a first conversation covers.
Questions readers ask
What is AI slop?
AI slop is a failed handoff whose work was materially generated, transformed, or scaled by AI, and which imposes avoidable interpretation, verification, repair, delay, or trust costs on its recipient. AI provenance alone is not enough; the handoff must also fail the recipient burden test.
What is the recipient tax?
The recipient tax is the portion of a recipient's burden that a proportionate, producer owned upstream control could likely have reduced at lower total system cost. Legitimate judgment, joint uncertainty, and inherent complexity are ordinary coordination, not tax.
Is AI generated content always slop?
No. Human work can be slop, AI assisted work can retain full authorship, and wanted synthetic entertainment that delivers what a willing audience expected is not slop. Provenance and quality are separate axes, and no study in this report's evidence set measures them as one number.
How much does workslop cost?
The one measured signal is a commercial self report: 40 percent of 1,150 surveyed US desk workers received workslop in the prior month, with nearly two hours estimated per incident and a modeled $186 per affected employee per month. The total economic recipient tax has not been measured, and this report does not invent it.
Can AI detectors identify slop?
No. Detectors estimate provenance, not usefulness, and current evaluations show inconsistent biases across systems, including against English language learners in some settings and no systematic bias in others. Detection is not authorship, and a detector score never establishes that work is slop.
What is the fastest way to act on this report?
Nominate one high volume or high consequence handoff and run the bounded diagnostic: define purpose and success before production, measure a baseline, apply proportional authorship and recipient controls, compare with fixed eligibility rules, then expand, revise, or stop.
The evidence base
Twenty eight ledgered claims sit behind this report, each carried with its population, timeframe, and limits. The cited work:
These names identify the published research cited above. They indicate sources, not endorsement, partnership, or any relationship with Paciva. Commercial measurements are labeled as commercial wherever they appear, and detector estimates are never converted into slop prevalence claims.
More research from Paciva
This report is one of six. Keep going with the rest of the series.
Paciva also runs the AI News Monitor, a free live record of how AI gets covered across 13 topics in every language. The Slop lane follows how AI generated slop spreads through daily work as it develops.