--- yoast: focus_keyphrase: "state of agents" related_keyphrases: - "AI agents 2026" - "agent comparison" - "AI agent management" - "delegating to AI agents" seo_title: "The State of Agents 2026: What Actually Changed %%sep%% Paciva" seo_title_chars: 56 slug: "state-of-agents-2026" meta_description: "A review of 20 agent offerings across 200 evidence cells found only 7 documented expansions in what agents are trusted to do. Here is what changed in 2026 and the human work that useful delegation still requires." meta_description_chars: 212 cornerstone: true cornerstone_reason: "Standing category theme; future State of Agents editions will link back to this page." social_image: "https://paciva.ai/resources/state-of-agents-2026/og.png?v=5" facebook_title: "Agents are getting better. What makes the work good enough?" facebook_description: "Twenty agent offerings, the changes that matter, and the human work behind useful delegation." twitter_title: "The State of Agents, 2026" twitter_description: "A review of 20 agent offerings across 200 evidence cells found only 7 documented expansions in what agents are trusted to do. Here is what changed in 2026 and the human work that useful delegation still requires." twitter_card: "summary_large_image" page_type: "Web Page" article_type: "Report" author_note: "Paciva" publisher_note: "Paciva" index: true follow: true canonical: "https://paciva.ai/resources/state-of-agents-2026/" breadcrumbs_title: "The State of Agents 2026" target_search_terms: - "state of AI agents 2026" - "AI agent comparison 2026" - "best AI agents for business" - "AI agent management burden" - "GPT-6 Astra benchmarks" - "AI agent review capacity" - "cost per accepted AI result" - "AI agent data custody" - "delegation to AI agents" measured_analysis: body_words: 4499 sentences: 269 avg_sentence_words: 16.7 flesch_reading_ease: 45.0 long_sentences_over_20_words_pct: 26 passive_voice_pct_heuristic: 8 paragraphs_over_150_words: 0 keyphrase_in_title: true keyphrase_in_first_paragraph: false keyphrase_density_note: "the exact phrase 'state of agents' appears in chrome and headings rather than body prose; 'agent' appears 30 times" predicted_yoast: green: ["keyphrase in SEO title", "keyphrase in slug", "meta description length", "text length", "internal links", "outbound links", "images with alt text"] orange: ["keyphrase density in body prose", "passive voice near 10 percent", "sentence length share over 20 words", "Flesch below 60: this is a technical research report and 45 is the measured value, not a target miss to fix with shorter words"] before_publish_human_checklist: - "Confirm the final slug; every canonical, OG and UTM link is built on https://paciva.ai/resources/state-of-agents-2026/" - "Upload og.png and og@2x.png beside index.html and confirm both URLs return 200" - "Re-scrape on LinkedIn Post Inspector, Facebook Sharing Debugger and X card validator after upload" - "Re-verify mutable AS05 and AS06 product pages if publication slips more than 48 hours" - "Resolve the X handle: this page uses @paciva_ai while the site footer links x.com/paciva" --- # Agents are getting better. What makes the work good enough? **The State of Agents · Q1 to September 2026.** Twenty agent offerings, the changes that matter, and the human work behind useful delegation. 20 offerings reviewed · 200 observable cells · 7 documented expansions · 0 outcome cells. Evidence cutoff September 4, 2026; Astra sources verified September 5, 2026; method version 2.1. ## The gain is real. So is the gap. In early September, Artificial Analysis ran GPT-6 Astra against a document task where every rubric criterion has to pass. Astra cleared all of them 33.2 percent of the time. Its predecessor, GPT-5.6 Sol, cleared them 28.2 percent of the time. That five-point gain is real, and someone other than the vendor measured it. It also means that on roughly two-thirds of those documents, something still failed. The same shape shows up in the products. Between January and early September, Glean added collaborative editing to its agents, goose added handoffs a person can resume, and OpenClaw added updater rollback with a diagnostic pass after updates. Real changes to what the software does. None of them measures whether the work got better, or what the people around the work now carry. Three things blur together and are worth keeping apart: what a model scores on a specified evaluation, what a product documents it can now do, and what happens to the work in an actual organization. The central tradeoff runs beneath it all. More generated work can remove drafting and reconstruction. It can also add briefing, checking, correcting, and coordinating. The net effect is a question to measure. A finished file is not necessarily ready for review, approval, or use. If you work alone, pick one recurring job you can judge yourself and count your own review and upkeep time inside the result. If you run a small business or a startup, give one person the whole workflow, including the parts exported to sales, finance, delivery, and customers. If you buy for an enterprise, evaluate the product with its data paths, permissions, review capacity, recovery, and procurement terms, because the product alone is not what you are deploying. Twenty offerings are covered, frozen as exact routes at a September 4, 2026 cutoff. Astra sources were refreshed September 5. ## A running example, and it is fiction One example runs through every section, and it is fiction. An agent prepares a customer proposal from call notes, the current price list, approved contract terms, and delivery capacity. Sales wants it persuasive. Finance needs the margin floor held. Delivery needs a schedule it can meet. The customer needs commitments that hold. There is no single golden proposal, which is the point. The arithmetic and the approved terms can be checked against a source. Tone, emphasis, and which tradeoff to offer require judgment, and prices and capacity can both change after the draft is written. Nothing below reports a trial of this workflow, because no trial was run. ## Choose the job before the product. For the purposes of this report, an agent is a system that takes a goal, chooses and executes tool actions, observes what came back, and adapts while keeping track of the task. That separates it from an assistant that answers questions and from automation that follows a fixed path. Durable memory, bounded permissions, and recovery are things to buy on, not tests a system has to pass before the word applies. Three things get sold as one and are not one: the model, the executor that takes the actions, and the workspace where people and agents coordinate. Astra is a model used inside an offering, not an offering. Buzz is a collaboration overlay and belongs entirely outside the executor comparison. The table below is the inventory as documented: twenty offerings, the work each is built for, how it deploys, the operating work it asks of you, and the shape of its bill. It carries no ranking or score, and the order reflects the study's stable display order. Read it to shortlist, not to choose. Shortlist on the job you actually have, the applications you already run, the technical capacity you can staff, what your data requires, and whether you can tell good output from bad in that domain. Two warnings travel with the table. A free license does not remove the operating bill. And the length of an upkeep list is not a burden score, because a long list of small tasks can cost less than a short list of hard ones. The categories in that table do real work. Managed generalists arrive with the route decided for you. Managed workflow platforms assume you have a process to encode and people to encode it. Software executors are aimed at repositories. Local runtimes and self-hosted systems hand you control and the operating burden in the same motion. A coding executor might build the proposal workflow perfectly well and still be the wrong interface for the salesperson who uses it every day. Company size is not the same variable as company stage, and neither settles this. A technical solo founder can run software a small service business would need help maintaining, and enterprise integration capacity does not remove the requirement to review the work or own the outcome. Regulation and data sovereignty cut across all of it. One correction is load-bearing. OpenAI introduced Workspace Agents on April 22, 2026, not at Astra's September launch. The Work and Codex documentation verifies Astra separately. Do not carry Astra's model identity, admin defaults, or benchmark results across to Workspace Agents, and treat the current rollout wording on the launch page as unsettled, because that page mixes updated availability language with older preview text. ## Seven documented changes, and what the count does not mean. The study fixed twenty routes against ten observables, which produces two hundred cells. Seven of those cells document an expansion between Q1 and the September cutoff, with exact dates on both ends and a matched claim. They land in four observables, not ten. Collaboration accounts for three. Glean moved from a shared agent library to multiplayer agents with collaborative editing. goose moved from subagent logs in the interface to concurrent notifications and a resumable handoff. Dify moved from restrictions inside invalid containers to forms that work inside loops and survive a refresh. Persistent work accounts for two: OpenHands moved from a retained sandbox to a local backend that remembers its mode and scopes a directory, and OpenClaw moved from a shared task ledger to per-agent directories and up to one hundred managed worktrees. Intervention accounts for one, where Hermes moved from reporting completion status to letting an owner list, steer, and stop delegations and see cost per delegation. Maintenance accounts for one, where OpenClaw added updater rollback and a diagnostic triage pass after updates. Here is what that seven does not say. It does not say seven of two hundred agents improved. It does not say product quality improved. It does not describe how common any of this is in the market. The number measures how much of this evidence the research could document under its own rules, and nothing else. ### What Astra changes Astra's document result is the clearest independent gain: the all criteria pass rate rises from 28.2 to 33.2 percent against Sol. Around it the independent picture is mixed. At launch, Artificial Analysis measured AA-Briefcase roughly 80 Elo higher and GDPval-AA v2 roughly 80 Elo lower, with presentation quality declining and no numeric delta given, and those two Elo figures do not sit on a common scale. The coding index rose about two points to 67 at roughly unchanged task cost in Codex at maximum effort, while general evaluation cost rose 75 percent. That 75 percent belongs to the v4.1.1 composition and must not be carried onto v4.2, which changed both the test composition and the Elo anchors, so v4.2 putting Astra four points above Sol is not a continuation of the v4.1.1 result. ARC Prize ran the same model through different execution setups and got different answers. At maximum effort the standard harness scored 62.7 percent for $26,098, while the provider adapter scored 98.6 percent for $17,332. A separate best-observed adapter run at high effort reached 99.9 percent for $18,817 at a different effort setting, so it is not a like-for-like comparison. The execution system moves the result. That establishes system sensitivity, not enterprise reliability, and not a privacy or memory advantage. OpenAI's own evaluations report execution gains from Sol to Astra: AutomationBench 18.1 to 41.4, OSWorld 2.0 65.7 to 72.6 on an offline partial score, and Terminal-Bench 4.0 37.3 to 57.9. These are vendor-reported, do not share a single definition of success, and none measures management savings. The search found no independent Astra field study measuring total human handling, downstream rework or full cost of ownership. That is a gap in the record, not evidence that those outcomes cannot improve. And one note keeps the timeline honest: June releases are visible in the September snapshot rather than being Q3 events, and Astra against Sol is a September comparison against a predecessor, not a measured movement from Q1 to Q3. ## Quality is a contract, not a score. Go back to the proposal. Quality there means work that named people can accept and use under the requirements in force right now. A golden reference answer is one way to evaluate that, and it is not available for most knowledge work. Three kinds of content sit inside the same document and need different handling. Checkable facts: the prices, the arithmetic, the approved terms. Negotiable preferences: tone, emphasis, which alternatives to present. Reserved decisions: discount exceptions, capacity commitments, legal concessions, permission to send. Name who decides each one, keep dissent visible rather than averaging it away, and stop consequential commitments when the people with authority have not resolved the conflict. An average rating cannot erase a prohibited term. A dated note from September 6 sharpens the rubric without adding a result. OpenAI's chief scientist draws a line between achieving a stated goal and exercising sound judgment when objectives are unfamiliar or in conflict, and observes faster progress on the capabilities easiest to measure. For a buyer that means judging the deliverable and the conduct that produced it. An accurate proposal can still fail if the agent exceeded its authority. ### Name the next step, not the quality Work arrives in one of three states. The label names what the recipient is being asked to contribute, not a mandatory stage it had to pass through. Early collaboration is valuable when people know they are being asked to help resolve uncertainty. The failure is presenting exploratory work as decision ready. Recipients can flag missing evidence or ask for review; they cannot veto every preference. Constraints that cannot be waived stay blocking, and the standard scales with the consequence. ### Finished is not the same as inspectable A practitioner transcript circulated during this work is worth using as an illustration and not as a ranking. It describes a low effort workbook produced without dedicated sources or checks sheets, a higher effort run that surfaced more decision relevant questions, and a third workbook that was easier to inspect. The original files, their correctness, their cost and their repeatability were not independently checked here. The buyer question underneath it is concrete. Can another person trace an important assertion back to its source, tell evidence apart from assumption, reproduce a calculation, see which checks ran and what those checks could not cover, and identify what would change the recommendation? A sources sheet proves none of that on its own, and neither does a PASS label. More sheets, more citations, more tokens and more polish do not establish better reasoning. The contrast with software is less clean than it first appears. In one review of algorithmic against holistic evaluation, maintainer written tests credited partial success on work that human review rejected, and none of fifteen manually reviewed pull requests was mergeable without further work. Deterministic tests can miss what acceptance requires. That is a failure mode, not a rate to carry to other models, products or tasks. ### Onboarding, and what feedback actually changes Write the packet before the first task: the purpose, the current sources and who owns them, examples of good and bad, the decision limits, the exceptions, the tools and permissions, who to ask, how corrections get recorded, when instructions expire, and how to stop, recover and retire the workflow. Instructions and memory are not training. Connecting an application does not tell the system which source outranks another. Human onboarding costs time too, so the comparison that matters is incremental work against the previous process rather than against zero. Make the packet relational and not merely documentary: who knows what, whose approval matters, which commitments exceed authority, and when to ask instead of proceed. Access to the company file store is not evidence of any of that. Sort incoming feedback into four buckets: a factual correction, an existing requirement, a proposed preference, or an authorized decision. An owner can adopt a preference as a scoped requirement. The most recent comment confers no authority by itself, and an exception needs both an authorized decision and a rule that permits exceptions at all. Follow one correction all the way through. Finance objects to a discount. That becomes an authorized rule update, the next proposal gets checked against it, and then it gets tested again after a later price change. Record where the correction came from, who authorized it, what it covers, when it expires and what it changed. The research here informs the design and does not prove an effect on agents. A meta analysis of 607 feedback effects across 23,663 observations found an average improvement of d = .41, and more than a third of feedback interventions reduced performance. In a separate study of 1,401 participants, human and AI feedback loops amplified bias in ways participants did not recognize. Feedback is a mechanism that moves in either direction, so treat a feedback loop as something to test rather than something to install. Fixing this proposal is not the same as improving the next proposal. Show the authorized correction reaching the rule or the scoped memory, tell the affected people it changed, and check later work for the same defect returning. Track repair time separately from evidence that the defect stopped. ### Standards move, so recheck against them Two older studies are useful as context, with their era and setting stated. In a 2026 field experiment, 758 consultants completed 12.2 percent more tasks 25.1 percent faster with higher quality on eighteen tasks inside the capability frontier, and scored 19 percentage points lower on one task outside it. In a study of 5,172 customer support agents across 133 teams and 3,006,395 chats, resolutions per hour rose 0.30, which is 15.2 percent, with the lowest skilled quintile gaining 0.5 per hour or 36 percent while the top skilled showed no significant gain and small significant declines in resolution rate and satisfaction. Neither study measures anything about Q1 to September 2026, and both used older models. Combine stable test tasks with samples of current work, and recheck after any meaningful change to inputs, requirements, model, tools or permissions. ## Count the work everyone carries. ### Premature circulation exports unfinished work Sales circulates a polished draft. Finance reconstructs the assumptions underneath it because they were never stated. Delivery challenges a promise it cannot meet. The approver cannot tell which objections are still open. A revised copy arrives and the discussion starts over. People who never generated the draft are now carrying its unfinished work. This is an illustrative failure pattern rather than a measured agent effect, and human authors produce it just as reliably. Start a small log rather than a measurement program. Record the result and its readiness state, the total human handling across every role it touched, and the principal defect or rework. Count incidental recipients, failed attempts and the time spent measuring. Keep delay and displaced attention separate from labor minutes instead of converting them into money. Management tax is all human handling with the agent minus comparable human handling in the process it replaced. It can be positive or negative. Positive does not by itself mean waste, because that difference includes productive judgment and necessary governance. With no matched baseline, describe the workload and claim no savings. Classify preventable reconstruction separately, because that is the part worth attacking. Now run the better version. The agent packages its sources, flags what it is uncertain about, routes the precise discount question to the person who owns that decision, and notifies the people whose commitments changed. Compare the reconstruction and follow along effort against the first version, and do not assume the saving. The target is preventable reconstruction, not participation. Work moved between teams inside the same organization is cost displacement and not automatically an externality. ### Coordination gets removed as well as added The same agent can answer a routine question from approved material before it ever reaches finance, and escalate when authority or uncertainty is genuinely unresolved. That displaces support and clarification work while other tasks add review, and both directions have to be measured. OpenAI reports reduced activity in an internal technical support channel and lower attendance at some office hours. That is a different setting from the advertising experiment discussed below, and fewer messages do not establish time saved or that every support need was met. ### Three scenarios, with the assumptions on the face of them At one dollar per generated attempt and sixty dollars an hour for review, twenty review minutes per attempt costs $42 per accepted result at 50 percent acceptance, $26.25 at 80 percent and $21 at 100 percent. Review dominates generation across that whole range. One reviewer with 120 minutes a day clears 24 outputs at five minutes each, twelve at ten minutes and six at twenty, so generation above that line becomes a queue rather than throughput. And at ten dollars per attempt, moving from full acceptance to half doubles the cost per accepted result from $10 to $20, with 40 percent acceptance taking it to $25 before any downstream loss. These are scenarios, not vendor measurements. A complete cost of ownership also carries setup, maintenance, integration, failure recovery, the recipient's work, switching and exit. A twelve month projection is not measured annual performance, and zero accepted results leaves the ratio undefined rather than large. ### More agent work is not the same as cheaper finished work Five numbers get collapsed into one and should not be: the unit inference price, the resource cost per attempt, the number of attempts consumed, total spend, and the fully loaded cost per accepted outcome. Cheaper attempts may encourage more experiments or larger assignments, which is a plausible mechanism rather than a measured price response. More usage can exhaust a subscription allowance without changing this month's bill at all. Cash paid, allocated subscription cost and API price equivalent usage are three different accounting bases and should never be summed. Tie retries, abandoned runs and parallel candidates back to the deliverable they belong to. Five candidates that produce one accepted proposal are one outcome, not five. Match costs and outcomes to the same cohort of work, because this week's bill and this week's acceptances often describe different jobs. Pair the capacity picture with queue age, active reviewer time, rework and unique accepted throughput. Repeated approvals of the same artifact are not new deliverables, and waiting is not labor. ### Match the effort to the stage, not to the prestige of the model Inspecting an artifact and overseeing behavior are different jobs. Sources and checks help a reviewer read a document. They cannot certify that an action the agent took was authorized. Count the work of checking actions separately from the work of checking outputs. Exploring cheaply, then going deeper only where the uncertainty is consequential, then preparing the work for its audience, is a workflow hypothesis worth testing rather than a prescribed sequence. Record the model, the effort setting, the tools, the checks, elapsed time, total inference spend and the human handoff work, then test whether the extra effort improved the decision or reduced review enough to pay for itself. More effort can still be wasted effort. And running two models to compare their answers is not independent verification, because shared sources and shared assumptions reproduce the same error. ### Shared visibility is not reduced coordination Grok Bot's persistent environment may reduce repeated briefing. Set that against credential management, memory upkeep, monitoring and teardown. For Buzz, look at whether explanation gets duplicated, how notifications land, and what reconciliation costs, holding the underlying agents and task constant. The question to test is whether the workspace delivers a relevant change and a precise question to the right person, or requires everyone to follow the whole thread. Measure notification load, participants per decision, repeated explanations and lost objections, and include the positive case, because agents can summarize changes, route questions, preserve dissent and track commitments. One experiment is directly on point. Across 2,234 participants producing 11,024 advertisements, working with AI agents raised ads per worker by 50 percent, raised task oriented messages by 25 percent, and raised delegation by 17 percent, while image quality and diversity suffered. Output and communication rose together. Messages are not minutes, and this is an advertising task rather than a proposal workflow. A survey of production deployments adds context with its denominators attached: across twenty interview cases and 306 raw responses covering 86 deployed or piloted systems, 41 of 60 relevant cases ran ten steps or fewer before a human intervened, and 23 of 31 used human in the loop evaluation. Continued human involvement is the norm in the systems that were studied. ## Custody belongs to the route, not the label. Follow the customer details, the prices and the negotiated terms through three arrangements. Vendor operation, where the provider runs everything. Local execution with cloud inference, where your machine drives the work and a remote model does the thinking. Local execution with local inference, where both stay with you. Dedicated hosting and restricted egress configurations sit between these and change the answer again. Compare the arrangements on patching, secrets, identity, backups, monitoring, support, recovery, export and deletion, because those are the costs that arrive after the decision. And note the one that catches people: local execution can still send data to models and tools. Local capability and practical locality are different things, and neither one means private. Privacy and intellectual property need separate answers. On privacy, ask about training use, retention, support access, subprocessors, residency and key control. On intellectual property, ask about rights to inputs and outputs, code and weight licenses, whether state can be exported, and what switching actually costs. Contract promises and observed behavior answer different questions, and a hosting location is not a custody finding. The practical version is a walk, not a questionnaire. Take one real document through the route you are buying and write down every place it comes to rest: the prompt and its attachments, the model provider's logs, the tool the agent called and whatever it retained, the trace layer, the backup, the support ticket that will quote it back six weeks from now. Every resting place is a separate retention, residency and deletion question. The cost picture has the same shape. A unit price, total consumption and the cost of an accepted outcome are three different meters, and a cheaper meter need not produce a smaller bill or cheaper finished work. Provider incentives run through seats, credits, token and tool usage, prerequisite services and switching costs. That is worth understanding without the blanket claims that providers only sell tokens or that every hyperscaler is unprofitable. Independent evaluations help here when their dates and execution conditions are stated, and some of them evaluate a model inside a named executor rather than the model alone. None of them certifies your workflow. The ARC result establishes that execution configuration matters. It establishes nothing about custody. Two limits on this section are worth stating plainly. Every custody route in the study is recorded as a documented route shell with the decisive custody fields not publicly demonstrated. Every authority and recovery row is recorded as not sufficient for a recommendation. Not publicly demonstrated means the evidence was not published, not that the capability is absent. ## Function, authority and reach are three separate axes. Return to the proposal one more time. An agent can produce, coordinate, allocate and handle exceptions, or support leadership, and each of those functions demands its own evidence before you delegate it. These functions do not form a ladder and they are not a promotion track. Authority and reach are separate axes crossing them: authority is what the agent may do without asking, running from reading to drafting to reversible action to consequential action, and reach is how far the effect travels, from a personal task to a team workflow to a cross functional portfolio. An agent can produce with high authority over a personal task and coordinate with almost none across a portfolio. Those are different risk positions and they need different evidence. Authority and reach raise the possible consequences of a mistake. They are risk dimensions, not achievement scores, and human accountability for consequential commitments stays explicit at every level. Describe progress as a larger demonstrated scope of useful work: corrections persist where they should, exceptions get handled, reviewers keep up, and results survive changing inputs. This report does not predict a date for autonomous managers, and no evidence here supports one. One founder may hold every human role in this picture. A small business splits them across a few people. An enterprise distributes them widely. A shared agent interface does not move any of those decision rights. Domain expertise includes knowing when work is ready, noticing what evidence is missing, and knowing whose input matters. If agents absorb junior production tasks, the organization needs an answer for how it will develop the reviewers it will need in five years. That is a workforce design question to plan for, not an established claim about deskilling. ## Adopt in proportion to the decision. Three situations call for three different next steps, and each one earns a different size of conclusion. The rule underneath them is that the evidence has to match the consequence, not the enthusiasm. Three to five runs will expose setup and measurement problems, which is useful. They cannot establish broad superiority. Formal confirmation belongs wherever the consequences require that assurance, and it is not a prerequisite for reversible personal use. Claims about improved quality or reduced human handling need a baseline from the current process. Claims about expert performance need an expert comparator. Expert parity is not a universal adoption gate. Expansion can chase quality, responsiveness, accessibility or control just as legitimately as savings. Name the benefit you are buying and the cost you will accept. Require evidence for a claimed cost advantage rather than insisting every adoption has to be cheaper. Stop or narrow the scope when errors escape review, when queues exceed capacity, when sources become unreliable, when recovery fails, or when recipients inherit rework they did not sign up for. Recheck after any material change. Rising activity with flat accepted throughput calls for diagnosis rather than an automatic verdict, because exploration or genuinely higher quality can justify the effort. Before adding agents or reviewers, try better evidence packets, reviewer assistance, selective checking, or cancelling redundant candidates. One more note from September 6: OpenAI reports diminishing reliance on monitoring the model's chain of thought. Provider access to internal reasoning is not the same thing as an explanation or an audit log you control. Keep permission controls, action records and bounded tests alongside artifact review, because none of them certifies safety alone. Then do the practical thing. Choose one recurring job. Name who accepts the work. Record the human work it takes today. Compare what changes with the agent in place. And before you widen the audience, state the readiness state, the evidence gaps and the decision you are asking for, then measure what the recipients spend rather than only what the operator saved. ## The twenty offerings | Offering | Vendor | Intended work | Deployment | Setup and continuing work | Cost components | | --- | --- | --- | --- | --- | --- | | Claude Cowork | Anthropic | cross-app knowledge work | vendor managed with mode-specific local and cloud state | plugins permissions review integration lifecycle | seat and usage meters | | Glean Agents | Glean | enterprise knowledge work | vendor managed enterprise graph | graph hygiene policies evals collaboration review | quote plus model and integration costs | | Gemini Enterprise app | Google Cloud | knowledge work | vendor managed suite | admin policies connectors projects inbox review | subscription and underlying route meters | | Workspace Agents | OpenAI | repeatable knowledge work | vendor managed; route terms apply | authoring enablement permissions evaluation review exceptions | subscription allowance plus credits and tool meters | | Computer | Perplexity | general knowledge work | vendor isolated cloud sandbox | connector policy memory review credit oversight | included credits shared pool optional PAYG | | Grok Bot | xAI | computer work | persistent vendor cloud VM | persistent identity logins approvals monitoring teardown | allowance tokens and enterprise quote | | Copilot Studio | Microsoft | custom business agents | Power Platform tenant | environment DLP identity inventory evals capacity | credits PAYG and prerequisite services | | Agentforce Platform | Salesforce | CRM service and sales | Salesforce org plus chosen model routes | data topics actions permissions evals escalation | license plus credits or unmetered add-on | | AI Agent Studio | ServiceNow | ServiceNow workflows | ServiceNow instance and connectors | specialist build ACLs gateway testing tracing incidents | quote prerequisites and services | | Agents | UiPath | business automation | UiPath cloud with route-specific connectors | CoE tools credentials traces queues human handoffs | platform user consumption and model meters | | Claude Code | Anthropic | software development | vendor service plus buyer execution environment | repo policy tools spend review correction | seat API tokens and CI review labor | | Cline VS Code extension | Cline | repository task | buyer IDE history and chosen model route | repo instructions approvals checkpoints subagents review | inference workstation CI and review labor | | Copilot cloud agent | GitHub | software development | GitHub repository and Actions sandbox | issue specification managed settings PR review correction CI | seat premium requests Actions and review labor | | OpenHands | OpenHands | software tasks | buyer Docker workspace and model route | repo setup sandbox permissions tests upgrades | inference sandbox CI hardware and labor | | goose | AAIF | general and software work | buyer workspace with chosen extensions and model | recipes permissions approvals extensions review | machine inference extensions and labor | | browser-use OSS Python worker | Browser Use | web operations | buyer browser and worker; target sites still contacted | credentials allowlists injection defense exception recovery | model browser proxy CAPTCHA and labor | | Hermes Agent | Nous Research | general personal work | buyer gateway profile and chosen dependencies | skills memory auth policies child agents recovery | high-context inference hardware channels and labor | | OpenClaw | OpenClaw | general personal work | buyer gateway state and chosen dependencies | role credentials patches backups evals incidents | hardware inference channels and operator labor | | Dify | LangGenius | agent workflows | buyer multi-service stack | tenancy secrets plugins backups HA migrations review | database cache vector workers model and labor | | n8n | n8n | bounded agent workflows | buyer workflow engine database and dependencies | workflow credentials queues upgrades exceptions | executions workers database APIs model and labor | ## The exhibits **Exhibit 1. Twenty offerings. Five buying problems.** Selected routes grouped by the work they execute · snapshot: September 4, 2026 Twenty offerings grouped into six managed generalists, four managed workflows, four software executors, four local runtimes, and two self-hosted workflow systems. Sources: Product documentation for the 20 offerings shown · September 2026 **Exhibit 2. The changes arrived at different times.** Seven eligible release comparisons · Q1 to Q3 through September 4, 2026 Seven dated comparisons connect Q1 documentation to later releases across Glean, OpenHands, goose, Hermes, OpenClaw and Dify. Sources: Glean, OpenHands, goose, Hermes, OpenClaw and Dify release notes · January to September 2026 **Exhibit 3. More work clears every requirement.** September 2026 comparison · professional document reasoning, not workplace acceptance GDP.pdf all-pass rate: Astra 33.2%, Sol 28.2%, Fable 5.1 26.2%. Astra gains five percentage points over Sol. Every criterion must pass; these are document benchmark results. Source: Artificial Analysis · September 4, 2026 · GDP.pdf / Index v4.2 **Exhibit 4. Better on one test. Worse on another.** Astra versus Sol · separate benchmark scales, not a combined quality score Separate panels: Astra gains about 80 Elo on AA-Briefcase and loses about 80 on GDPval-AA v2 versus Sol. Elo scales are separate; presentation quality also declined. Source: Artificial Analysis · September 3, 2026 launch evaluation · v4.1.1; approximate Elo changes **Exhibit 5. The model name is not the whole system.** ARC-AGI-3 Semi-Private · standard harness versus OpenAI-native provider adapter ARC reports standard harness max effort at 62.7% for $26,098, provider adapter max effort at 98.6% for $17,332, and provider adapter best observed high effort at 99.9% for $18,817. Same model, different execution conditions; not business-task cost or privacy evidence. Source: ARC Prize · September 3, 2026 · Closed-world games; costs cover the full evaluation **Exhibit 6. The invoice has more than one meter.** Six concrete configurations show where subscription, consumption, and labor enter. Six configurations compare entry, variable, and human operating costs. Local software still incurs hardware, inference and maintenance costs. A conceptual strip separates unit price, total consumption including retries and candidates, and cost per qualifying result. Count human work once, including work displaced. Sources: OpenAI, Anthropic, GitHub, Microsoft, OpenClaw and Dify documentation · September 2026 **Exhibit 7. Review time can dominate token cost.** Illustrative calculation · $1 generation per attempt; labor valued at $60 per hour At twenty review minutes per attempt, cost per accepted result is $42 at 50 percent acceptance, $26.25 at 80 percent, and $21 at 100 percent. These are scenarios, not product measurements. Calculation: ($1 + minutes × $1/minute) ÷ acceptance rate. Excludes setup and downstream costs. **Exhibit 8. Faster generation can outrun review.** Illustrative capacity · one reviewer has 120 minutes available per day Review capacity plateaus at 24, 12 or 6 outputs per day depending on whether each review takes 5, 10 or 20 minutes. Excess arrivals become backlog if retained. Track queue age, rework and unique accepted throughput; reviews are not accepted deliverables. Calculation: min(arrivals, 120 ÷ review minutes). No rework; review is not the same as acceptance. **Exhibit 9. Follow the data beyond the agent.** Architectural comparison · arrows show possible task-data paths, not volume Three architecture lanes contrast vendor execution, local execution with cloud inference, and local execution with local inference. External tool calls can send data outside all three. Illustrative deployment patterns · Check retention, logging and support access for your chosen service. **Exhibit 10. More output came with more coordination.** Separate research context · Ju & Aral, February 2026 preprint; not a Q1 to Q3 trend In an advertising experiment, human-AI teams produced 50 percent more ads per worker, sent 25 percent more task-oriented messages and delegated 17 percent more work. Source: Ju & Aral, Collaborating with AI Agents (2026 preprint) · 2,234 participants; advertising tasks. **Exhibit 11. Execution gains vary by task.** OpenAI-reported September comparisons · benchmark-specific scoring, not one success rate OpenAI reports Sol to Astra: AutomationBench 18.1 to 41.4, OSWorld 2.0 65.7 to 72.6, Terminal-Bench 4.0 37.3 to 57.9. Scores have different definitions; not Q1-to-Q3 outcome evidence. Source: OpenAI · September 3, 2026 · OSWorld uses offline partial score; research evaluation settings **Exhibit 12. A quality drop changes the economics.** Illustrative calculation · hold generation and human handling at $10 per attempt At a fixed ten dollars per attempt, cost per accepted result rises from ten dollars at full acceptance to twenty dollars at half acceptance and twenty-five dollars at forty percent acceptance. Calculation: $10 ÷ acceptance fraction. Assumes stable attempt cost; no downstream loss or setup costs. **Case note. Intensive agent work can consume a lot.** Provider-reported usage · two thresholds, not an exact spending comparison OpenAI reports median researcher usage above $600 per day by mid-August 2026 and current 90th-percentile usage above $7,000 per day in its September 6 report. API-price equivalents, not actual bills. Values are lower bounds; the display does not encode a ratio. Source: OpenAI · Research acceleration · September 6, 2026 · API-price basis; detailed aggregation unverified ## The evidence accounting 20 selected routes by 10 fixed observables produce 200 cells with exclusive statuses: 7 counted documented expansions, 27 documented but not counted, 12 qualified or noncomparable, and 154 with no demonstrated pair, summing to exactly 200. A non-additive coverage statistic, 44 cells with any source from both periods, overlaps these statuses and never joins their sum. No offer meets the four-comparable-expansion-cell threshold; no paired cell carries outcome evidence. The seven counted cells are MG016, MG135, MG146, MG162, MG175, MG177 and MG186, across collaboration (three), persistent state (two), identity and authority (one), and maintenance and recovery (one). Written by Paciva, which builds an executive assistant in this category. Paciva is not among the twenty offerings and was not screened. The screening gates required public availability and documented route evidence at the September 4, 2026 cutoff. Full interactive version: https://paciva.ai/resources/state-of-agents-2026/