Research Report · September 2026

The State of Agents2026

Agents are getting better. What makes the work good enough?

Twenty agent offerings, the changes that matter, and the human work behind useful delegation.

Try Pax
Two hundred agent figures in eight ranks, seven lit. Two hundred route and observable cells are drawn as two hundred figures in eight receding ranks. Seven figures are lit, one for each cell that documents an expansion between Q1 and September 4, 2026. The remaining one hundred and ninety three cells are documented but not counted, qualified, or have no demonstrated comparison pair. Counted expansion means documented operating machinery, not measured improvement in quality, burden, or outcomes. OpenClaw v2026.9.1, maintenance and recovery. Fail closed scans to updater rollback and post update triage. OpenClaw v2026.9.1, persistent state. Shared task ledger to per-agent directories and up to 100 managed worktrees. OpenHands v1.16.0, persistent state. Retained sandbox to a local backend that remembers mode and scopes directory. Dify Community v1.17.0, collaboration. Container restrictions to forms that work inside loops and survive refresh. Hermes Agent v2026.8.31, identity and authority. Completion status to list, steer, stop and per delegation cost. Glean Agents, collaboration. Shared agent library to multiplayer agents and collaborative editing. goose v1.49.0, collaboration. Subagent logs to concurrent notifications and resumable handoff.
20
offerings reviewed
200
observable cells
7
documented expansions
0
outcome cells
Evidence cutoff September 4, 2026Sources verified September 5, 2026Method version 2.118 minute read

Two hundred route and observable cells. Seven document an expansion.

01

The gain is real. So is the gap.

In early September, Artificial Analysis ran GPT-6 Astra against a document task where every rubric criterion has to pass. Astra cleared all of them 33.2 percent of the time. Its predecessor, GPT-5.6 Sol, cleared them 28.2 percent of the time. That five-point gain is real, and someone other than the vendor measured it. It also means that on roughly two-thirds of those documents, something still failed.

The same shape shows up in the products. Between January and early September, Glean added collaborative editing to its agents, goose added handoffs a person can resume, and OpenClaw added updater rollback with a diagnostic pass after updates. Real changes to what the software does. None of them measures whether the work got better, or what the people around the work now carry.

Three things blur together and are worth keeping apart: what a model scores on a specified evaluation, what a product documents it can now do, and what happens to the work in an actual organization.

The central tradeoff runs beneath it all. More generated work can remove drafting and reconstruction. It can also add briefing, checking, correcting, and coordinating. The net effect is a question to measure. A finished file is not necessarily ready for review, approval, or use.

If you work alone, pick one recurring job you can judge yourself and count your own review and upkeep time inside the result. If you run a small business or a startup, give one person the whole workflow, including the parts exported to sales, finance, delivery, and customers. If you buy for an enterprise, evaluate the product with its data paths, permissions, review capacity, recovery, and procurement terms, because the product alone is not what you are deploying.

Twenty offerings are covered, frozen as exact routes at a September 4, 2026 cutoff. Astra sources were refreshed September 5.

A running example, and it is fiction

One document, four directions of pull. Sales wants it persuasive, finance needs the margin floor held, delivery needs a schedule it can meet, and the customer needs commitments that hold. The example is fictional, and no trial was run.

One example runs through every section, and it is fiction. An agent prepares a customer proposal from call notes, the current price list, approved contract terms, and delivery capacity. Sales wants it persuasive. Finance needs the margin floor held. Delivery needs a schedule it can meet. The customer needs commitments that hold.

There is no single golden proposal, which is the point. The arithmetic and the approved terms can be checked against a source. Tone, emphasis, and which tradeoff to offer require judgment, and prices and capacity can both change after the draft is written. Nothing below reports a trial of this workflow, because no trial was run.

02

Choose the job before the product.

For the purposes of this report, an agent is a system that takes a goal, chooses and executes tool actions, observes what came back, and adapts while keeping track of the task. That separates it from an assistant that answers questions and from automation that follows a fixed path. Durable memory, bounded permissions, and recovery are things to buy on, not tests a system has to pass before the word applies.

Three things get sold as one and are not one: the model, the executor that takes the actions, and the workspace where people and agents coordinate. Astra is a model used inside an offering, not an offering. Buzz is a collaboration overlay and belongs entirely outside the executor comparison.

The table below is the inventory as documented: twenty offerings, the work each is built for, how it deploys, the operating work it asks of you, and the shape of its bill. It carries no ranking or score, and the order reflects the study's stable display order. Read it to shortlist, not to choose.

Model
The intelligence that reasons over the task
Astra is one of these
Executor
The system that takes the actions and holds the permissions
What you actually operate
Workspace
Where people and agents coordinate the work
Buzz is an overlay here
Three things sold as one. Benchmarks describe the first, your bill and your permissions live in the second, and your coordination costs live in the third.
OfferingIntended workDeploymentSetup and continuing workCost components
Claude CoworkAnthropic · EnterpriseCross-app knowledge workVendor managed with mode-specific local and cloud statePlugins permissions review integration lifecycleSeat and usage meters
Glean AgentsGlean · EnterpriseEnterprise knowledge workVendor managed enterprise graphGraph hygiene policies evals collaboration reviewQuote plus model and integration costs
Gemini Enterprise appGoogle Cloud · BusinessKnowledge workVendor managed suiteAdmin policies connectors projects inbox reviewSubscription and underlying route meters
Workspace AgentsOpenAI · EnterpriseRepeatable knowledge workVendor managed; route terms applyAuthoring enablement permissions evaluation review exceptionsSubscription allowance plus credits and tool meters
ComputerPerplexity · Enterprise MaxGeneral knowledge workVendor isolated cloud sandboxConnector policy memory review credit oversightIncluded credits shared pool optional PAYG
Grok BotxAI · EnterpriseComputer workPersistent vendor cloud VMPersistent identity logins approvals monitoring teardownAllowance tokens and enterprise quote
Copilot StudioMicrosoft · StandaloneCustom business agentsPower Platform tenantEnvironment DLP identity inventory evals capacityCredits PAYG and prerequisite services
Agentforce PlatformSalesforce · Employee Agent licenseCRM service and salesSalesforce org plus chosen model routesData topics actions permissions evals escalationLicense plus credits or unmetered add-on
AI Agent StudioServiceNow · Enterprise PlusServiceNow workflowsServiceNow instance and connectorsSpecialist build ACLs gateway testing tracing incidentsQuote prerequisites and services
AgentsUiPath · EnterpriseBusiness automationUiPath cloud with route-specific connectorsCoE tools credentials traces queues human handoffsPlatform user consumption and model meters
Claude CodeAnthropic · EnterpriseSoftware developmentVendor service plus buyer execution environmentRepo policy tools spend review correctionSeat API tokens and CI review labor
Cline VS Code extensionCline · v4.1.17Repository taskBuyer IDE history and chosen model routeRepo instructions approvals checkpoints subagents reviewInference workstation CI and review labor
Copilot cloud agentGitHub · EnterpriseSoftware developmentGitHub repository and Actions sandboxIssue specification managed settings PR review correction CISeat premium requests Actions and review labor
OpenHandsOpenHands · Community v1.16.0Software tasksBuyer Docker workspace and model routeRepo setup sandbox permissions tests upgradesInference sandbox CI hardware and labor
gooseAAIF · v1.49.0General and software workBuyer workspace with chosen extensions and modelRecipes permissions approvals extensions reviewMachine inference extensions and labor
browser-use OSS Python workerBrowser Use · v0.13.8Web operationsBuyer browser and worker; target sites still contactedCredentials allowlists injection defense exception recoveryModel browser proxy CAPTCHA and labor
Hermes AgentNous Research · v2026.8.31General personal workBuyer gateway profile and chosen dependenciesSkills memory auth policies child agents recoveryHigh-context inference hardware channels and labor
OpenClawOpenClaw · v2026.9.1General personal workBuyer gateway state and chosen dependenciesRole credentials patches backups evals incidentsHardware inference channels and operator labor
DifyLangGenius · Community v1.17.0Agent workflowsBuyer multi-service stackTenancy secrets plugins backups HA migrations reviewDatabase cache vector workers model and labor
n8nn8n · Community v2.37.9Bounded agent workflowsBuyer workflow engine database and dependenciesWorkflow credentials queues upgrades exceptionsExecutions workers database APIs model and labor

Twenty selected routes in the studyโ€™s stable display order. No ranking, no score. This is not a price quote, an observed setup-time comparison, or proof of segment suitability.

Shortlist on the job you actually have, the applications you already run, the technical capacity you can staff, what your data requires, and whether you can tell good output from bad in that domain. Two warnings travel with the table. A free license does not remove the operating bill. And the length of an upkeep list is not a burden score, because a long list of small tasks can cost less than a short list of hard ones.

The categories in that table do real work. Managed generalists arrive with the route decided for you. Managed workflow platforms assume you have a process to encode and people to encode it. Software executors are aimed at repositories. Local runtimes and self-hosted systems hand you control and the operating burden in the same motion. A coding executor might build the proposal workflow perfectly well and still be the wrong interface for the salesperson who uses it every day.

Company size is not the same variable as company stage, and neither settles this. A technical solo founder can run software a small service business would need help maintaining, and enterprise integration capacity does not remove the requirement to review the work or own the outcome. Regulation and data sovereignty cut across all of it.

One correction is load-bearing. OpenAI introduced Workspace Agents on April 22, 2026, not at Astra's September launch. The Work and Codex documentation verifies Astra separately. Do not carry Astra's model identity, admin defaults, or benchmark results across to Workspace Agents, and treat the current rollout wording on the launch page as unsettled, because that page mixes updated availability language with older preview text.

03

Seven documented changes, and what the count does not mean.

The study fixed twenty routes against ten observables, which produces two hundred cells. Seven of those cells document an expansion between Q1 and the September cutoff, with exact dates on both ends and a matched claim. They land in four observables, not ten.

Collaboration accounts for three. Glean moved from a shared agent library to multiplayer agents with collaborative editing. goose moved from subagent logs in the interface to concurrent notifications and a resumable handoff. Dify moved from restrictions inside invalid containers to forms that work inside loops and survive a refresh. Persistent work accounts for two: OpenHands moved from a retained sandbox to a local backend that remembers its mode and scopes a directory, and OpenClaw moved from a shared task ledger to per-agent directories and up to one hundred managed worktrees. Intervention accounts for one, where Hermes moved from reporting completion status to letting an owner list, steer, and stop delegations and see cost per delegation. Maintenance accounts for one, where OpenClaw added updater rollback and a diagnostic triage pass after updates.

Here is what that seven does not say. It does not say seven of two hundred agents improved. It does not say product quality improved. It does not describe how common any of this is in the market. The number measures how much of this evidence the research could document under its own rules, and nothing else.

AreaDocumented changeBuyer question
CollaborationGlean collaborative editing; goose resumable handoffs; Dify review forms that survive refreshDoes shared work reduce repeated explanation or lost feedback?
Persistent workOpenHands remembers workspace mode; OpenClaw separates agent work areasIs context current, and can parallel work be reconciled?
InterventionHermes adds steering, stopping and delegation-cost visibilityCan the owner intervene before a mistake spreads?
MaintenanceOpenClaw adds updater rollback and post-update diagnosticsDoes recovery work in this deployment?
Showing 200 of 200 cells.

Click any cell for its documented Q1 and September states, the claim the ledger permits, and the inference it prohibits. The 154 nearly empty cells are findings, not rendering gaps: no demonstrated comparison pair exists there.

What Astra changes

Astra's document result is the clearest independent gain: the all criteria pass rate rises from 28.2 to 33.2 percent against Sol. Around it the independent picture is mixed. At launch, Artificial Analysis measured AA-Briefcase roughly 80 Elo higher and GDPval-AA v2 roughly 80 Elo lower, with presentation quality declining and no numeric delta given, and those two Elo figures do not sit on a common scale. The coding index rose about two points to 67 at roughly unchanged task cost in Codex at maximum effort, while general evaluation cost rose 75 percent. That 75 percent belongs to the v4.1.1 composition and must not be carried onto v4.2, which changed both the test composition and the Elo anchors, so v4.2 putting Astra four points above Sol is not a continuation of the v4.1.1 result.

ARC Prize ran the same model through different execution setups and got different answers. At maximum effort the standard harness scored 62.7 percent for $26,098, while the provider adapter scored 98.6 percent for $17,332. A separate best-observed adapter run at high effort reached 99.9 percent for $18,817 at a different effort setting, so it is not a like-for-like comparison. The execution system moves the result. That establishes system sensitivity, not enterprise reliability, and not a privacy or memory advantage.

OpenAI's own evaluations report execution gains from Sol to Astra: AutomationBench 18.1 to 41.4, OSWorld 2.0 65.7 to 72.6 on an offline partial score, and Terminal-Bench 4.0 37.3 to 57.9. These are vendor-reported, do not share a single definition of success, and none measures management savings.

The search found no independent Astra field study measuring total human handling, downstream rework or full cost of ownership. That is a gap in the record, not evidence that those outcomes cannot improve. And one note keeps the timeline honest: June releases are visible in the September snapshot rather than being Q3 events, and Astra against Sol is a September comparison against a predecessor, not a measured movement from Q1 to Q3.

04

Quality is a contract, not a score.

Go back to the proposal. Quality there means work that named people can accept and use under the requirements in force right now. A golden reference answer is one way to evaluate that, and it is not available for most knowledge work.

Three kinds of content sit inside the same document and need different handling. Checkable facts: the prices, the arithmetic, the approved terms. Negotiable preferences: tone, emphasis, which alternatives to present. Reserved decisions: discount exceptions, capacity commitments, legal concessions, permission to send. Name who decides each one, keep dissent visible rather than averaging it away, and stop consequential commitments when the people with authority have not resolved the conflict. An average rating cannot erase a prohibited term.

A dated note from September 6 sharpens the rubric without adding a result. OpenAI's chief scientist draws a line between achieving a stated goal and exercising sound judgment when objectives are unfamiliar or in conflict, and observes faster progress on the capabilities easiest to measure. For a buyer that means judging the deliverable and the conduct that produced it. An accurate proposal can still fail if the agent exceeded its authority.

Name the next step, not the quality

Work arrives in one of three states. The label names what the recipient is being asked to contribute, not a mandatory stage it had to pass through.

Early collaboration is valuable when people know they are being asked to help resolve uncertainty. The failure is presenting exploratory work as decision ready. Recipients can flag missing evidence or ask for review; they cannot veto every preference. Constraints that cannot be waived stay blocking, and the standard scales with the consequence.

EXPLORATION

Open questions, provisional assumptions, known gaps. The reader shapes the problem.

REVIEW

Fundamentals checked, evidence traceable. The reader challenges and resolves named issues.

APPROVED FOR DECISION

Objections dispositioned by the owner. The reader decides within their authority.

The stamp names the next contribution being asked for, not a quality score and not a mandatory sequence.
StateWhat recipients receiveWhat they are asked to do
ExplorationOpen questions, provisional assumptions, known gapsShape the problem or supply missing expertise
ReviewChecked fundamentals, traceable evidence, unresolved choicesChallenge the analysis and resolve named issues
ApprovalObjections resolved or dispositioned by the authorized owner, permitted exceptions documented, current versionMake the specified decision within their authority

Finished is not the same as inspectable

A practitioner transcript circulated during this work is worth using as an illustration and not as a ranking. It describes a low effort workbook produced without dedicated sources or checks sheets, a higher effort run that surfaced more decision relevant questions, and a third workbook that was easier to inspect. The original files, their correctness, their cost and their repeatability were not independently checked here.

The buyer question underneath it is concrete. Can another person trace an important assertion back to its source, tell evidence apart from assumption, reproduce a calculation, see which checks ran and what those checks could not cover, and identify what would change the recommendation? A sources sheet proves none of that on its own, and neither does a PASS label. More sheets, more citations, more tokens and more polish do not establish better reasoning.

The contrast with software is less clean than it first appears. In one review of algorithmic against holistic evaluation, maintainer written tests credited partial success on work that human review rejected, and none of fifteen manually reviewed pull requests was mergeable without further work. Deterministic tests can miss what acceptance requires. That is a failure mode, not a rate to carry to other models, products or tasks.

Onboarding, and what feedback actually changes

Write the packet before the first task: the purpose, the current sources and who owns them, examples of good and bad, the decision limits, the exceptions, the tools and permissions, who to ask, how corrections get recorded, when instructions expire, and how to stop, recover and retire the workflow.

Instructions and memory are not training. Connecting an application does not tell the system which source outranks another. Human onboarding costs time too, so the comparison that matters is incremental work against the previous process rather than against zero. Make the packet relational and not merely documentary: who knows what, whose approval matters, which commitments exceed authority, and when to ask instead of proceed. Access to the company file store is not evidence of any of that.

Sort incoming feedback into four buckets: a factual correction, an existing requirement, a proposed preference, or an authorized decision. An owner can adopt a preference as a scoped requirement. The most recent comment confers no authority by itself, and an exception needs both an authorized decision and a rule that permits exceptions at all.

Follow one correction all the way through. Finance objects to a discount. That becomes an authorized rule update, the next proposal gets checked against it, and then it gets tested again after a later price change. Record where the correction came from, who authorized it, what it covers, when it expires and what it changed.

The research here informs the design and does not prove an effect on agents. A meta analysis of 607 feedback effects across 23,663 observations found an average improvement of d = .41, and more than a third of feedback interventions reduced performance. In a separate study of 1,401 participants, human and AI feedback loops amplified bias in ways participants did not recognize. Feedback is a mechanism that moves in either direction, so treat a feedback loop as something to test rather than something to install.

Fixing this proposal is not the same as improving the next proposal. Show the authorized correction reaching the rule or the scoped memory, tell the affected people it changed, and check later work for the same defect returning. Track repair time separately from evidence that the defect stopped.

Standards move, so recheck against them

Two older studies are useful as context, with their era and setting stated. In a 2026 field experiment, 758 consultants completed 12.2 percent more tasks 25.1 percent faster with higher quality on eighteen tasks inside the capability frontier, and scored 19 percentage points lower on one task outside it. In a study of 5,172 customer support agents across 133 teams and 3,006,395 chats, resolutions per hour rose 0.30, which is 15.2 percent, with the lowest skilled quintile gaining 0.5 per hour or 36 percent while the top skilled showed no significant gain and small significant declines in resolution rate and satisfaction. Neither study measures anything about Q1 to September 2026, and both used older models.

Combine stable test tasks with samples of current work, and recheck after any meaningful change to inputs, requirements, model, tools or permissions.

05

Count the work everyone carries.

Premature circulation exports unfinished work

Sales circulates a polished draft. Finance reconstructs the assumptions underneath it because they were never stated. Delivery challenges a promise it cannot meet. The approver cannot tell which objections are still open. A revised copy arrives and the discussion starts over. People who never generated the draft are now carrying its unfinished work. This is an illustrative failure pattern rather than a measured agent effect, and human authors produce it just as reliably.

Start a small log rather than a measurement program. Record the result and its readiness state, the total human handling across every role it touched, and the principal defect or rework. Count incidental recipients, failed attempts and the time spent measuring. Keep delay and displaced attention separate from labor minutes instead of converting them into money.

Management tax is all human handling with the agent minus comparable human handling in the process it replaced. It can be positive or negative. Positive does not by itself mean waste, because that difference includes productive judgment and necessary governance. With no matched baseline, describe the workload and claim no savings. Classify preventable reconstruction separately, because that is the part worth attacking.

Now run the better version. The agent packages its sources, flags what it is uncertain about, routes the precise discount question to the person who owns that decision, and notifies the people whose commitments changed. Compare the reconstruction and follow along effort against the first version, and do not assume the saving. The target is preventable reconstruction, not participation. Work moved between teams inside the same organization is cost displacement and not automatically an externality.

One early draft, five desks of rework. A polished looking document circulates before its assumptions settle, and every recipient inherits a growing stack of copies and corrections.

Coordination gets removed as well as added

The same agent can answer a routine question from approved material before it ever reaches finance, and escalate when authority or uncertainty is genuinely unresolved. That displaces support and clarification work while other tasks add review, and both directions have to be measured. OpenAI reports reduced activity in an internal technical support channel and lower attendance at some office hours. That is a different setting from the advertising experiment discussed below, and fewer messages do not establish time saved or that every support need was met.

Three scenarios, with the assumptions on the face of them

At one dollar per generated attempt and sixty dollars an hour for review, twenty review minutes per attempt costs $42 per accepted result at 50 percent acceptance, $26.25 at 80 percent and $21 at 100 percent. Review dominates generation across that whole range. One reviewer with 120 minutes a day clears 24 outputs at five minutes each, twelve at ten minutes and six at twenty, so generation above that line becomes a queue rather than throughput. And at ten dollars per attempt, moving from full acceptance to half doubles the cost per accepted result from $10 to $20, with 40 percent acceptance taking it to $25 before any downstream loss.

These are scenarios, not vendor measurements. A complete cost of ownership also carries setup, maintenance, integration, failure recovery, the recipient's work, switching and exit. A twelve month projection is not measured annual performance, and zero accepted results leaves the ratio undefined rather than large.

$1
20 min
$60/hr
50%
120 min
40/day
Cost per accepted result
$42
Exhibit 7 at these defaults
Reviews cleared per day
6
Exhibit 8 at these defaults
Backlog added per day
34
Arrivals above capacity
$10 attempt at this acceptance
$20
Exhibit 12 at these defaults

These are scenarios, not product measurements. No setup, downstream or failure costs are included. Nothing you set here leaves your browser.

More agent work is not the same as cheaper finished work

Five numbers get collapsed into one and should not be: the unit inference price, the resource cost per attempt, the number of attempts consumed, total spend, and the fully loaded cost per accepted outcome. Cheaper attempts may encourage more experiments or larger assignments, which is a plausible mechanism rather than a measured price response. More usage can exhaust a subscription allowance without changing this month's bill at all. Cash paid, allocated subscription cost and API price equivalent usage are three different accounting bases and should never be summed.

Tie retries, abandoned runs and parallel candidates back to the deliverable they belong to. Five candidates that produce one accepted proposal are one outcome, not five. Match costs and outcomes to the same cohort of work, because this week's bill and this week's acceptances often describe different jobs. Pair the capacity picture with queue age, active reviewer time, rework and unique accepted throughput. Repeated approvals of the same artifact are not new deliverables, and waiting is not labor.

Match the effort to the stage, not to the prestige of the model

Inspecting an artifact and overseeing behavior are different jobs. Sources and checks help a reviewer read a document. They cannot certify that an action the agent took was authorized. Count the work of checking actions separately from the work of checking outputs.

Exploring cheaply, then going deeper only where the uncertainty is consequential, then preparing the work for its audience, is a workflow hypothesis worth testing rather than a prescribed sequence. Record the model, the effort setting, the tools, the checks, elapsed time, total inference spend and the human handoff work, then test whether the extra effort improved the decision or reduced review enough to pay for itself. More effort can still be wasted effort. And running two models to compare their answers is not independent verification, because shared sources and shared assumptions reproduce the same error.

Shared visibility is not reduced coordination

Grok Bot's persistent environment may reduce repeated briefing. Set that against credential management, memory upkeep, monitoring and teardown. For Buzz, look at whether explanation gets duplicated, how notifications land, and what reconciliation costs, holding the underlying agents and task constant.

The question to test is whether the workspace delivers a relevant change and a precise question to the right person, or requires everyone to follow the whole thread. Measure notification load, participants per decision, repeated explanations and lost objections, and include the positive case, because agents can summarize changes, route questions, preserve dissent and track commitments.

One experiment is directly on point. Across 2,234 participants producing 11,024 advertisements, working with AI agents raised ads per worker by 50 percent, raised task oriented messages by 25 percent, and raised delegation by 17 percent, while image quality and diversity suffered. Output and communication rose together. Messages are not minutes, and this is an advertising task rather than a proposal workflow.

A survey of production deployments adds context with its denominators attached: across twenty interview cases and 306 raw responses covering 86 deployed or piloted systems, 41 of 60 relevant cases ran ten steps or fewer before a human intervened, and 23 of 31 used human in the loop evaluation. Continued human involvement is the norm in the systems that were studied.

Illustration: a team discussing work together

See the whole formation on your own work

Pax is an executive assistant for individuals and a chief of staff to organizations. The fastest way to judge it is on a recurring job you already own.

06

Custody belongs to the route, not the label.

Follow the customer details, the prices and the negotiated terms through three arrangements. Vendor operation, where the provider runs everything. Local execution with cloud inference, where your machine drives the work and a remote model does the thinking. Local execution with local inference, where both stay with you. Dedicated hosting and restricted egress configurations sit between these and change the answer again.

Compare the arrangements on patching, secrets, identity, backups, monitoring, support, recovery, export and deletion, because those are the costs that arrive after the decision. And note the one that catches people: local execution can still send data to models and tools. Local capability and practical locality are different things, and neither one means private.

Privacy and intellectual property need separate answers. On privacy, ask about training use, retention, support access, subprocessors, residency and key control. On intellectual property, ask about rights to inputs and outputs, code and weight licenses, whether state can be exported, and what switching actually costs. Contract promises and observed behavior answer different questions, and a hosting location is not a custody finding.

The practical version is a walk, not a questionnaire. Take one real document through the route you are buying and write down every place it comes to rest: the prompt and its attachments, the model provider's logs, the tool the agent called and whatever it retained, the trace layer, the backup, the support ticket that will quote it back six weeks from now. Every resting place is a separate retention, residency and deletion question.

The cost picture has the same shape. A unit price, total consumption and the cost of an accepted outcome are three different meters, and a cheaper meter need not produce a smaller bill or cheaper finished work. Provider incentives run through seats, credits, token and tool usage, prerequisite services and switching costs. That is worth understanding without the blanket claims that providers only sell tokens or that every hyperscaler is unprofitable.

Independent evaluations help here when their dates and execution conditions are stated, and some of them evaluate a model inside a named executor rather than the model alone. None of them certifies your workflow. The ARC result establishes that execution configuration matters. It establishes nothing about custody.

Two limits on this section are worth stating plainly. Every custody route in the study is recorded as a documented route shell with the decisive custody fields not publicly demonstrated. Every authority and recovery row is recorded as not sufficient for a recommendation. Not publicly demonstrated means the evidence was not published, not that the capability is absent.

One document, six resting places. The prompt and its attachments, provider logs, the tool it called, the trace layer, the backup, and the support ticket six weeks later. The third stop is where most routes surprise people, and every stop is its own retention, residency and deletion question.
07

Function, authority and reach are three separate axes.

Return to the proposal one more time. An agent can produce, coordinate, allocate and handle exceptions, or support leadership, and each of those functions demands its own evidence before you delegate it.

These functions do not form a ladder and they are not a promotion track. Authority and reach are separate axes crossing them: authority is what the agent may do without asking, running from reading to drafting to reversible action to consequential action, and reach is how far the effect travels, from a personal task to a team workflow to a cross functional portfolio. An agent can produce with high authority over a personal task and coordinate with almost none across a portfolio. Those are different risk positions and they need different evidence.

Authority and reach raise the possible consequences of a mistake. They are risk dimensions, not achievement scores, and human accountability for consequential commitments stays explicit at every level.

Describe progress as a larger demonstrated scope of useful work: corrections persist where they should, exceptions get handled, reviewers keep up, and results survive changing inputs. This report does not predict a date for autonomous managers, and no evidence here supports one.

One founder may hold every human role in this picture. A small business splits them across a few people. An enterprise distributes them widely. A shared agent interface does not move any of those decision rights.

Domain expertise includes knowing when work is ready, noticing what evidence is missing, and knowing whose input matters. If agents absorb junior production tasks, the organization needs an answer for how it will develop the reviewers it will need in five years. That is a workforce design question to plan for, not an established claim about deskilling.

Functionproduce, coordinate, advise
Authorityread, draft, reversible action, consequential action
Reachpersonal task, team workflow, cross functional portfolio
One agent occupies one point in this space per workflow. Moving out on any axis raises the consequences of a mistake, so each move needs its own evidence.
FunctionExampleWhat must be demonstrated
ProduceDraft from approved inputsAccepted work under current requirements
CoordinateGather finance and delivery decisionsAccurate handoffs, preserved objections, manageable review
Allocate and handle exceptionsSuggest priorities across competing proposalsSound choices under change, escalation and recovery
Support leadershipCompare markets and strategic commitmentsEvidence quality, alternatives, uncertainty, useful challenge
Authority and reach are risk dimensions, not achievement levels.

Function

Authority

Reach

08

Adopt in proportion to the decision.

Three situations call for three different next steps, and each one earns a different size of conclusion. The rule underneath them is that the evidence has to match the consequence, not the enthusiasm.

Personal, reversible draftingLowest assurance
Next step
Small supervised pilot on familiar tasks, with a time log that includes your own review and upkeep.
What the evidence then supports
A limited personal workflow decision. Nothing more.
Repeated work used by othersBounded assurance
Next step
Representative tasks, named reviewers, downstream checks, and a recovery exercise before anyone depends on it.
What the evidence then supports
A bounded operational decision inside the conditions you tested.
Consequential action or a broad comparative claimHighest assurance
Next step
Stronger controls, appropriate expertise, and a study sized to the claim you intend to make.
What the evidence then supports
Only the specific scope and claims you validated.

Three to five runs will expose setup and measurement problems, which is useful. They cannot establish broad superiority. Formal confirmation belongs wherever the consequences require that assurance, and it is not a prerequisite for reversible personal use. Claims about improved quality or reduced human handling need a baseline from the current process. Claims about expert performance need an expert comparator. Expert parity is not a universal adoption gate.

Expansion can chase quality, responsiveness, accessibility or control just as legitimately as savings. Name the benefit you are buying and the cost you will accept. Require evidence for a claimed cost advantage rather than insisting every adoption has to be cheaper.

Stop or narrow the scope when errors escape review, when queues exceed capacity, when sources become unreliable, when recovery fails, or when recipients inherit rework they did not sign up for. Recheck after any material change.

Rising activity with flat accepted throughput calls for diagnosis rather than an automatic verdict, because exploration or genuinely higher quality can justify the effort. Before adding agents or reviewers, try better evidence packets, reviewer assistance, selective checking, or cancelling redundant candidates.

One more note from September 6: OpenAI reports diminishing reliance on monitoring the model's chain of thought. Provider access to internal reasoning is not the same thing as an explanation or an audit log you control. Keep permission controls, action records and bounded tests alongside artifact review, because none of them certifies safety alone.

Then do the practical thing. Choose one recurring job. Name who accepts the work. Record the human work it takes today. Compare what changes with the agent in place. And before you widen the audience, state the readiness state, the evidence gaps and the decision you are asking for, then measure what the recipients spend rather than only what the operator saved.

09

Questions readers ask.

What actually changed between Q1 and September 2026?

Seven of 200 route and observable cells document an expansion with exact dates: three in collaboration, two in persistent work, one in intervention controls, one in maintenance. The count measures what the research could document under its own rules, not market-wide improvement.

Which agent is the best one to buy?

The report refuses to rank the twenty offerings, because no evidence in the package supports a universal winner. Every rank in the study is suppressed. Shortlist by the job, your applications, your technical capacity, your data requirements and your ability to judge results.

What is the management tax?

All human handling with the agent minus comparable human handling in the prior process. It can be positive or negative, and positive does not alone mean waste. With no matched baseline, describe workload and claim no savings.

Does running an agent locally keep my data private?

Not by itself. Local execution can still send data to models and tools, and every custody claim in the study is a documented route shell with the decisive fields not publicly demonstrated. Custody belongs to the exact end-to-end route, never to a hosting label.

What did GPT-6 Astra actually improve?

Independently: the GDP.pdf all-criteria pass rate rose from 28.2 to 33.2 percent against Sol, and one knowledge-work evaluation improved while another declined. Vendor-reported computer-use benchmarks rose. No independent field study measured total human handling, rework or full cost of ownership.

How many test runs do I need before rolling an agent out?

Three to five runs expose setup and measurement problems; they cannot establish broad superiority. Match the evidence to the decision: a personal reversible workflow needs a supervised pilot and a time log, while a broad comparative claim needs a study sized to that claim.

What is a readiness state?

A label that names the next contribution a recipient is asked to make: exploration, review, or approval. It replaces the guess. The failure it prevents is presenting exploratory work as decision ready and exporting the unfinished parts to people who never generated it.

Who wrote this, and what is the Paciva connection?

Paciva, which builds an executive assistant in this category. Paciva is not among the twenty offerings and was not screened; the screening gates required public availability and documented route evidence at the September 4 cutoff.

10

The evidence base, and what it refuses to say.

7
counted documented expansions
27
documented, not counted
12
qualified or noncomparable
154
no demonstrated pair

The four statuses are exclusive and sum to exactly 200. A separate, non-additive statistic, 44 cells with any source from both periods, overlaps these statuses and never joins their sum. No offer meets the four-comparable-expansion-cell threshold, and no paired cell carries outcome evidence.

SY

Synthesis claims with permitted and prohibited readings

AS

Astra benchmark and vendor sources, dated

SC

Published research with transfer limits

MG

Route and observable movement cells

CL

Atomic offer claims from product documentation

Seven publication gates hold throughout: no universal best-agent rank; no cost per accepted task without a fixed task and a complete human-cost ledger; no reduced management tax without a baseline; no effective privacy, authority or recovery from documentation alone; no business-outcome improvement inferred from feature releases; no fully local, open source or private without exact route evidence; and not publicly demonstrated never reads as absent.

Logos and names indicate citation, not endorsement.

Written by Paciva, which builds an executive assistant in this category. Paciva is not among the twenty offerings and was not screened. The screening gates required public availability and documented route evidence at the September 4, 2026 cutoff. Method version 2.1; evidence cutoff September 4, 2026; Astra sources verified September 5, 2026.
Illustration: a professional reaching an outcome

Get started with Pax

Less noise, more signal, and a running record of what your team actually decided.

Get started with Pax