--- title: Privacy Follows Capability authors: Jeremy Mays, Frederick Townes publisher: Paciva published: 2026-09-12 canonical: https://paciva.ai/resources/privacy-follows-capability/ license: Published for public use with attribution --- # Privacy Follows Capability > The useful AI coworker reads mail, remembers meetings, searches files, connects to accounts, and acts for you. Every new capability widens the privacy surface. The answer is not weaker AI. It is the least exposing route that can still do the work, with identity and secrets kept inside the smallest necessary boundary. | | | |---|---| | Vendors and surfaces reviewed | 9 companies, 24 named surfaces | | Ledgered claims | 74 claims, 7 publication gates | | Time-sensitive rows rechecked | September 12, 2026 | | Reading time | About 30 minutes | Bracketed identifiers such as [R-01] are claim identifiers from the report's fact check ledger. The evidence base at the end lists the primary sources. ## 1. The privacy problem starts when the AI becomes useful **Privacy risk grows with what an AI can see, remember, infer, connect to, and do. Private AI therefore routes each task by sensitivity, required capability, and permitted authority.** Ask an assistant to handle your day and look at what handling it requires. Mail, calendar, documents, messaging, contacts, the open web, memory of what happened last week, and permission to send or change something on your behalf. The request is three words. The system boundary it opens is enormous, and the product feels simpler the larger that boundary gets. [R-01] Compare two things that both get called AI. The first is a blank prompt box. You paste what you choose, it answers, you close the tab. The second is a coworker that watches what arrives, remembers what matters, and acts without being asked twice. The second is more useful for exactly one reason: it has more context and more authority. That is not a flaw in the design. That is the design. Most privacy conversations then collapse five separate controls into one word. Five questions sit behind that word. Does your content train a model. Who receives it. How long is it retained. Who may review it. What is the system authorized to do. Each carries its own answer, and a strong answer to one tells you nothing about the rest. A vendor that promises no training has not promised no retention, no review, no egress, or no action. [R-02] Avoidance has to start before anything goes wrong, because the failure modes here are not reversible. Unnecessary disclosure cannot be made necessary after the fact. Aggregation can reveal a relationship, a diagnosis, a negotiation, or a strategy that no single document contained. Durable memory extends the window in which a mistake can still hurt you, and untrusted content that reaches memory can influence decisions made much later. [L-04, R-05] That last one deserves separating from the rest. Where the data sits and what the system may do are independent choices. A model running on hardware you own, wired to send mail and move money without asking, is a larger risk than a hosted model that can only draft. Placement is custody. Authority is blast radius. Treating them as one decision is how a privacy-first deployment produces a worse incident than the one it was built to prevent. [R-07] The rule is not "minimize everything." Every capability needs a paired control, and every consequential route needs evidence a person can inspect afterward. > The assistant becomes indispensable at the same speed it becomes nosy. ### Exhibit 1. Five verbs, and the failure each one carries Every capability needs a paired control. The first column is what the system does, the second is what it reaches, and the third is how the same capability fails. - **See.** Messages, files, screens, records, and accounts. Failure mode: Receives more context than the task required. - **Remember.** Chats, indexes, logs, screenshots, caches, backups. Failure mode: Turns a temporary task into durable state. - **Infer.** Identities, relationships, health facts, strategy. Failure mode: Ordinary fragments become sensitive together. - **Connect.** Integrations, tools, subprocessors, other models. Failure mode: One task quietly creates several processors. - **Act.** Sending, paying, filing, deleting, publishing. Failure mode: Private context becomes an irreversible consequence. *Source: the NIST software and AI agent identity and authorization concept paper, a draft rather than binding law, plus the vendor documentation in the evidence base. Verified August 25, 2026. An operating recommendation, not a measured model.* > Read across, not down. Answering one row tells you nothing about the rows beside it. ### Diagram. Each verb the system gains widens what it can reach The column heights are relative, not measured. They carry one idea: capability and exposure move together, so the control has to move with them. > The last column is the one procurement forgets. Authority is what turns exposure into something nobody can take back. ## 2. The brand is not the privacy boundary **A brand name covers far more than one data contract. The unit that governs your task is the exact product surface, not the logo.** Under one brand you will find a consumer app, a managed workspace, a direct API, a hyperscaler route, a connector, and an agent, and they do not share rules. The unit that governs your task is the exact product surface, the granted scope, the retained state, the model route, the human access path, the tool authority, and the deletion trigger. Start with the assistants that behave like coworkers, because they are the ones expanding all five verbs at once. Viktor works inside Slack. Its privacy policy, last updated July 31, 2026, says that on joining a channel it may pull up to about 90 days of that channel's history for context, and that separately granted Slack scopes let it search broader private content. The same policy says Viktor trains no foundation model, its own or a third party's, on customer data. It names one internal exception: service-specific quality models, such as a router that picks which AI provider serves a request. It also says AI providers may temporarily retain data under their own API retention policies, which is a separate question from training. Deletion from active production systems takes about 30 days, and encrypted backups age out on a roughly 35-day rotation. A Viktor blog post states seven business days where the policy states about 30. Both are published here, and which one governs a given surface is a procurement question rather than a contradiction. Its security page documents approval for money movement, code pushes, and customer email. [V-01, V-02, V-03, V-04, V-05] Catch runs inside an inbox. Its security page says AI vendors are contractually prohibited from retaining or training on request data, and, two paragraphs later, that Catch uses your data to continuously improve the experience for you and others. Both statements can be true at once, because the first governs third parties and the second governs Catch, and the mechanism behind the second is not described. Integrations request minimum permissions and can be revoked at any time, and actions happen within granted permissions and with approval. Catch works in the background of an inbox, and it can send email, send texts, and place calls. No numeric deletion schedule and no human access rule appear on the security page, and a separate privacy policy page exists at a published address but did not render for review. That is an open research gap on our side as much as a gap in what is published, and it belongs on a procurement list rather than in a finding about Catch. [C-01, C-02, C-03] Town connects to work accounts, keeps memory, and runs routines. Its safety documentation, updated September 11, 2026, makes three statements. AI training is off by default and can be switched on from account settings. Google integration data is fully off-limits for training. No team member reaches the production database, except when you approve support access or an executive authorizes a one-time look at a single session during a security incident. Deletion is immediate as a soft delete, with permanent removal within 30 days, subject to legal and security exceptions. Its privacy policy, last updated July 21, 2026, still describes using personal information for research and development including to train its AI models, with an opt-out toggle in account settings. Two dated documents, seven weeks apart, describe the same control from opposite directions. The newer one is the trust page. Which one governs your contract is a question worth asking before signing it, not after. [T-05, T-01] Now the frontier providers, where the same brand carries different rules per surface. OpenAI runs consumer ChatGPT, managed workspaces, the API, apps, and agent mode. Commercial products and the API are not used for training by default. On the API, abuse monitoring logs are retained up to 30 days, zero data retention excludes customer content from those logs and forces the store parameter off, and both zero data retention and modified abuse monitoring require prior approval. Some endpoints are not eligible at all: conversations, agents, and assistants retain application state until it is deleted. Two named conditions sit alongside the strongest setting. Eyes Off content is excluded from human review unless applicable law requires it, which is a protection with a legal carve-out rather than a gap. Safety Retention runs the other way, and permits retaining and reviewing content that classifiers flag. Agent artifacts keep their own schedule: agent chats, browsing history, and screenshots persist until deletion, and removal can take up to 90 days. Agent mode reaches accounts and acts externally, and OpenAI's own material describes some, but not all, prompt injection mitigations, which is the honest way to say that layered defenses reduce the risk rather than removing it. [O-03, O-04, O-05, R-03, O-05] Anthropic runs consumer Claude, the commercial API, connectors, computer use, and Cowork. Commercial inputs and outputs are not used for training by default, and API data normally deletes within 30 days while Enterprise chat persists until deleted or governed by configured retention. Then capability changes the contract. Four models are designated Covered Models: Claude Mythos 5.1 and Claude Fable 5.1 from August 31, 2026, and Claude Mythos 5 and Claude Fable 5 from June 9, 2026. On those models, prompts and completions are retained for at least 30 days and then automatically deleted, and zero data retention is not available in workspaces, Enterprise organizations, or third-party platforms. Announced on September 1, 2026, Enterprise Frontier Safeguards will roll out in phases beginning in fall 2026, and eligible customers can use zero data retention with Claude Fable 5 and Claude Fable 5.1 for a limited time as a transition to it. Read that hinge carefully. Capability and retention are bound together here rather than set separately, and the announced architecture changes who holds the monitoring data rather than whether monitoring happens. [A-02, A-03, A-04, A-06, R-06] Google runs the consumer Gemini app, Workspace, the paid API, and Vertex AI. Qualifying Workspace editions do not use customer data for training without permission, and content is not human reviewed or used for training outside the domain without permission. Retention is a different question, and the schedules do not match. Administrators set Workspace prompt and response retention to 3 months, 18 months, 3 years, or indefinite. The Gemini app runs up to 36 months, with an 18-month default. With history off, new chats still persist for up to 72 hours. On the paid API, approved zero data retention is not the absence of all operational records. Four kinds survive it. Search grounding and Maps grounding each store prompts, context, and output for 30 days, with no opt-out. The Interactions API stores state unless you set store to false. A Live API session handle holds conversation state for up to 24 hours, and explicit caching persists to whatever expiry you set. [G-02, G-06, G-07] xAI runs Grok on X, standalone Grok, the API, Enterprise, and Grok Bot, and those are different governing terms, not one product. The API excludes inputs and outputs from training unless you give explicit permission, and separately retains requests and responses for 30 days for abuse auditing. Those are two facts, and readers routinely collapse them into one. Team-wide zero data retention removes durable persistence and disables eight things in exchange: per-API-key request logging, the stateful Responses, Files, Collections, and Batch APIs, deferred completions, stored image and video outputs, and voice agent conversation history. Grok Bot adds a persistent cloud computer with a browser, filesystem, and terminal, plus configurable approvals. [X-01, X-04, X-07, X-08, X-05] What each vendor documents about a named surface on a named date is the most a buyer can verify from outside. Where a company publishes less than another, that is a question to ask, not a finding to publish. ### Diagram. Where the published record stops One cell per surface per control, read straight off the comparison table further down this section. A green cell means a published answer on a named page. An amber cell means the answer is not published, which is a question for the vendor rather than a finding about them. Paciva publishes this report and appears in it on the same terms. *Source: the same vendor pages as the comparison table below, on the dates each cell there states. A missing answer is a research finding about what is published, never a claim about what happens.* > Read the amber cells as your call list. Every one of them has an answer. It just is not on a public page. ### Exhibit 2. Seven vendor surfaces, four controls, one date on every cell Compare one control at a time. This is not a score. Cells marked Read carry the date the page was fetched for this report. Cells marked Verified publish at their ledger date and were not re-fetched. A missing answer renders as a procurement question, never as an accusation. Paciva is listed on the same terms as everyone else, and Paciva publishes this report. Leaving ourselves out of a table we built would have been the dishonest option, so the row is here, sourced to public documents anyone can check, and marked. | Surface | Training default | Retention | Human review | Action authority | |---|---|---|---|---| | Viktori Slack assistant | No foundation model training on customer data, first or third party. Internal service-specific quality models are named as the exception. Policy dated July 31, 2026. Read September 12, 2026. | About 30 days from active systems, about 35 days for encrypted backups. A vendor blog states seven business days; both are published here. AI providers may temporarily retain data under their own API retention policies. Read September 12, 2026. | Encryption and access controls are documented. No published rule for customer content. Ask. Verified August 25, 2026. | Approval documented for money movement, code pushes, and customer email. Verified August 25, 2026. | | Catchi Inbox agent, email, text, calls | AI vendors contractually prohibited from retaining or training. Catch itself uses data to improve the experience for you and others. First-party improvement mechanism undefined. Ask. Read September 12, 2026. | No numeric schedule on the security page. A separate privacy policy page did not render for review. Open research gap. Ask. Read September 12, 2026. | No rule on the security page. The separate privacy policy page did not render for review. Open research gap. Ask. Read September 12, 2026. | Minimum permissions per integration, revocable at any time. Actions within granted permissions and with approval. Read September 12, 2026. | | Towni Work assistant with memory and routines | Safety docs: off by default, opt in from settings, Google integration data fully off-limits for training. Privacy policy: research and development including AI training, with an opt-out. Two dated documents. Ask which governs the contract. Safety docs September 11, 2026. Policy July 21, 2026. | Soft delete immediate, permanent removal within 30 days, legal and security exceptions. Read September 12, 2026. | No production database access except user approved support, or an executive authorized one-time look at a single session during a security incident. Read September 12, 2026. | Read-only, approval required, and advance-authorized autonomous modes. Verified August 25, 2026. | | OpenAIi API, commercial products, agent mode | Commercial products and the API are not used for training by default. Verified August 25, 2026. | Abuse logs up to 30 days. Zero data retention excludes content and forces store off, by prior approval. Conversations, agents, and assistants are ineligible and keep state until deleted. Agent chats, browsing history, and screenshots persist until deletion, and removal can take up to 90 days. Verified September 11, 2026. | Eyes Off excludes human review unless applicable law requires it. Safety Retention permits review of classifier flagged content. Verified September 11, 2026. | Agent mode reaches accounts and acts externally. Vendor material describes some, but not all, prompt injection mitigations. Verified September 11, 2026. | | Anthropici API, Enterprise, connectors, computer use | Commercial inputs and outputs are not used for training by default. Verified August 25, 2026. | API normally deletes within 30 days, verified August 25, 2026. Four Covered Models retain prompts and completions for at least 30 days, and zero data retention is unavailable in workspaces, Enterprise organizations, or third-party platforms. Covered Models read September 12, 2026. | Enterprise Frontier Safeguards, announced September 1, 2026, is described as changing who holds monitoring data, in phases from fall 2026. No separate human review rule on the reviewed pages. Ask. Read September 12, 2026. | Connectors and computer use add stored context, screenshots, visible applications, external actions, and configurable approval boundaries. Verified August 25, 2026. | | Googlei Workspace, Gemini app, paid API, Vertex AI | Qualifying Workspace editions: no training outside the domain without permission. Paid API traffic excluded from product improvement. Verified September 11, 2026. | Workspace 90 days to indefinite as administrators set it. Gemini app up to 36 months, 18-month default, 72 hours with history off. Search and Maps grounding retain 30 days with no opt-out. Verified September 11, 2026. | Qualifying Workspace content is not human reviewed outside the domain without permission. Consumer surfaces carry their own review and exception rules. Verified September 11, 2026. | Vertex eligibility and controls vary by model, feature, grounding, cache, state, and logging. No single action-authority rule spans these surfaces on the reviewed pages. Ask per surface. Verified August 25, 2026. | | xAIi API, Grok surfaces, Grok Bot | The API excludes inputs and outputs from training without explicit permission. Consumer Grok surfaces carry separate terms. Verified August 25, 2026. | 30 days for abuse auditing on ordinary API use, verified August 25, 2026. Team-wide zero data retention removes durable persistence and disables eight documented features. Feature list read September 12, 2026. | Consumer terms name human, safety, legal, feedback, and de-identification exceptions. No API specific rule on the reviewed pages. Ask. Verified August 25, 2026. | Grok Bot adds a persistent cloud computer with a browser, filesystem, and terminal, plus configurable approvals. Bots for one user share that computer. Verified August 25, 2026. | | Pacivai Pax, the assistant this report's authors build | A contractual non-training commitment, with a published per-provider table. The Gemini free tier is excluded at the routing layer because it may be used for training, and xAI is listed as still being confirmed. AI Use Statement v0.1.0, effective January 1, 2026. Read September 12, 2026. | A published schedule by data category. OAuth tokens purged within 24 hours of disconnect. After termination, a 30-day export window, deletion from active systems within 30 days, and from backups within 90 days after that. Voice runs three tiers: never stored, 24 hours, or a customer-set 30-day default. Decision Log runs 30 days in-app and 6 years as an audit trail. Data Retention Schedule v0.1.0, effective January 1, 2026. Read September 12, 2026. | A GDPR Article 22(3) route. A customer or an affected third party can contest a Pax Action and demand human review by a Paciva employee, with a written disposition within 10 business days. Read September 12, 2026. | Draft and review by default. Sends and write actions pass a Confirmation Gate. Waivers are tool-scoped rather than blanket, and every action is written to a Decision Log. Read September 12, 2026. | *Source: each vendor's own privacy policy, security page, trust documentation, and API documentation, on the dates shown. The Paciva row is sourced to the Paciva Trust Center in the same way, and readers should weigh it knowing the authors work there. This is a record of published statements about named surfaces, not an independent control test, and that limit applies to the Paciva row first.* > Six operations get treated as one and are not: opting out of training, revoking an integration, canceling a subscription, deleting a source, deleting the service, and deleting the state derived from it. Each has its own trigger and its own exceptions. [R-11] ## 3. Privacy has a capability price **The strongest privacy setting is not always the route that can finish the job. That is a design constraint to work inside, not a reason to send more data quietly.** The tradeoffs are documented, not hypothetical. On Bedrock, retention is a policy you set, and a model that requires retention reports itself unavailable under an insufficient mode; with the account or project set to none, invoking such a model is blocked and returns an error rather than quietly downgrading your setting. On Covered Models, the retention rule is the price of the capability, qualified by eligible interim zero data retention and the announced Enterprise Frontier Safeguards path. On xAI, the team-wide zero data retention mode costs eight named features, including every stateful and batch surface. On Vertex AI, eligibility moves with grounding, state, cache, model, and monitoring choices, so two teams on the same provider can hold different contracts. [H-01, A-04, A-06, X-08, G-04, H-06] Every entry above is a number or a named list in vendor documentation, which means procurement can ask for it in writing. A vendor that documents which features a retention setting disables has already told you what your team loses on the day you switch it on, and comparing that list against what your workflows depend on turns the risk into a number instead of a worry. Ask which of those features your current workloads touch. Ask what happens to a request when the setting and the model disagree, because the difference between a blocked call and a quietly downgraded one is the difference between a policy and a preference. Ask what event forces a requalification. The first two answers usually sit in the vendor's documentation already, which leaves one question that needs a person. [R-10] Local placement carries a second price. A customer-controlled model changes custody, and whether it still meets the task's quality, speed, context, and tool requirements is a question the task has to answer rather than a given. Sanitization carries a third. Removing identity and relationships can remove the meaning the work depends on, and a summary that no longer knows who is angry with whom is not a cheaper version of the task. It is a different task. [L-05, R-08] So the decision rule is narrow: use the minimum disclosure and authority that still meets the task's capability, quality, and authority requirements. And the failure path has to be stated out loud before anyone is under deadline pressure. If sealing damages the utility the work requires, the answer is a qualified controlled model, a person, or a blocked route. It is never a quieter external fallback. [R-08] None of that can be settled by reputation. Whether a sealed form still supports the work is a task-specific question. Run the same task both ways, compare the output, and decide one workload at a time. Nobody can answer it for you in advance, and it is exactly the claim that has to be tested rather than assumed. [R-08] A route is not useful merely because it is private, and it is not qualified merely because it is capable. Both have to be true for the task in front of you. [R-06] ### Exhibit 3. Four places where a privacy choice trades against a capability Each entry quotes vendor documentation rather than an estimate. The closing line of the Vertex entry is this report's synthesis, not a vendor statement. - **Team-wide zero data retention.** Removes durable persistence of prompts and outputs across the team. - **Bedrock retention mode set to none.** One of five account or project modes: none, default, aws_review, provider_data_share, and inherit. - **Covered Models.** Claude Mythos 5.1 and Claude Fable 5.1 from August 31, 2026. Claude Mythos 5 and Claude Fable 5 from June 9, 2026. - **Approved Vertex AI zero data retention.** Granted against a configuration, not against a company. *Source: xAI API security FAQ, Amazon Bedrock data retention and abuse detection documentation, Anthropic Covered Models retention, Vertex AI zero data retention. Read September 12, 2026, except Vertex, verified August 25, 2026.* > The price is published. Not one of these costs is an estimate, which is what makes them usable in a contract conversation. ## 4. Watermarking answers the output question **Provenance marking is a real mechanism pointed at a real problem. It is pointed at the wrong end of the pipe for privacy.** Some providers mark what their models produce. Anthropic has announced watermarked text in future Claude models, with older models covered over time. The signal carries no identifying information and traces to no person, organization, or chat. Detection weakens on small samples and on factual passages, where the model has fewer word choices to encode. Google documents SynthID across supported image, audio, text, and video, including text in the Gemini app and web experience, which is not a claim about the Gemini API. OpenAI documents C2PA and SynthID for supported images and SynthID for supported audio, with no deployed text watermark claim in the reviewed material. Grok documents a visible mark on generated images and video, and nothing invisible was found. [W-01, W-02, W-03, W-04] The asymmetry is the part worth keeping. A positive signal can support a likely origin on a compatible path. A missing signal establishes nothing, because text can be rewritten, translated, shortened, or produced by a model that marks nothing at all. A mark does not prove truth, ownership, integrity, sole authorship, or who saw the prompt. [W-05] Disclosure duties are where provenance does real work. The EU transparency obligation in Article 50 applies from August 2, 2026, and the Commission's guidance describes a transition to December 2 for the Article 50(2) marking and detection duty on systems already in service before August 2, which is narrower than a general delay. Exact applicability depends on the statute and on use-specific exceptions, which makes it a question for counsel rather than for a routing policy. For a buyer, that turns provenance into a narrow procurement item rather than a privacy control. Ask which modalities a provider marks today. Ask whether the mark is visible or statistical, and whether detection is available to you or only to the provider. Ask what happens to the signal when content is edited, translated, shortened, or passed through a second model. Then file the answers somewhere other than your privacy file, because they answer a different question. Disclosure duties, content policy, and platform labeling all draw on provenance. Confidentiality draws on retention, review, route, and authority, and not one of those four appears anywhere inside a watermark. [W-06, W-05] ### Exhibit 4. Which end of the pipe the mark sits on A watermark is applied to what the model writes. Nothing about that touches what the model read. - **Anthropic.** Announced statistical text watermarking for future Claude models. No identifying information, and weaker on small samples and factual passages. - **Google.** SynthID across supported image, audio, text, and video, including text in the Gemini app and web experience. Not a claim about the Gemini API. - **OpenAI.** C2PA and SynthID for supported images, SynthID for supported audio. No deployed text watermark in the reviewed material. - **xAI.** A visible mark on generated images and video. No invisible, C2PA, or text watermark claim was found. - **Input, before the model.** Who received your prompt; How long it is retained; Whether a person can read it; Whether it trained anything; Who you are - **Output, after the model.** A likely origin on a compatible path; A disclosure duty; Truth, ownership, or integrity; Sole authorship; Anything at all when the signal is absent *Source: each provider's own provenance documentation, verified between August 25 and September 12, 2026. A negative research finding means the claim was not confirmed in the reviewed sources on the review date.* > Watermarking can answer part of "where did this output come from?" It does not answer "who saw the input?" ## 5. A pre-egress sealed hyperscaler route **Follow one task from private ingress to private return. This is a qualification exercise on paper, not a tested deployment.** A customer renewal briefing contains a person's name, email address, and phone number, an account number, order history, a commercial complaint, and a confidential pricing concession. The team wants frontier reasoning on it. The first two questions are whether frontier capability is genuinely required and whether the task still means anything once identity is removed. The policy position is fixed before the work starts. Recognized plaintext PII does not go to a hyperscaler. Recognized means detected within the declared entity taxonomy, languages, modalities, and parser support, which is a bounded claim rather than a promise about every possible identifier. Unsupported input, incomplete extraction, residual recognized PII, or lost task utility sends the work to a controlled local model, human handling, or a blocked route. It never triggers a quieter external fallback. [R-08] Now the part that makes the ordering matter. Provider side filtering is real and useful, and it is not the first boundary. AWS documents that its sensitive information filters can block or mask supported PII, and documents exactly where the original survives: the input field in CloudWatch model invocation logs always contains the original unmodified request regardless of guardrail intervention, and the match field returned in guardrail trace output contains the original PII value rather than the masked one, by design, so an application can use the detection result. In tool use workloads, PII the model writes into tool call arguments, PII in tool results your application returns, and PII in the tool definitions you supply are neither blocked nor masked. That is defense in depth, not proof of pre-egress removal. [H-03] Three hyperscalers publish controls worth pinning, and each leaves a qualification behind. AWS Bedrock offers five retention modes, IAM, private networking, encryption, logging controls, and guardrails. The legacy provider sharing mode does no provider sharing today. Microsoft Foundry offers stateless inference excluded from base model training, geography choices, private networking, managed identity, role-based access control, encryption, and modified abuse monitoring. Stateful Responses and Assistants store history, and flagged samples can be human reviewed. Google Vertex AI promises no training without permission and offers configurable abuse monitoring exceptions, cache controls, residency, customer managed keys, service controls, and access transparency. Eligibility stays specific to the model, feature, grounding, state, cache, logging, and account configuration. [H-01, H-02, H-04, H-05, G-04, H-06] Run the renewal brief through all six gates and look at what the external provider holds. A request written about a customer, with opaque tokens where the name, the email address, the phone number, and the account number used to be, and an instruction to draft a renewal position. The provider can reason about the commercial history, because the history is what the reasoning needs. It holds no recognized plaintext identifier, because the map that would restore them never left the local system. What it may still hold is a rare fact, a date, or a role specific enough to point back at a person, which is what gate three exists to test. The draft comes back, gets inspected for leaked identity and unknown tokens, gets rehydrated locally, and lands in front of a person who decides whether to send it. What the route removed was recognized plaintext identity. Whether it kept the capability the task needs is the other thing gate three tests, one workload at a time. [H-07, R-08] Write "without sending recognized plaintext PII," never "without PII risk." Detection has a supported scope and a miss rate, local devices and maps and logs and backups remain sensitive, and confidential non-PII can still stop the send. > If the provider performs the first redaction, the provider received the thing you meant not to send. ### Exhibit 5. Six gates, and what each one is allowed to fail into Author proposed route. AWS documents filtering limits, not validation of this design. Failed or uncertain checks stay controlled, human handled, or blocked. - **Sealed payload.** Opaque task-scoped tokens where supported identifiers were, plus the instruction. Draft-only authority. - **Rehydration map.** Stays local. Kept out of prompts, traces, analytics, crash reports, backups, and tools. - **Controlled or human only.** Unsupported input, incomplete parsing, residual recognized PII, unresolved confidentiality, or damaged utility. - **The send.** Rehydration happens locally after verification, and only for the authorized destination. - **01. Qualify the task inside the private boundary.** Record source, owner, purpose, sensitivity, quality obligation, requested action, and prohibited content before anything moves. - **02. Seal and verify before egress.** Replace supported identifiers with opaque, task-scoped tokens, and keep the rehydration map out of prompts, traces, analytics, crash reports, backups, and tools. Reject unsupported media, incomplete parsing, residual recognized PII, unknown tokens, or malformed schema rather than proceeding. [P-01, P-02] - **03. Test residual confidentiality and utility.** The sealed payload can still carry the pricing concession, and a rare fact can still identify a person with every name removed. Separately, confirm the model can do the work without the relationships you just removed. If either test fails, the task stays controlled or human only. [L-04, R-08] - **04. Pin the exact route.** Provider, model, version, endpoint, account, region, retention mode, grounding, state, cache, logging, safety exception, and contract. This is the level at which zero data retention actually exists, because it is assembled from all of those rather than granted by a brand. Requalify after any material change, and block silent fallback. [H-08, R-10] - **05. Limit authority.** Only the required source, destination, tool, and action. The renewal case is draft only. Sending the message requires a person. [R-07] - **06. Verify, rehydrate, and record inside the private boundary.** Inspect raw output for PII, for direct sealed values, for unknown tokens, for prompt leakage, and for schema failure. Only after verification succeeds may identity be restored locally, and only for the authorized destination. Failed or uncertain verification keeps the output controlled or stops the task without restoring identity. [P-03] *Source: Bedrock sensitive information filter documentation for the provider side limits, and repository inspection at a pinned commit for the control shapes. Repository inspection supports code shape, not deployment, coverage, performance, or production maturity.* > Design conditions, not demonstrated outcomes. If every check passes, the hyperscaler sees tokens rather than recognized plaintext identity, and capability, privacy, and authority have been qualified together rather than one at a time. [H-07] ### Exhibit 6. Sealing, demonstrated in your own browser Type or edit the text on the left. Detection runs in this page, and nothing you type is sent anywhere, which is the point the exhibit is making. *Pattern-level detection over names, email addresses, phone numbers, and account identifiers only. Closing the tab ends it.* > Look at what is still there. The tokens took the person, including the bare first name later in the same paragraph. The 12 percent concession crossed the boundary untouched, because no identifier appears anywhere in it. That is what gate three is for, and it is why removing personal data is the floor rather than the finish line. ## 6. PII is the floor. Custody is the choice **Delete the names from an acquisition plan. It remains an acquisition plan.** Direct identifiers are the part everyone knows to remove. The rest of what matters in professional work has no name in it at all: inventions, source code, designs, pricing, negotiating position, legal analysis, unpublished research, contracts, and client material. Masking is necessary and it is not sufficient, because rare facts, specific dates, locations, roles, and relationship graphs re-identify people and expose strategy on their own. [L-04] Legal categories stay distinct while you do this. Trade secrets, privilege, contractual confidentiality, and client records carry different duties, and external disclosure can affect protections as well as obligations. WIPO's guidance is to keep trade secrets under organizational control where possible, limit access, mask or obfuscate data, and use private instances with contractual use and destruction safeguards when an external model is genuinely needed. Where a duty is in play, that is a question for counsel rather than for a routing policy. [L-02] Open models enter here, and the vocabulary needs care. The Open Source Initiative's definition requires the freedoms to use, study, modify, and share, along with the information needed to modify the system. Weights you can download do not by themselves meet that bar. Open weights and open source are different claims, and neither one is by itself about your privacy. [L-01] The conditional advantage is real and narrow. A properly licensed model under your control can remove an external inference recipient from the path, but only when the runtime, telemetry, dependencies, network egress, storage, logs, and backups are controlled too. Local placement does not remove security, licensing, memory, tool, patching, deletion, or quality duties. It moves them to you. [L-05, L-05] A controlled deployment is therefore an operating envelope rather than a download, and the envelope has seven walls. Verify the license and artifact provenance against a trusted signed manifest. Block network egress by default and allowlist dependencies. Keep memory, logs, caches, and backups encrypted and bounded. Grant least privilege over sources, tools, and destinations. Defend against prompt injection and poisoned memory, because untrusted content reaching durable memory is a security problem rather than a quality one. Audition the model on the real task rather than on a leaderboard. Keep patch, rollback, reproducibility, and deletion procedures that someone has tested. That list is a recommendation rather than a standard, and NIST keeps data privacy, intellectual property, cybersecurity, and value-chain risk as separate categories for the same reason. [R-05, L-03] Placed properly, open models take their seat beside the other routes rather than replacing them: controlled local work where custody decides the outcome, and local detection and verification wrapped around sanitized frontier work, so the strongest reasoning stays available without identity traveling with it, with human handling or a block where no qualified path exists. The portfolio is the point, because a single route cannot satisfy every combination of sensitivity, capability, and duty. [L-05] ### Exhibit 7. What survives a perfect masking pass Assume every direct identifier is removed correctly. Here is what the recipient still holds. - **The strategy.** Pricing floors, concession limits, negotiating position, acquisition targets, and the reason behind a deadline. - **The invention.** Source code, designs, unpublished research, methods, and the specific failure that made the next version work. - **The relationships.** Who reports to whom, who is in conflict, who was excluded from the meeting. A graph re-identifies people without a single name. - **The rare fact.** One unusual date, role, location, or event can identify a person more reliably than their name would. - **The duty.** Privilege, trade secret protection, contractual confidentiality, and client obligations do not travel with the identifiers. *Source: WIPO trade secrets guidance on digital objects, NIST AI 600-1 risk categories, and report synthesis. Verified August 25, 2026. Guidance, not legal advice.* > Masking is a floor, not a finish line. Removing PII and running locally do not by themselves create compliance with anything. ## 7. Placement, authority, and evidence **Close with a decision system rather than an adjective. Choose where data may go, decide what the system may do, then record the evidence without creating a new sensitive archive.** Those conditions produce four routes. Controlled local for sensitive or confidential work that fits a qualified customer-controlled model and runtime; it stops or moves to a person when quality or controls fail. Sanitized external when frontier capability is needed, recognized PII can be sealed, residual confidentiality is cleared, and the exact route is qualified; it fails closed to local, human only, or blocked. Governed external for approved non-sensitive content on a pinned external route; it stops on policy drift, new sensitivity, or an unqualified feature. Human only or blocked when PII cannot be sealed safely, utility depends on identity, confidentiality is unresolved, or no route meets the duty; that decision does not get downgraded later because the deadline moved. Notice what those four routes have in common. Each one names its own failure in advance and in writing. That is the part organizations skip, and skipping it is what produces incidents, because a route with no documented fallback does not stop when it fails. Whoever is closest to the deadline widens it instead. Writing the fallback down before anyone is under pressure turns a judgment call made at speed into a decision already made calmly, by people who had time to think about it. [R-08] Placement settled, authority is still open, and it takes three levels chosen independently of where the data sits. Town documents read-only, approval required, and advance-authorized modes, which is the shape this takes in a shipping product, and advance authorization is delegated authority rather than an absence of authority. A local model with broad write access is not a conservative configuration. [T-04, R-07] The receipt itself stays minimal. Record the task's purpose and authority, the supported input and verification results, the confidentiality and utility decision, the exact route and the policy date, the outcome, the expiry, and the fallback. Distinguish provider claims from contractual terms, customer configuration, local observations, and independent tests, because those are four different kinds of evidence and they fail differently. The receipt must not contain plaintext PII, rehydration maps, credentials, secrets, or unnecessary source content. It records what was claimed and configured. It does not prove provider retention, deletion, human access behavior, or operating effectiveness, and no log establishes any of those. Expire it and requalify when the model, endpoint, feature, grounding, cache, connector, policy, or contract changes. [R-09, R-10] > Privacy does not require weak AI. It requires the strongest qualified capability, with less than everything. ### Exhibit 8. Route decider Answer four questions and the page walks the same gates the report describes. The receipt it writes contains no personal data by construction. *Source: the routing rules in this report, assembled from R-06 through R-11, H-08, and L-04 through L-05. Verified August 25, 2026. An operating recommendation, not a compliance determination.* > Two of the four routes are external. The one that starts with personal data reaches a provider only after sealing, and only because sealing happened first rather than because the provider promised something. ### Exhibit 9. Eight questions for the buyer Every one of these has an answer. Most procurement processes never ask for it, and the posture that results rests on published summaries rather than on terms. - **01.** What can the system see, remember, infer, connect to, and do by default? - **02.** Which exact product surface, model, endpoint, and feature rules govern this task? - **03.** Are no training, no retention, no review, no logging, and no action evidenced separately? - **04.** Which source, derived state, safety record, cache, screenshot, log, or backup survives deletion, revocation, or cancellation? - **05.** Does recognized plaintext PII leave before the first filter runs? - **06.** Can confidential non-PII, or damaged task utility, force controlled or human handling? - **07.** Which actions require fresh approval, and which can be authorized in advance? - **08.** What event expires the route and requires requalification? *Source: report synthesis across R-06 through R-11, H-08, and L-04 through L-05. Verified August 25, 2026.* ## Questions readers ask **Does a no training promise mean my data is private?** No. Training, transmission, retention, human review, and action authority are five separate controls with five separate answers. A vendor that promises no training has not promised no retention, no review, no egress, or no action. **Is zero data retention a company level setting?** No. It is a property of an exact route, assembled from the model, endpoint, feature, state, cache, grounding, logging, safety exception, account setting, and contract. Eligibility attaches to a configuration, so one provider can serve two teams under different terms. **Does running a model locally make it private?** Local placement can remove an external inference recipient. It does not remove supply chain, authority, memory, logging, licensing, or quality risk, and it moves those duties to you. **Is open source the same as open weights?** No. The Open Source Initiative's definition requires the freedoms to use, study, modify, and share, plus the information needed to modify the system. Downloadable weights alone do not meet that bar. **Can a watermark tell me who wrote something?** No. Anthropic states that the text watermark it has announced for future Claude models will carry no identifying information. A positive signal can support a likely origin on a compatible path, and a missing signal establishes nothing. **Is a provider side PII filter enough?** Provider-side filtering runs after the data has arrived, so it is a second layer rather than the boundary. AWS documents where originals survive in logs and in trace output, by design. **What happens when sealing breaks the task?** The work moves to a qualified controlled model, to a person, or to a blocked route. It never falls back to sending more data externally because the deadline moved. **Is removing personal data enough for business confidentiality?** No. Pricing, strategy, source code, unpublished research, and relationship graphs carry no names and still expose the thing worth protecting. Masking is the floor. ## Evidence base Provider policy supports the stated rule for the named surface and date, not operating effectiveness. Product documentation supports a bounded vendor claim, subject to the applicable contract and account configuration. An independent standard supports the cited framework, not a vendor implementation. Repository inspection supports the code shape at a pinned commit, not deployment, coverage, performance, or production maturity. Synthesis is analysis, presented as analysis. Contracts and configured account controls can supersede public defaults, and a negative research finding means only that a claim was not confirmed in the reviewed sources on the review date. - [NIST agent identity concept paper](https://www.nccoe.nist.gov/sites/default/files/2026-02/accelerating-the-adoption-of-software-and-ai-agent-identity-and-authorization-concept-paper.pdf) - [NIST AI 600-1](https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf) - [OWASP memory attack surface](https://genai.owasp.org/2026/05/13/memory-is-a-feature-it-is-also-an-attack-surface/) - [OSI Open Source AI Definition](https://opensource.org/ai/open-source-ai-definition) - [WIPO trade secrets guide](https://www.wipo.int/web-publications/wipo-guide-to-trade-secrets-and-innovation/en/part-vii-trade-secrets-and-digital-objects.html) - [European Commission Article 50 FAQ](https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act) - [OpenAI API data controls](https://developers.openai.com/api/docs/guides/your-data) - [OpenAI enterprise privacy](https://openai.com/enterprise-privacy/) - [ChatGPT agent](https://help.openai.com/en/articles/11752874) - [Anthropic Covered Models retention](https://privacy.claude.com/en/articles/15425996-data-retention-practices-for-covered-models) - [Enterprise Frontier Safeguards](https://www.anthropic.com/news/enterprise-frontier-safeguards) - [Anthropic text watermark](https://www.anthropic.com/news/claude-text-watermark) - [Google Workspace privacy hub](https://knowledge.workspace.google.com/admin/generative-ai/generative-ai-in-google-workspace-privacy-hub) - [Gemini API zero data retention](https://ai.google.dev/gemini-api/docs/zdr) - [Vertex AI zero data retention](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/vertex-ai-zero-data-retention) - [Google SynthID](https://deepmind.google/models/synthid/) - [xAI API security FAQ](https://docs.x.ai/developers/faq/security) - [Grok Bot approvals and privacy](https://docs.x.ai/grok-bot/approvals-security-and-privacy) - [Bedrock data retention](https://docs.aws.amazon.com/bedrock/latest/userguide/data-retention.html) - [Bedrock sensitive information filters](https://docs.aws.amazon.com/bedrock/latest/userguide/guardrails-sensitive-filters.html) - [Microsoft Foundry data privacy](https://learn.microsoft.com/en-us/azure/foundry/responsible-ai/openai/data-privacy) - [Viktor privacy policy](https://viktor.com/legal/privacy) - [Catch security and privacy](https://www.catchagent.ai/security) - [Town trust and safety docs](https://www.town.com/docs/safety) - [Paciva Trust Center, AI Use Statement](https://trust.paciva.ai/legal/ai) - [Paciva Trust Center, Data Retention](https://trust.paciva.ai/legal/retention) ## About this report Written from a fact-check ledger of 74 claims with published evidence classes and verification dates, and seven publication gates that name wordings this report is not permitted to use. Time-sensitive vendor rows were rechecked on September 12, 2026, and rows that were not re-fetched publish at their ledger date, which each cell states. Where official pages conflict, both statements are published with their dates and the reconciliation is named as a procurement question rather than resolved in the reader's favor. Paciva makes Pax, an executive assistant for individuals and a chief of staff to organizations. The routing architecture in this report is how that kind of assistant has to handle sensitive work. Frederick Townes builds Paciva, which is building toward a managed system for supported PII sealing, policy-based local and external model routing, fail-closed verification, and local rehydration. The architecture is the thesis here; universal coverage, performance, and production maturity remain separate proof obligations. P-04 Send the link, quote the exhibits, or take the buyer checklist into your next vendor call. Canonical HTML version: https://paciva.ai/resources/privacy-follows-capability/