Can AI agents act without being trusted?
An AI assistant that answers a question can make a mistake. An AI agent that reads files, delegates tasks and calls tools can turn a mistake into an action.
That difference matters as organizations move beyond chat interfaces. Agents can gather information, coordinate with other agents, prepare decisions and update systems. These abilities make them useful, but they also raise a practical question: who decides what an agent is allowed to do when it reaches the point of action?
The question is becoming harder to ignore. NIST has identified agent identity and authorization as areas requiring attention, while OWASP describes “excessive agency” as a risk that arises when an AI system has more functionality, permission or autonomy than its task requires.
The architectural approach is to separate the agent's reasoning from the controls that govern access and changes. Agents remain flexible in how they investigate a task, while their ability to act depends on authorization.
Autonomy needs a boundary
Consider an agent asked to evaluate suppliers and prepare an operational request within a defined budget. It may need to read a catalogue, compare options, ask a specialist agent to check constraints and produce a recommendation. Those steps benefit from flexibility. The agent might take a route its designers did not anticipate.
Submitting the request is different. That action has an external effect. It should depend on the requester’s authority, the available budget, the chosen supplier and the current approval rules. The agent’s confidence in its own recommendation cannot establish any of those facts.
The same distinction applies to less visible actions: writing a file, changing a database record, sending a message or invoking another service. An agent can propose an action. A separate control decides whether it is permitted. A constrained tool then carries it out.
This separation preserves the useful part of autonomy. The agent can still investigate and plan. It simply cannot make its own plan sufficient authority to act.
Identity and authorization are organizational controls
Agent security includes access-control problems that organizations already face. An employee may be correctly signed in yet have no reason to read another team's records, approve their own expense or export a customer list. A service account may be legitimate but far too powerful for the task it performs. Controls protect against mistakes, compromised credentials and deliberate misuse as well as model error.
Authentication establishes who is present; authorization determines what that identity may do in this context. The two questions must remain distinct when an agent acts on someone's behalf. A valid user session does not automatically authorize every action the agent can imagine. Nor should an agent borrow a broad service credential that hides which person initiated the work. The system needs to retain both identities: the accountable human or service principal and the agent that proposed the operation.
This also makes human oversight more precise. A reviewer should see the intended effect, target and relevant constraints before approving a consequential step. Approval of a task in general is not a permanent grant for every later tool call. Policy can permit routine actions automatically, require review above a threshold and reject actions outside the mission altogether.
The underlying ideas are established enterprise practice. NIST's zero-trust architecture treats access as a decision about a subject and a resource, enforced at a boundary rather than inferred from network location or a previous login. Agents make that discipline more urgent because they can discover and attempt actions that were not listed in the original user request.
Following one request through the system
For the supplier task, the permitted outcome is a draft based on the relevant catalogue. Submitting an order remains outside the mission.
The agent compares options and asks a specialist to check the budget. The specialist reads figures but cannot modify the draft. Its findings support the recommendation without giving either agent additional permission.
The primary agent then proposes creating the draft. A separate control evaluates user rights, task scope, destination and budget constraints. It can permit the operation, require a human checkpoint or halt the process if authorization fails.
If permitted, a constrained worker receives the approved inputs and destination, writes the draft and returns a receipt. The request, decision and observed result remain connected. If the agent proposes submitting the order when its mission permits only preparing a draft, the authorization service rejects the request as out of scope. No order is submitted, and the refusal is recorded alongside the proposal and policy decision. The agent can continue preparing the authorized draft.
From an idea to an accountable action
The supplier example rests on decisions made before the agent starts work. Human governance defines the mission, its beneficiaries, acceptable outcomes and boundaries, together with budgets, operating limits and who may stop or revise the task. Governance assigns responsibility and defines the agent’s role and delegable capabilities. Identity services authenticate the requester. These constraints must be enforced by the platform, independently of prompt guidance.
The harness is the agent’s working environment. It supplies instructions, memory, task context and permitted tools, allowing the agent to investigate, coordinate specialists and revise its plan.
The authorization engine applies policies that define which operations and resources the verified principal may use for the mission. Those policies remain outside the agent’s control when it revises its plan.
The execution environment exposes only the resources needed for the authorized operation. For a draft, that could mean a fixed output location, limited file access, no purchase endpoint and a short execution window. More sensitive operations may need stronger sandboxing, tighter network access or a human checkpoint before execution.
Delegation and shared infrastructure make these boundaries harder to maintain, as the following sections show.
Why multi-agent systems make this more important
Delegation can make the supplier task more capable, but it also makes authority harder to follow. Each specialist should receive a defined subtask, only the data and tools it needs, and a clear expiry or completion point. Creating another agent must not expand the authority granted to the mission.
Information needs similar care. A catalogue returned by a tool may contain text that looks like an instruction. Another agent may confidently report an incorrect conclusion. The system should retain the origin of that material and distinguish data supplied for analysis from instructions issued by an authorized person. Prompt injection matters precisely because that distinction can otherwise disappear inside an agent's working context.
The harness can validate proposal formatting and flag suspicious content, but cannot self-authorize a request. A specialist's finding remains evidence. It cannot expand permissions or establish whose data may be used.
One platform, several tenants
An organization rarely has a single undifferentiated pool of users and data. Teams, subsidiaries and customers may share an agent platform while expecting strict separation. In a multi-tenant service, each tenant has its own documents, policies, credentials, budgets and records. A useful agent must know whose context it is operating in without gaining access to everyone else's.
Consider two customers using the same supplier evaluation agent. Customer A's proposal must not incorporate Customer B's catalogues, budget figures, conversation history or cached search results. A specialist created during A's request must not carry A's access into a later request for B. Even when the model and infrastructure are shared, the authority and data context must remain distinct.
Tenant identity therefore cannot be a label supplied by the model in a tool argument. It must come from a verified user or service identity and current membership. It has to travel through delegation, retrieval, tool calls, asynchronous work and audit events. At every boundary, the system checks that the resource actually belongs to that tenant and that the acting principal may use it. Data layer controls can reinforce those checks, but an internal service call is not automatically safe.
Isolation also has an availability dimension. One tenant should not consume all shared model calls, worker capacity or budget. Tenant aware quotas and usage records make both cost and operational impact accountable. OWASP's multi-tenant security guidance highlights these same concerns across authorization, storage, caches and shared capacity.
This is an architectural requirement, not a claim that a single user prototype has already solved multi-tenancy. Adding a tenant label at the end cannot repair data that has crossed the wrong boundary. The execution path must enforce separation when the proposal becomes an action.
What sits beneath the boundary
With mission, delegation and tenant context established, action proposals pass through a gateway to a policy decision point. It evaluates verified context: the principal, agent, tool, resource, tenant and operational limits. The corresponding enforcement point blocks or permits the call at a mandatory boundary.
For the supplier request, the proposal can be a JSON object containing the operation create_draft, the destination resource, supplier identifier, amount and idempotency key. The gateway validates its schema and resolves the destination against trusted resource metadata. It attaches the authenticated principal, tenant membership and delegation scope from server-side context. Those authority fields must not become trustworthy merely because the model included them in its arguments.
The resulting decision is bounded by purpose, resource, duration and policy rules. Permission to create one draft for Tenant A is not a general purpose pass to write elsewhere or place an order. The isolated worker must verify that its actual operation matches the decision before executing.
The decision should identify the authorized operation and resource, the relevant parameters, its expiry and the policy revision used. A worker on another service can receive a signed decision or retrieve a protected decision record. In either case, it verifies the binding to the actual request. A signature protects the decision against alteration. It does not establish that the mission is still active or that the resource has not changed.
The agent can adapt its plan while each consequential step is evaluated against rules it cannot modify. Unavailable policy services, missing identity or ambiguous resources produce no implicit permission. A manager's instruction cannot override restrictions beyond their authority.
At the point of effect, the resource service should recheck that the mission and authorization remain valid, and validate the actual inputs. Where supported, it should combine this check with the mutation. Resource version checks can detect changes since approval. If authorization and mutation cannot be coordinated atomically, a race remains and must be handled explicitly. Outputs should be checked and recorded after execution.
The execution path must also prevent a direct route around those checks. Otherwise, a well designed policy service would govern only the calls that happen to use it. Restricted file access, network access and tool credentials must therefore support the same boundary as the gateway.
Uncertain outcomes must stop the workflow until they have been reconciled. An idempotency key needs durable enforcement. Store it with a digest of the request and the operation's result. Enforce uniqueness in the durable store and reject reuse with different parameters. For a database write, committing that record with the mutation avoids a separate bookkeeping gap. For an external service, use its supported idempotency mechanism and reconcile uncertain outcomes. A successful worker retry alone cannot prove that only one effect occurred.
A useful record links the original request, delegation, evidence used, proposal, decision, execution and observed result. It should make refusals as visible as successful actions. Otherwise, teams cannot tell whether a control worked or an agent simply never tried the prohibited path.
Audit is not a substitute for prevention. It is how operators investigate incidents, verify that policy was enforced and improve the system. Continuous tests should cover normal work, unauthorized access, misleading data, service failures and attempts to repeat an action. An uncertain result after a crash must remain visible in the record rather than becoming an invented success.
These controls do not require an infallible model. They prevent the agent’s reasoning or retrieved content from becoming authority on their own. Their effectiveness still depends on correctly scoped identities, policies and enforcement. An authorized action may still be factually wrong or unsuitable, so validation and human review remain necessary where the consequences warrant them.
Implementing the controls with existing tools
Agent Fortress applies this control pattern in a local proof of concept built from existing tools and small adapters. It uses a deliberately simple scenario: preparing one pizza-order draft from synthetic data.
Keycloak authenticates the operator through OpenID Connect with Authorization Code and PKCE. A small adapter connects that login to AnythingLLM’s Simple SSO, and AnythingLLM calls Hermes through its OpenAI-compatible API. Hermes coordinates bounded specialist tasks, with hooks restricting each role’s tools and delegation. This is a single-operator integration: session headers and agent labels do not establish a distinct cryptographic identity for every delegated agent.
Agentgateway enforces the LLM and MCP entry points. The model route uses a service key. The tool route checks a Keycloak JWT and restricts the available tools. External authorization calls a separate Python policy decision point (PDP), combining AGT’s agent-control-specification runtime with Open Policy Agent (OPA). AGT applies the manifest contract, OPA evaluates Rego rules against structured proposals. The service validates JSON schemas and returns decisions with scope, expiry and policy version. Missing authorization or unavailable policy services cause refusal.
The gateway also inspects messages and tool outputs with its native regex guards and blocks a tested secret pattern in model responses. These checks catch the demonstration’s hostile instruction and secret-shaped output. They establish a specific inspection path, with no claim of general resistance to prompt injection or comprehensive data loss prevention.
Execution belongs to a narrow Python worker. Before writing the fixed draft destination, it validates the proposed result, checks budget and allergies, and obtains a fresh PDP decision. An operator can persistently revoke the task, and the worker records an idempotency key and draft digest to recognize retries and reject conflicting reuse. An uncertain receipt stops execution. The local file write is atomic, but authorization, the file and its receipt are separate operations: revocation and crash windows still require explicit handling.
The worker runs in Docker as a non-root process, with a read-only root filesystem, limited mounts, no published port and resource caps. The optional Helm deployment runs the core services in kind under WSL, using Calico NetworkPolicies to deny traffic by default and permit defined flows. In the Kubernetes profile, Agentgateway has no direct egress permission on port 443. Provider traffic goes through the broker, whose network rule allows any destination on that port. The Compose profile retains an egress limitation. These controls do not demonstrate microVM isolation or an Anthropic-only network allowlist.
PostgreSQL adds row level security for structured task data, with policies exercised through a restricted runtime account. That test is a separate data-layer demonstration: the MCP tools currently read JSON fixtures. Serving multiple tenants would require verified tenant context throughout identity, delegation, storage and enforcement, beyond this single-user setup.
In the latest Kubernetes profile, an internal Anthropic broker retrieves the provider key for each request from Keycloak’s file based Vault through a custom, access-controlled extension. Hermes and Agentgateway do not receive that provider key. The broker holds it temporarily in memory and uses internal HTTP, so this remains a local prototype. The real provider was reached, but generation is still unverified because the required workspace identifier was missing.
OpenTelemetry and Tempo support traces. Loki and Alloy collect logs. Prometheus and Grafana provide operational visibility. Application audit events correlate the request, policy decision and worker result, including refusals. MinIO Object Lock adds retention for archived evidence, with a historical retention test that was not repeated in the final delivery run. Collection remains best effort, and local retention does not protect against an administrator controlling the host. Optional Falco tests detect a worker syscall canary. They do not establish a supported host-wide deployment under WSL.
The report also considers Cedar, LiteLLM, HashiCorp Vault, LLM Guard or NeMo Guardrails, Wazuh, and E2B with Firecracker/KVM. Those choices belong to the reference design rather than the deployed POC. In particular, E2B is not active: the WSL sandbox lifecycle test uses a Docker simulation. The implementation demonstrates a bounded control pattern while leaving stronger isolation and production assurance to be verified.
Is this too much architecture for an agent?
The controls should follow the effects an agent can produce and the infrastructure already available. Reading public information, accessing tenant data and changing external systems require different safeguards. Teams should first identify those effects, then strengthen the boundary as the consequences grow. This does not require a security decision for every sentence the model generates.
The first stage is visibility: identify the user and tenant, list the tools the agent can call, and record each attempted action. This often reveals overly broad credentials or tools that expose write functions to a read-only task. Removing unnecessary capabilities is a quick improvement even before a dedicated policy service exists.
The next stage is enforcement. Put a check on the routes that reach sensitive data or create external effects. Define rules in terms of concrete operations and resources: read this collection, create one draft, spend no more than this amount. Keep high impact approval with a person where the risk warrants it. An agent can continue working on safe parts of the task while a consequential action waits for a decision.
The final stage is assurance: test refusals as carefully as successes, exercise unavailable services and repeated requests, and verify that audit records connect decisions to effects. The controls should evolve when the agent gains new tools, serves another tenant or handles more sensitive data.
This progression avoids attempting to perfect every safeguard before the first deployment while retaining controls beyond prompt guidance or a login screen. The architecture grows at the boundary where the agent's capabilities grow.
The local proof of concept exercises a concrete effect: writing one draft file from synthetic data. Its recorded checks cover bounded delegation, denied tool calls, policy service failure, targeted hostile inputs, duplicate requests, revocation before the write and correlated audit events. It performs no purchase or other external business operation.
These results support the feasibility of the control pattern within the tested local scenario. Production readiness still requires stronger identity across delegated agents and tenants, verified network boundaries for the selected deployment, revocation and recovery tests around concurrent effects, reliable audit coverage and broader adversarial testing. The real provider path has reached Anthropic, but successful generation remains unverified. Those limits should remain visible alongside the successful checks.
These limits also matter when discussing sovereignty. Keeping policy, identity and execution under an organization's control improves governance. It does not automatically mean that all processing or data stays on premises. If an external model provider handles a request, that data flow still needs to be understood and governed.
Making useful autonomy governable
Agents need flexibility to investigate a task. Organizations need to make consequential steps identifiable, limited and reviewable without having to predict every step of the agent's reasoning.
An agent can explore several paths to a solution. It can coordinate specialists and adapt when new information arrives. But when it asks to cross a boundary into a real system, someone other than the agent must be able to answer: is this action authorized, within scope and safe to execute here?
This separation between reasoning and authority gives organizations a way to use capable agents while keeping decisions about access and real-world effects in accountable hands. Its effectiveness depends on the boundaries enforced by the implementation.