AGNTCON Europe 2026 Conference Report
GNTCon + MCPCon Europe 2026 was held in Amsterdam. It is a practitioner conference dedicated to building, securing and operating AI agents in production, and to the Model Context Protocol that most of them depend on. The programme is weighted towards teams that have already shipped: nearly every session was built around a system running with real users, and around the failures that surfaced once it was. Research contributions are rare; production war stories, reference architectures and post-mortems are the norm.
The speaker mix combined vendor platform teams, enterprise engineering teams and independent open-source maintainers. Several sessions ended by open-sourcing the tooling presented.
Keywords: AI agents, multi-agent systems, MCP, stateless MCP, WebMCP, agent identity, delegated authorization, identity broker, ReBAC, Zanzibar, AAuth, OAuth2 token exchange, AI control plane, agent observability, governance, agent sandboxing, autonomous remediation, software factory, agentgateway.
The Big Trends
Thirteen themes recurred across the two days, independently of the speaker or the vendor. They are listed roughly in order of how often they came up.
- Organising discoverability for agents. Agents discover their tools, data and peers at runtime rather than at build time, which makes discoverability an infrastructure problem rather than a documentation problem. The day-two keynote on delivery mechanisms opened on precisely this question, and registries of approved agents and MCP servers appeared repeatedly as the first building block of any governance story: you cannot govern what you cannot enumerate.
- The control plane is becoming the centre of gravity. Authentication and authorization were the entry point, but the discussion consistently widened. The control plane is taking on security, financial control (spend caps and cost attribution), policy management, and availability — including the question nobody has a clean answer to: when a model in the chain stops responding, what should happen to the workflow? This is related to observability but distinct from it. Observability tells you what happened; the control plane decides what is allowed to happen and what happens next even in case of failure or when running out of credits.
- Authorization must be scoped, never inherited. Delegating a user's full privilege set to an agent was treated as unacceptable by every security-oriented speaker. The disagreement is only about the mechanism. Solutions presented included the AAuth draft protocol, Zanzibar-style relationship-based access control, curated registries with per-agent identity, and MCP gateways enforcing scope at the call. Zalando presented a production identity broker issuing grants along a User → Agent → target → scope chain, with effective access defined as the intersection of what the user may do and what the agent is allowed to do. The shared-robot-account shortcut was explicitly named as an anti-pattern, because it pools the permissions of every user behind one agent.
- Multi-agent parallelism is distributed systems. Once agents run in parallel, they hit the exact problem set of distributed computing: synchronisation, waiting on a peer that has died, high availability, durable state, partial failure and recovery. Several speakers reached this conclusion from different starting points. The practical consequence is that the answers largely already exist in distributed-systems engineering, and agent frameworks are currently rediscovering them.
- The integration surface is unsettled: MCP, WebMCP or CLI. There is no settled answer on how agents should reach systems, and a full keynote was devoted to the three competing doors. An MCP server exposes tools over a client-server protocol and carries the context of the system it fronts. WebMCP makes the web page itself the tool provider, running in the browser tab with the user's session. A CLI leverages the terminal and the context of the working tree. The choice has real consequences for access control, observability and latency, and the clearest guidance offered all conference was a rule of thumb: use MCP when the call crosses an organisational boundary, and a direct tool call otherwise.
- The protocol layer is maturing, and it went stateless. The 2026-07-28 MCP specification removed the initialization handshake and the protocol-level session, which is the largest revision since the protocol launched. The motivation was operational rather than conceptual: sessions pinned a client to one server instance and forced sticky routing or a shared session store on anyone running a remote server. This was the subject of a keynote and is the single most consequential change for anyone operating MCP infrastructure.
- Governance, observability and active monitoring of agents. Monitoring final outputs is no longer sufficient. Governance, observability and active monitoring were presented as one continuum covering operational health, cost, compliance and audit, and quality. "Active" is the operative word: several speakers argued for intervening at runtime rather than reporting after the fact.
- Composability of agents, and the policy problem it creates. Composing agents inside one domain is now tractable. Composing them across domains is not, and the blocking question is policy: whose policy applies when an agent from one organisation calls an agent or tool belonging to another, and how are those policies reconciled at runtime? The Zalando session came closest to an answer by evaluating two gates per call — a delegation grant at the broker and a tool policy at the gateway — but only within a single organisation.
- Packaging agents for production. Getting an agent to work is now the easy part. Packaging it — dependencies, sandboxing, secrets, durable state, restart behaviour, registration, versioning and rollback — is where teams lose time. Zalando's answer was to make registration a declarative Kubernetes resource reconciled against the platform, so that product teams never maintain their own registration scripts.
- Determinism has to be reintroduced deliberately. A rule written in a prompt is a request, not an enforcement mechanism, and instructions get lost as the context window fills. Several speakers independently used near-identical phrasing: hooks and callbacks are how you put deterministic logic inside a non-deterministic flow. The software-factory session put it most bluntly as "fight slop with deterministic code".
- Coding agents scale poorly to large collaborative work. Coding agents driven by a prompt work well for small, local, well-bounded tasks. As soon as the work requires collaboration on something large — multiple modules, cross-cutting changes, shared context between several agents or several people — the approach becomes complicated and the context requirement grows faster than the benefit. This was the opening question of the software-factory session and the implicit premise of the vulnerability-remediation session.
- Rethinking the software factory. The consequence of the previous point: the software factory itself has to be rethought for an era of coding agents and autonomous agents. Branching strategy, review, CI gating, environment provisioning and merge policy were all designed around human throughput. Two sessions argued that the binding constraint is not whether an agent can produce a change but whether the surrounding system can absorb it without degrading.
- Evaluation and CI/CD remain unsolved. Because the component producing the output is non-deterministic, conventional regression testing does not transfer. No speaker claimed to have solved this, and several named it explicitly as open, particularly for integration into CI/CD pipelines. The SlopCodeBench results presented on day two sharpened the problem: agents pass checkpoint tests while the code underneath steadily degrades, so pass-rate testing actively conceals the failure.
Overall assessment of the event: the level was high and resolutely practical. Speakers presented systems in production with the bugs included, which is unusual and valuable. The recurring weakness was evaluation: almost every speaker acknowledged not having a satisfactory way to test their system, and no session filled that gap. The two days also complemented each other well — day one asked how to control agents, day two asked how to reach them and how to organise around them.
The Kubernetes community had a prominent presence at AGNT CON, bringing a strong open-source atmosphere to the event.
The Keynotes
1. Beyond MCP gateways: an AI control plane for securing and governing agents (Day 1)
Positioning
The keynote set the frame for the conference by asking how an organisation secures and governs AI agents it does not fully trust. Its central claim is that the MCP gateway — currently the default answer in most organisations — is a checkpoint, not a governance system. A gateway can intercept and filter tool calls, which is useful, but it leaves three questions unanswered: what agents exist, what they have actually done, and what happens when one misbehaves.
The proposal is to treat this as a control plane, in the sense the term carries in networking and Kubernetes: a central layer holding policy and configuration, separate from the agents executing the work. That framing is what made the talk a good opener, because it is the structure into which most of the other sessions fit — including Zalando's platform on day two, which is the same architecture built out inside one company.
MCP (Model Context Protocol): the open protocol through which agents discover and call external tools and data sources. A gateway sits between the agent and the MCP servers and can inspect, filter or block individual tool calls.
The three-layer model
- Curate — decide what exists. A registry of approved agents and MCP servers, plus an identity for each agent. Nothing enters the environment without being registered, which is what makes the inventory question answerable and connects directly to the discoverability trend.
- Monitor — see what happens. Hooks, filters and audit logs. Hooks and filters act at call time; audit logs make activity reconstructable afterwards. This is the layer that answers "what has this agent done?".
- Isolate — contain the blast radius. Agent sandboxes, a network proxy, and a credential store. The agent runs in a disposable environment, its egress traffic passes through a proxy enforcing destination policy, and it never holds raw secrets.
The three layers are complementary rather than alternatives. Curation without monitoring produces an inventory nobody watches; monitoring without isolation produces an excellent record of an incident you could not prevent.
Worked example: Discobox as the isolation layer
Discobox was presented as a concrete coding-agent sandbox. It gives an agent its own disposable machine with passwordless sudo, a desktop and a browser, so the agent is not blocked by a permission prompt on every action. Several boxes run side by side, one per task, and work returns through git: each box clones the repository, the agent commits, and the commits are merged back or opened as a pull request.
What makes it relevant is that it implements all three isolation controls at once. Each environment is a container inside a virtual machine, giving two layers of isolation. All outbound traffic passes through a proxy using a per-box mTLS identity, with allowed destinations set by policy and every request logged and attributed to one box. Secrets never enter the environment: the box holds a placeholder and the proxy substitutes the real credential only on requests to the domain that credential is bound to. Privileged requests are checked against grants written in plain language before a credential is issued, and no credential is issued without a logged verdict.
Project: https://discobox.ai/
What I learned, and how much innovation it brings
The individual components are not new. Registries, audit logs, egress proxies and credential brokers are standard enterprise security building blocks. The contribution is the assembly: arguing that these have to be one product rather than four tools bolted together, and that the agent case makes the assembly urgent because agents choose their own actions at runtime.
The keynote was vendor-adjacent and the control-plane framing serves a product category, so parts of it read as a category pitch. That said, it did not oversell: the speaker was explicit that the gateway alone is insufficient, which is an argument against the simpler product most vendors sell. Relevance is high, and the three-layer model is directly reusable as a checklist for assessing internal maturity.
2. Stateless MCP: the 2026-07-28 specification
Background: MCP and the stateful client-server model
The speaker opened by revisiting the MCP architecture. MCP is a client-server protocol. A host application — a chat client, an IDE assistant or an agent framework — runs an MCP client, and each MCP server exposes tools, resources and prompts that the client can list and invoke. The protocol runs over two transports: STDIO, where the server is a child process on the same machine, and Streamable HTTP for remote servers.
Until this year the protocol was stateful. Every connection opened with an initialize and initialized handshake in which client and server exchanged protocol version and capabilities, and over HTTP the server issued an Mcp-Session-Id header that every subsequent request had to carry. That session pinned the client to whichever server instance answered first, so running a remote MCP server at any scale meant sticky routing, a shared session store, or both. This is the model that worked well for one developer on one machine and broke on cloud-native infrastructure.
Streamable HTTP brought real advantages: remote connectivity, a single URL to connect and authenticate against, and simpler distribution through connectors and plugins. But three practical problems remained: shared state between client and server, connection management, and compatibility with ordinary HTTP infrastructure such as load balancers, proxies and caches.
The other problem is chattiness. The published Hugging Face figure is that a single tool call can generate over one hundred MCP protocol messages, most of it negotiation and lifecycle traffic rather than useful work. At Hugging Face's scale this directly affects the conversion rate between clients that connect and clients that go on to make a tool call.
[The session notes record roughly 73 messages per tool call. The figure published in Hugging Face's own material is "over 100".]
What stateless MCP changes
The 2026-07-28 specification makes the protocol core stateless: every request is self-contained and any server instance can answer it. The changes the speaker highlighted:
- No handshake. The initialize and initialized handshake is removed (SEP-2575). Protocol version, client info and client capabilities now travel in a _meta field on every request, and version mismatches return an explicit error rather than failing at connection time.
- No protocol session. The Mcp-Session-Id header and the protocol-level session are removed (SEP-2567). List endpoints no longer vary per connection. Servers that genuinely need cross-call state now mint an explicit handle and the client passes it back as an ordinary tool argument — the protocol is stateless while the application stays as stateful as it needs to be.
- Discovery on demand. A new server/discover method, which servers must implement, advertises supported protocol versions, capabilities and identity. Clients may call it before anything else for up-front version selection, or use it as a backward-compatibility probe.
- Cacheable lists. tools/list, resources/list, prompts/list and resources/read now carry ttlMs and cacheScope hints (SEP-2549), with deterministic ordering. Clients can cache tool catalogues instead of re-listing everything on every reconnect.
- HTTP-native routing. Mcp-Protocol-Version, Mcp-Method and Mcp-Name are promoted to HTTP headers mirroring the JSON-RPC body, so gateways, proxies and load balancers can route, rate-limit and audit without parsing the body. A mismatch between header and body is rejected.
- Improved elicitation via multi-round-trip requests. A server that needs more input mid-call returns an input_required result; the client collects the answer and retries the original request with it attached. This replaces server-initiated requests that depended on holding an SSE connection open, which is what made elicitation incompatible with a stateless design.
The problems this resolves are operational. A stateless server runs behind a plain load balancer, on serverless platforms and at the edge, with no session affinity. Losing an instance no longer invalidates sessions pinned to it, so recovery becomes an ordinary request retry. And removing the handshake eliminates a full round trip before the first tool result. The speaker reported traffic reductions between 50% and 95% depending on the workload, which is consistent with removing the lifecycle exchange and caching tool catalogues, though the figure is workload-dependent.
Deployment guidance
- Migrate to the latest SDK. The wire contract changed; the SDKs absorb most of it.
- Enable list caching. This is where most of the traffic reduction comes from, particularly for an orchestrator spawning sub-agents: tools/list traffic drops from one fetch per sub-agent per server to one fetch per server.
- Ask whether you actually need discover or subscribe. server/discover exists for clients that must know a server's version and capabilities before their first real call, or that need a compatibility probe; a client that simply calls tools/list and tools/call and handles errors does not need it. Subscriptions push notifications when a resource changes and require a long-lived connection, which is exactly what a stateless deployment is trying to avoid. With cacheable lists and TTL hints, polling is usually cheaper, so subscriptions should be reserved for genuinely live resources.
The Hugging Face MCP server
The second half covered Hugging Face's own MCP server, a single endpoint in front of models, datasets, storage, papers, GPU-backed applications and compute. The design point is that it moves the computation to the Hugging Face side: rather than the agent pulling a dataset or a model down and processing it locally, it invokes a tool that runs the work where the data and the GPUs already are. The server is open source, and Hugging Face published a write-up of the transport trade-offs they faced building it, which is the clearest available explanation of the Streamable HTTP communication patterns.
Endpoint: https://huggingface.co/mcp
Build write-up: https://huggingface.co/blog/building-hf-mcp
Assessment
High relevance, moderate innovation. The specification change is not conceptually novel — it is the well-understood move from stateful sessions to self-contained requests that HTTP itself made decades ago — but it is the single most consequential change for anyone operating MCP infrastructure, and the keynote was the clearest explanation of it available. The deployment guidance was specific enough to act on, and the honest treatment of what stateless does not solve (application state still exists, it just moves behind explicit handles) kept it from being a marketing slot.
3. Three doors to one tool: MCP, WebMCP and CLI
The problem
The keynote started from discoverability: how does an agent find out which actions and capabilities are available to it, and how should a system owner expose them? The speaker framed this as urgent by noting that agent traffic has overtaken human traffic on the internet, so the question of how capability is published is no longer a niche integration concern.
The three doors
- MCP server. A server exposing tools, resources and prompts over the client-server protocol described in the previous keynote. Its advantage is that it carries the context of the system it fronts: the server sits next to the data and the business logic and can expose exactly the operations that make sense there. Its cost is that it is a separate component to build, deploy, authorize and operate.
- WebMCP[^1]. The page registers structured tools through a browser API (navigator.modelContext), and an agent discovers and calls them directly. The tools run client-side in the tab, with access to the DOM and to the user's already-authenticated session, so there is no separate server to deploy and no separate auth layer to build. Sensitive tools require human-in-the-loop confirmation; read-only tools can be marked to skip it. In practice this means a user can ask a model to perform an action on a site — adding someone to a group in an application, for example — and the model executes it in the context of the page, on the user's behalf.
- CLI. Invoking capability through the command line. Its advantage is that it leverages the terminal and the context of the working tree, which is already where a coding agent operates and already carries the developer's environment and credentials.
The open questions
The speaker closed on two considerations that apply across all three doors. Access control is the first: each door inherits a different security model — an MCP server needs its own authorization, WebMCP inherits the browser's origin model and the user's session, and a CLI inherits whatever the shell environment already holds. The second is composability between agents and the tools or APIs they call, which is the same policy problem raised in the day-one sessions and answered most concretely by Zalando.
Assessment
This was the most useful framing session of the conference, because it named a decision most teams are making implicitly. The three doors are not competitors so much as three different answers to "where does the context live" — in the backing system, in the page, or in the working tree — and that is the question to ask when choosing. WebMCP is the genuinely new element and is worth tracking: it removes both the deployment and the authentication burden for anything a user can already do in a web application, which is a large fraction of enterprise internal tooling.
Favourite Talks
Nine sessions stood out across the two days, selected on the strength of their original contribution rather than the profile of the speaker. They cover the software factory, delegated access, coordination observability, architecture choices, tracing, plan-first code review, regulated production deployments, and evaluation flywheels. The remaining sessions are summarised in the next section and were not materially weaker, only narrower.
1. State of the software factory (HumanLayer)
Why this talk
It asked the question the rest of the conference circled around: anyone can turn a ticket into a PR prompt in a coding agent, but how do we meaningfully collaborate on big things? And unlike most sessions on this topic, it brought evidence that the naive answer does not work.
The factory, before and after
The speaker first laid out the conventional software factory: vision, then a product manager, then engineering, then a backlog of work items, then someone builds the thing, a pull request, production, users, and finally complaints and feature requests flowing back to the product manager.
The agentic factory is the same loop with one box replaced. "Someone builds the thing" becomes orchestration, harness, sandbox and model. Everything else — the intake, the review, the deployment, the feedback path — is unchanged, and that is the point. Substituting the build step does not by itself make the loop faster, because the loop was never bottlenecked only there.
The evidence: agents degrade code quality over time
The speaker's central claim is that agents are bad at maintaining code quality over time. Pull request after pull request, quality and readability decline, and maintainability becomes the real constraint. SlopCodeBench was presented as the empirical support for this.
The benchmark contains 20 language-agnostic problems spanning 93 checkpoints. An agent implements the first specification from scratch, then repeatedly modifies and extends its own prior code as the specification evolves. Each checkpoint specifies only observable behaviour at a CLI or API boundary, so the architecture is left entirely to the agent — which is what distinguishes it from earlier iterative benchmarks that constrain design decisions too tightly to measure them. Two trajectory-level signals are tracked: verbosity, the fraction of redundant or duplicated code, and structural erosion, the share of complexity mass concentrated in high-complexity functions.
- No agent solved any problem end to end across 11 models. The highest checkpoint solve rate was 17.2%.
- Quality degrades steadily: structural erosion rises in 80% of trajectories and verbosity in 89.8%.
- Measured against 48 open-source Python repositories, agent code is 2.2 times more verbose and markedly more eroded.
- Tracking 20 of those repositories through their git history shows human code staying flat while agent code deteriorates with each iteration. This is the finding that matters: degradation is not a property of software evolution in general, it is specific to the agent.
- A prompt-intervention study shows that quality-aware prompting improves initial verbosity and erosion but does not slow the rate of degradation, improve pass rates, or reduce cost. You cannot prompt your way out of it.
The conclusion the authors draw, and the one the speaker used: pass-rate benchmarks systematically undermeasure extension robustness. An agent can pass every checkpoint while producing code that is progressively harder to extend, which means the standard evaluation actively conceals the failure.
Reference: Orlanski, G., Roy, D., Yun, A., Shin, C., Gu, A., Ge, A., Adila, D., Sala, F., and Albarghouthi, A. "SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks", arXiv:2603.24755, March 2026. Project page: scbench.ai. Code: github.com/SprocketLab/slop-code-bench. Note that the work is from the University of Wisconsin-Madison, not the University of Virginia as recorded in the session notes, and that a later revision reports somewhat different figures (15 agents, best checkpoint rate 14.8%, erosion in 77% and verbosity in 75.5% of trajectories, against 473 repositories). Cite the version you rely on.
For balance, the paper has drawn substantive criticism worth noting if the report is circulated: problems that frontier agents could solve in a single shot were removed from the pool, agents were given no handoff document or memory between sessions, and the prompt pushed only for correctness with no quality signal. Critics argue this measures how the tools were used rather than how good they are. The counter-argument is that the human baseline was measured on the same metrics and did not degrade.
Roughly a dozen practices
The speaker closed with a list of practices for running the agentic factory. The ones captured:
- Plan before you build, using the software factory pattern: build in hours, then validate for days with your users.
- Do not overplan. The counterweight to the previous point.
- Eliminate expected pain with agents — automate the friction you already know is coming.
- Find a way for the model to report what it actually did, rather than what it intended to do.
- Help agents test their own work.
- Fight slop with deterministic code. The same principle as hooks and callbacks elsewhere in the conference, applied to code quality.
- Feed user feedback into the factory.
- Feed incidents into the factory: wire in monitoring and the actions to take on an alert.
- Let models experiment.
- Loops need forward pressure and back pressure. Forward pressure produces changes; back pressure from tests, review and incidents is what stops quality degrading.
- Compound over time — learn from past sessions rather than starting cold each time.
Assessment
One of the most important session of the conference, because it is the only one that argued against the prevailing optimism with data rather than anecdote. The practical consequence is that a coding-agent strategy needs an explicit answer to code-quality degradation — deterministic gates, refactoring cycles, or architectural constraints imposed outside the agent — and that answer cannot be a better prompt.
2. Economies of scale for MCP and agents: why you need an identity broker (Zalando)
Why this talk
It is the production answer to the problem day one raised repeatedly, delivered by a team running it internally, with the code open sourced. It also reframes delegated access as a platform-economics problem rather than a security problem, which is a more useful framing for getting it funded.
The pain it solves
The argument starts from repetition rather than from security. Connecting a tool is only the start: every MCP integration has to discover authorization, validate token audiences, and handle scopes and errors. Frameworks such as FastMCP can implement the full OAuth 2.1 ceremony the MCP specification mandates, but that means configuring and maintaining it for every server you deploy, plus secure token storage — in effect running a fleet of OAuth2 authorization servers. The problem is the repeated implementation, not OAuth itself.
The speakers drew the explicit parallel with microservices. Controls such as authentication, transport security, retries and rate limiting used to live in each service; product teams repeated the work and the implementations drifted. Moving those controls into gateways and identity services let product teams focus on domain logic. The same shift is now due for agents: runtime, routing, delegation, registration and observability are shared concerns being rebuilt by every team. Moving them into the platform means a fixed initial investment plus maintenance, after which each new agent reuses the security properties and the infrastructure.
Why the obvious shortcuts fail
Attribution. An agent is not a person and not a fixed script. A human delegates intent, the agent has its own identity, and the tool performs the action; memory and model inform decisions but confer no authority. The action has to be attributed to both the user and the agent.
The robot account. A dedicated account holding persistent credentials does not solve delegated access; it just moves the credentials. Worse, a shared agent account pools the permissions of every user behind it, so the current user effectively gains the combined permissions of everyone else. This was named explicitly as an anti-pattern.
The model proposed instead: identify the agent and delegate the user's access, with effective access being the intersection of what the user may do and what the agent is allowed to do, plus a live delegation grant. Delegation is a graph rather than a single credential — user, agent, target and scope — and every edge stays attributable.
How it works
The broker keeps the existing identity provider. A trusted proxy authenticates the user and forwards that identity; the broker answers the separate question of what has been delegated, resolving the user, agent, service and live grant, and managing provider sessions, consent and scoped access centrally. It sits in the call path, driven by an infrastructure gateway such as agentgateway, between an agent and an MCP server or between agents.
The token exchange for a single tool call runs as follows. The agent sends a tool request carrying a user-agent token. The gateway performs an RFC 8693 token exchange against the broker, presenting the subject token, the resource and a signed client assertion. The broker verifies the gateway and checks for a live grant, denying the exchange if none exists. It then retrieves the matching provider token, decrypting it from the token vault, and returns it to the gateway only. The gateway replaces the Authorization header and forwards the request to the target. The agent never receives a provider token at any point.
Consent is a user-facing surface, not a configuration file. The screen shows validated agent metadata so the user knows which product and operator is asking, describes the requested access as a permission set in business language rather than raw provider scope strings, and records the user, agent and permissions with an expiry where appropriate. Users can review and revoke grants afterwards rather than only at first login.
Two other properties are worth noting. Registration is declarative: a Kubernetes custom resource describes an agent or MCP integration and an operator maps its lifecycle to broker API calls, so product teams never maintain separate registration scripts. And revocation is immediate — withdrawing the user-agent grant means the next token exchange is denied, with no deployment and no waiting for a token to expire.
Policy: two gates per call
This is where the session connects to the wider authorization discussion. Two independent checks must both pass for a call to proceed: the broker authorizes the token exchange against the grant and its policy, covering user, agent and target; and the gateway evaluates an Open Policy Agent policy against the specific tool and its arguments. A denial at either gate stops the call.
OPA (Open Policy Agent): a CNCF general-purpose policy engine. Policies are written in the Rego language and evaluated at a decision point outside application code, which is what allows the tool policy to be changed without redeploying the agent or the MCP server.
The speakers were candid that a policy gate is not yet a human approval flow. Today the system offers consent, revocation, token exchange and allow-or-deny tool policy, and a missing delegation can trigger an MCP elicitation. Holding a risky call, sending an approval URL through elicitation and resuming after approval is in progress; binding an authenticator approval to a specific high-risk transaction via CIBA backchannel authentication, and carrying the user's intended task as policy input so that calls outside that intent can be flagged, are on the roadmap.
Code: https://github.com/zalando-incubator/agentic-identity-broker
Assessment
High innovation in assembly rather than in primitives — RFC 8693 token exchange, OAuth2 consent and OPA are all established — and the strongest engineering session of the conference. Two ideas generalise regardless of whether you adopt the broker. First, effective access as the intersection of user permission and agent permission, with a live grant required, is the correct mental model and is expressible in most existing systems. Second, keeping provider tokens out of the agent entirely, with substitution happening at a trusted gateway, is the same pattern Discobox applies to sandboxed coding agents, arrived at independently. When two unrelated teams converge on "the agent never holds the credential", that is a signal.
3. Autonomous vulnerability remediation in production (Thomson Reuters)
Why this talk
The only session presenting autonomous changes merged into production at scale rather than proposed, and the practical counterpart to the software-factory session: it shows what the surrounding engineering has to look like before an agent-authored change can be trusted.
Contribution
AI is finding security issues faster than any team can fix them by hand. Letting AI fix them is the obvious next step, but pushing automated changes into production carries a very high trust bar. Over the past year the team built and shipped two autonomous remediation agents raising pull requests against hundreds of production repositories. Both are being open sourced.
Once a vulnerability is reported, time-to-resolve is the metric that matters. A general-purpose AI assistant already helps when the fix is cheap and local. The loop degrades as the fix spreads beyond a single location — to application level, across modules, or into higher complexity — because each step outward requires substantially more context about the codebase.
The framing sentence from the session: a valid fix does not scale by itself. Producing a technically correct patch is necessary but not sufficient. The patch must also build, pass tests, not break callers, and be reviewable enough that a human merges it.
SAST agent (static application security testing): handles findings in first-party code, where the vulnerability is in code the organisation owns and can edit directly.
SCA agent (software composition analysis): handles findings in third-party libraries, where the fix is usually a version bump and the difficulty is the breaking changes that come with it, especially in legacy code.
Both follow the same two-phase loop: first plan, by understanding the context of the repository; then generate, where the model writes the fix and a set of checks judges it. The judging step is what separates this from an assistant — the agent evaluates its own output against build and test evidence before a human sees it.
What produced the trust
The speaker was explicit that trust came from unglamorous engineering rather than better prompting: teaching the agents how to build and test an application, tracing data flows so that risky fixes are rejected rather than attempted, and resolving breaking changes when modernising legacy code. The team also analysed patterns across rejected pull requests, which is the feedback loop that moved hundreds of PRs to merged.
CVE (Common Vulnerabilities and Exposures): the public catalogue of known security vulnerabilities. Growing CVE volume is the structural argument for automating remediation.
Assessment
High innovation, and the most credible answer to the coding-agent scaling problem. Read alongside the software-factory session, it is an existence proof that the degradation problem is tractable: the difference is that every change here passes through deterministic build, test and data-flow gates before a human ever sees it. The generalisable lesson is that the hard part of automated remediation is not generating the patch but the surrounding evidence that lets a reviewer trust it. That is a software-factory investment, not a model investment.
4. Debugging the distributed mess: observability for multi-agent systems (Databricks)
Why this talk
The only session offering concrete, computable metrics for coordination quality rather than a general argument that observability matters.
Contribution
Most agent failures do not happen in the model; they happen in the handoff. When a supervisor agent passes flawed context downstream — a wrong tool output, misrouted state, or a plausible hallucination — the receiving agent has no way to detect it and continues confidently on a corrupted foundation.
In MCP-based systems this is structural rather than accidental: tool call responses become shared context across agents that never communicate directly, so one bad result upstream poisons every agent that touches it downstream. Traditional end-to-end testing misses this entirely, because the final output can still look reasonable.
Failures can originate in four places: the orchestration layer, the agent prompts, agent execution, and contention between agents (semaphores, deadlocks, waiting on a peer that never responds). Output monitoring is therefore insufficient. In a multi-agent system you have to monitor routing (which agent was selected, and was that right), handoff (what context actually crossed the boundary) and latency at each step rather than end to end. The triage distinction that matters: a model error, a routing error and context corruption are three different faults with three different fixes, and only per-boundary tracing tells them apart.
Four dimensions of governance
- Operational health.
- Cost management.
- Compliance and audit.
- Quality and value.
Evaluating the handoff
Two metrics were proposed for the handoff itself, which is usually not measured at all.
- Utilisation: how much of what was passed downstream was actually used. The example given was 3 facts used out of 40 passed, indicating the upstream agent is over-producing context.
- Coverage: needed and present, divided by needed. The example was a downstream agent missing 2023 revenue that it required. Low coverage indicates the upstream agent is under-producing.
The two are independent. A handoff can be simultaneously bloated and incomplete, and only measuring both reveals it.
Evaluating coordination: the claim ledger
Agents emit structured claims consisting of entity, value, confidence and source. You then measure whether conflicts between claims are detected and resolved correctly. This surfaces three failure modes that normally go unmonitored:
- Silent contradiction: two agents assert incompatible values and the conflict is never detected.
- Wrong resolution: the conflict is detected, but the losing claim was the correct one.
- Fabricated merge: the resolved value was never reported by any agent, having been synthesised during the merge.
Tooling shown: MLflow
The demonstration used MLflow for agent monitoring, in the CrewAI flavour. What MLflow provides for this use case:
- Tracing. end-to-end observability for agent systems, recording inputs, outputs, intermediate steps, metadata, latency and token usage at each step. It is OpenTelemetry-compatible and supports the GenAI semantic conventions, so traces export into an existing observability stack rather than locking into one vendor.
- Broad framework coverage. automatic instrumentation for the common frameworks (CrewAI, LangGraph, the OpenAI Agents SDK, DSPy, Pydantic AI and others) alongside manual instrumentation, which matters given how often teams switch frameworks.
- Queryable structured attributes. business metadata such as user_id and session_id on each span turns single-trace inspection into population-level queries: all traces where a given tool failed for a given user segment over a time window. This is what makes the utilisation and coverage metrics computable rather than anecdotal.
- Evaluation and monitoring with reusable judges. built-in and custom LLM judges and scorers run against traces during development, and the same judges and scorers are reused for production monitoring, so the quality metric stays consistent across the lifecycle. Custom scorers are where the handoff and claim-ledger metrics from this talk would be implemented.
- Human feedback loop. a Review App collects structured feedback from domain experts, turning production traces into evaluation datasets.
- Prompt versioning and trace replay. a failing production trace can be replayed against a changed prompt or agent version.
On Databricks specifically, traces can be stored in Unity Catalog as OpenTelemetry Delta tables, making them queryable in SQL and governed by the same access controls as the rest of the data estate.
Related: Databricks Unity Gateway
Unity Gateway is Databricks' control plane for enterprise AI, announced at Data + AI Summit in June 2026, when the earlier AI Gateway was folded into Unity Catalog. It extends Unity Catalog's governance model beyond data and AI assets to the runtime interactions between models, agents, MCP services, skills and tools, which makes it a commercial instance of the control-plane pattern from the day-one keynote — and the vendor counterpart to what Zalando built in-house.
- Registration and governance in one place. internal models, external model providers, MCP services and reusable skills are registered in Unity Catalog and governed with the same permissions and audit model used for data.
- MCP governance with on-behalf-of execution. fine-grained control over which agents reach which external systems, with the MCP call executing under the requesting user's exact permissions rather than a shared service identity — the same on-behalf-of principle Zalando implements through token exchange.
- Traffic management. a unified API across providers with fallbacks, rate limits and load balancing, plus automatic selection of the model and harness best suited to a task.
- Runtime guardrails. PII masking and prompt-injection detection applied at the interaction, plus identity-aware controls on trace data.
- Cost controls. hard spend caps, budgets, rate limits and alerts, with spend attributed across users, teams, applications, agents and models. This maps onto the financial-control dimension of the control plane.
- Observability. end-to-end traces across LLM and MCP calls with a full audit trail. Some controls were still in beta at the time of the announcement.
Product page: https://www.databricks.com/product/artificial-intelligence/unity-gateway
Assessment
Genuinely innovative on the measurement side. Utilisation, coverage and the claim ledger are the only concrete proposals of the conference for quantifying coordination quality, and they are framework-agnostic — implementable as custom scorers on any tracing stack. The Databricks product content was present but contained, and the core argument stands without it: evaluating final outputs is not evaluating a multi-agent system.
5. When NOT To Use an Agent: Choosing Between Workflows, Services, and Agent Systems (Jigyasa Grover, Uber & Rishabh Misra, Atlassian)
Why this talk
The core message is that agents are not the default architectural panacea and should be avoided in favor of deterministic workflows unless full agentic autonomy is needed in the case of high-complexity, high-error-tolerance domains (e.g., open-ended research, writing).
Contribution
Building agentic systems often stems from inflated expectations and pressure to adopt autonomous patterns, yet deployed agents frequently exhibit severe failure modes in production. In real-world incidents, autonomous agents have caused cascading operational failures—such as recursive loops, severe latency drift, boundary violations, and unconstrained action execution that bypassed safeguards—because their primary objective is unconstrained task completion. Over-relying on autonomous prompt decision-making creates an "evaluation abyss" where models appear successful in isolated tests but degrade rapidly in production environments.
Addressing these architecture flaws requires moving away from pure autonomy toward structured, deterministic composition where control flow remains explicitly managed in code. By treating LLMs as constrained components rather than monolithic decision-makers, teams can dramatically reduce error rates and improve system predictability.
Architectural Solutions & The Router Pattern
To mitigate agent unreliability, systems should implement a Router Pattern that replaces monolithic decision-making agents with lightweight, confidence-based classifiers. In this architecture, an initial routing layer analyzes user queries using rules, embeddings, and fast models to direct requests into specialized, deterministic pipelines:
- Grounded Retrieval Search: For information extraction and context-backed answering.
- Reasoning LLMs: For complex, multi-step logical synthesis when required.
- Action Tools / APIs: For deterministic, scoped execution of programmatic commands.
Additionally, employing cost-tier routing ensures simple intents use fast, inexpensive models, while complex tasks leverage deeper reasoning models. Strict schema enforcement at system boundaries ensures that invalid decisions fail immediately rather than propagating through execution loops.
Observability & Telemetry Framework
Comprehensive observability and granular tracking are essential for running composable agentic workflows reliably. Every execution step must generate structured telemetry, including:
- Exact prompt/response pairs and token counts per step (crucial for loop detection).
- Nested execution trees with per-step timing to pinpoint exact performance bottlenecks rather than relying solely on end-to-end latency.
- Tool call parameters and schema validation outcomes.
- Decision confidence scores and explicit flags for human-in-the-loop review gates.
Evaluation Abyss: A phenomenon where an LLM-powered system achieves high passing scores on static test suites and benchmarks, but fails unpredictably or degrades rapidly when exposed to the broad distribution of real-world production inputs.
Assessment
This session provides a much-needed, highly pragmatic counterweight to industry hype, delivering actionable engineering guardrails for production AI systems. Its generalizable lesson is that software engineering fundamentals—deterministic control flow, strict schema validation, and granular observability—must take precedence over probabilistic autonomy. For our engineering team, adopting a router-and-pipeline architecture offers a reliable blueprint for building predictable, cost-bounded, and easily debuggable AI-augmented applications.
6. From Opaque To Observable: Tracing Multi-Agent OpenClaw Workflows With OpenTelemetry (Jordan Augé, Cisco Systems)
Why this talk
It provides a concrete, production-ready solution for achieving fine-grained observability and tracing across complex multi-agent workflows.
Contribution
As multi-agent systems fan out across complex workflows involving parallel branches, subagent handoffs, external tools, and model queues, traditional monitoring falls short, leaving engineering teams blind to execution paths. When a failure occurs, standard logs make it nearly impossible to differentiate whether the root cause stems from agent reasoning errors, infinite tool loops, or external infrastructure bottlenecks like webhook timeouts. Without native end-to-end tracing, operating autonomous architectures at scale becomes an intractable black box.
To solve this opacity, Jordan Augé discussed the development of an open observability plugin for OpenClaw—often referred to as InsightClaw—which stitches fragmented execution paths into a unified OpenTelemetry trace model. By separating semantic reasoning failures from operational infrastructure issues, this telemetry framework gives engineering teams the precise visibility needed to debug complex multi-agent interactions.
OpenTelemetry Instrumentation & Signal Architecture
The telemetry architecture combines three distinct signal paths to capture the complete lifecycle of multi-agent execution:
- Typed Lifecycle Hooks: Captures explicit state transitions across request, agent, tool, and response flows.
- Diagnostics Events: Monitors operational telemetry including model usage, cost, queue metrics, and stuck-session signals.
- Provider SDK Auto-Instrumentation: Automatically instruments underlying GenAI model provider calls for comprehensive token and latency tracking.
Session Semantics & Overhead Mitigation
To handle long-lasting agent sessions that break standard telemetry conventions, the system introduces robust session semantics and performance safeguards:
- Runtime Correlation Keys (
openclaw.session.key): Maintains execution lineage across multi-turn handoffs and parallel agent forks. - Heuristic Session Boundaries: Governs session lifecycles using idle timeouts or explicit resets rather than rigid time windows.
- In-Plugin Caching & Aggregation: Prevents high-volume telemetry payloads from crashing the runtime by aggregating data locally before export.
OpenTelemetry: An open-source observability framework providing standardized APIs and SDKs to collect and export telemetry data such as traces, metrics, and logs.
Assessment
This session delivers practical value by bridging the gap between abstract agent execution and infrastructure health. Its generalizable lesson is that native end-to-end tracing using open standards like OpenTelemetry is non-negotiable for operating multi-agent systems reliably in production, enabling teams to isolate semantic reasoning loops from infrastructure failures with precision.
7. Pull Requests Are Dead, Long Live Peer Review (Dylan Ratcliffe, Overmind)
Why this talk
It is brutally honest about the evolution of the job of software engineering in the age of AI and calls for optimism.
Contribution
When AI code generation scales to produce the vast majority of a team's codebase, traditional pull request workflows break down entirely. Human reviewers are forced to audit massive diffs generated by a machine that simply agrees with any prompt, turning code review into an exhausting audit of automated extrapolation rather than a collaborative discussion between engineering peers. This shifts the engineering bottleneck away from writing code and onto reviewing opaque machine outputs, causing widespread reviewer fatigue and degrading overall software quality.
To restore effective review, teams must shift human oversight left from the code diff to the implementation plan. Instead of treating the AI as a peer programmer, engineers collaborate to define a clear, rigorous intent and architectural plan before any code is generated. An agentic coding system then executes against that plan until the output matches the specified design. Reviewers evaluate the human-written recipe rather than the final meal, treating code as a disposable commodity that can be easily regenerated if it deviates from the agreed-upon specification.
Adopting this plan-first workflow transforms team productivity by establishing strict guardrails around agentic code generation and preserving human attention for meaningful design decisions.
Preserving Human Attention
Protecting human attention is the single most critical requirement for scaling software engineering in the age of AI. When engineers are flooded with thousands of lines of machine-generated boilerplate diffs, cognitive overload sets in quickly, leading to rubber-stamp approvals, missed edge cases, and severe burnout. By shifting review effort upstream to implementation plans and automating execution via Until loops, organizations protect their engineers from tedious syntax auditing. This preserves precious human attention for high-value architectural design, creative problem-solving, and strategic decision-making.
Shift Left: Plan-First Code Generation Architecture
Moving engineering oversight upstream fundamentally alters how systems are built and validated:
- Intent Specification: Humans define strict architectural recipes, requirements, and boundaries before execution begins.
- Automated Execution Loops: Agentic coding systems act as disciplined executors rather than autonomous decision-makers, following the plan iteratively.
- Disposable Code Lifecycle: Because code can be regenerated instantly from a sound plan, teams discard broken or non-compliant implementations without hesitation.
The "Until" Loop Architecture
To replace open-ended, unpredictable "do-while" agent loops, the speaker introduced the "Until" loop pattern as a rigorous execution control mechanism. Unlike autonomous loops that run indefinitely toward a vague objective, an Until loop operates against a deterministic, machine-readable completion condition defined in the upstream plan. It uses automated test suites, static analysis, and judge agents to continuously evaluate the generated code's adherence to specs. The loop iterates strictly until all assertions pass or a hard iteration threshold is hit, ensuring the system fails fast and predictably rather than drifting into infinite loops or generating invalid code.
Agentic Code Generation: The automated production of software source code by AI systems operating against predefined plans or natural-language prompts.
Assessment
This presentation provides a visionary yet pragmatic roadmap for engineering teams navigating the rise of AI-generated code.Human expertise must shift from authoring and auditing syntax to designing robust intent and architecture. Embracing this shift protects engineering well-being, eliminates review fatigue, and positions teams to harness AI as a powerful lever for software delivery.
8. When Agents Run Healthcare: Building Reliable Agentic Systems in Highly Regulated Environments (Janosch Woschitz, BARMER)
Why this talk
It provides a realistic, production-tested blueprint for embedding autonomous agents into high-stakes, heavily regulated enterprise environments while successfully applying everything the entire conference highlighted.
Contribution
The public statutory health insurance system in Germany faces a dual challenge: demographic change is creating a shortage of skilled professionals while operational pressure continues to rise due to an aging population. At BARMER—one of Germany's major public statutory health insurance providers—healthcare organizations must automate high-volume processes without compromising reliability, governance, or trust. Drawing on real-world production insights from Janosch Woschitz, Senior AI Architect at BARMER KI, deploying agentic systems in such heavily regulated environments requires moving beyond isolated chatbots or copilots to embed agents directly into core enterprise workflows while maintaining strict compliance, auditability, and resilience.
To meet enterprise standards, architectures must separate agent runtimes from deterministic services, utilizing robust workflow orchestration and strict schema boundaries to prevent invalid operations. Because models inherently exhibit non-deterministic behavior, production systems cannot rely solely on prompt instructions to enforce compliance. Instead, engineering teams must implement programmatic guardrails, including cryptographically verifiable audit logs, fine-grained role-based access control, and seamless human-in-the-loop escalation paths for high-stakes decisions. These patterns ensure that every agent action is traceable, reproducible, and bounded by hard code constraints rather than probabilistic reasoning.
Real-World Healthcare Use Case & Core Architecture at BARMER
To address high-volume administrative burdens and service demands, BARMER developed production-grade agentic workflows to automate complex insurance processing tasks, such as handling member inquiries, claims verification, and medical service approvals:
- Core Use Case: Automating multi-step case management workflows that require synthesizing member history, insurance policy rules, and medical guidelines while ensuring zero compliance drift.
- Modular Architecture: The system separates the agentic reasoning layer from core insurance mainframes, utilizing a structured orchestration engine to coordinate task execution.
- Execution Flow: Incoming requests are ingested and parsed through strict schema validators, routed to specialized agent workers equipped with scoped API tools, and evaluated against programmatic business rules before any state change occurs.
Deterministic Architecture & Compliance Guardrails
Isolating non-deterministic reasoning from core business logic ensures that critical state transitions and regulatory mandates are enforced strictly in code:
- Deterministic Separation: Decouples probabilistic model execution from deterministic application backends, preventing unstructured agent outputs from directly modifying core databases.
- Programmatic Guardrails: Enforces hard boundaries via schema validation, role-based access control (RBAC), and automated compliance checks rather than relying on prompt-level instructions.
- Failure Boundary Design: Anticipates agent failure modes—such as infinite tool loops or boundary violations—by enforcing strict input/output schemas and immediate failure propagation.
Operational Traceability & Human-in-the-Loop Integration
Maintaining production stability in regulated domains requires continuous visibility and structured intervention points:
- Comprehensive Auditability: Maintains immutable, structured logs of every agent turn, tool invocation, and decision path to satisfy stringent regulatory reporting and compliance requirements.
- Human-in-the-Loop Escalation: Implements clear, programmatic review gates that intercept sensitive or ambiguous actions, routing them to human operators before execution.
Deterministic Separation: The architectural practice of isolating non-deterministic model reasoning from deterministic application logic to ensure critical business rules are enforced strictly in code.
Assessment
This session delivers exceptional practical value for engineering teams building AI systems in strict regulatory spaces. It highlights that probabilistic models must be securely wrapped in deterministic guardrails and rigorous audit frameworks to achieve production viability. For our engineering team, adopting strict separation of concerns and programmatic escalation gates provides a reliable pattern for deploying compliant, high-assurance agentic workflows.
9. Evaluating Agents at Scale: From 50 Examples to a Production Flywheel (Bauke Brenninkmeijer, Orq.ai)
Why this talk
It answers the remaining critical question at the end of the conference: who's judging the judge.
Contribution
Evaluating autonomous AI agents at scale breaks traditional software testing paradigms because non-deterministic outputs and complex reasoning paths make static assertions ineffective. When systems rely on open-ended multi-step execution, simple unit tests cannot capture semantic correctness or contextual nuance. Organizations are left struggling to scale quality assurance beyond a handful of manual test cases, creating a dangerous blind spot as agents move from controlled pilot environments into unpredictable production workflows.
To solve this scaling bottleneck, engineering teams must implement an evaluation flywheel powered by LLM-as-a-judge patterns and active learning loops. As demonstrated in practical data-analysis workflows answering dozens of complex business cases, the architecture begins with a small, curated set of evaluation examples. An automated jury of LLM judges evaluates the agent's outputs, flagging discrepancies where judges disagree on whether a run passes or fails. Rather than discarding these edge cases, the system isolates these "flipped" data points—instances of high ambiguity—and routes them directly to human domain experts for review.
The human resolution of these boundary questions is then codified back into the evaluation logic, systematically training and refining the LLM judges over time. This creates a self-sustaining flywheel where human judgment is captured, scaled, and transformed into an automated pass/fail evaluator that can be executed deterministically on every code or prompt change. By systematically closing the loop between production failures, judge disagreements, and human-in-the-loop validation, teams can continuously improve evaluation alignment without bottlenecking engineering velocity.
The Evaluation Refinement Flywheel (Step-by-Step)
- Bootstrap with Small Curated Datasets: The process starts with a small, highly curated set of about 50 evaluation examples, such as business-case queries answered by data-analysis agents, establishing an initial baseline for system behavior.
- Automated Jury Evaluation: An automated panel or jury of multiple LLM judges evaluates the agent's outputs and execution traces, assessing semantic correctness and functional constraints across predefined evaluation criteria while measuring consistency across multiple scoring runs to detect variance.
- Isolating Judge Disagreement ("Flipped" Cases): The system continuously monitors judge scoring to isolate the grey zone—high-ambiguity boundary cases where individual judges disagree, scores cluster near uncertainty thresholds, or verdicts flip unexpectedly.
- Human-in-the-Loop Consensus Routing: Each specific boundary question arising from the grey zone is systematically routed to human domain experts who inspect the execution context and provide authoritative ground-truth decisions.
- Judge Codification & Alignment: Human expert resolutions are codified back into the evaluation rubrics to achieve strict evaluation alignment, systematically training and refining the LLM judges so their automated scoring mirrors human intent.
- Continuous Automated Regression Suite: Human judgment is thus captured, scaled, and transformed into a self-sustaining test suite that runs deterministically as an automated regression gate on every code or prompt commit.
Strategic Data Sampling Strategies
The evaluation pipeline relies on three distinct sampling mechanisms to ensure robust coverage across production and edge-case scenarios:
- Random Sampling of General Traffic: Captures a representative cross-section of everyday user queries to evaluate how the agent performs on standard, unconstrained inputs.
- Historical Production Failure Traces: Harvests past system errors, edge cases, and failed execution runs to ensure that known failure modes are continuously tested against regression.
- Domain-Expert Curated Edge Cases: Targets highly specialized or high-risk scenarios manually crafted by domain experts to stress-test system boundaries and safety limits where automated traffic might fall short.
Integrating the Flywheel into Production & Regression Testing
To operationalize the evaluation flywheel in a production-grade system, the pipeline must be embedded directly into the CI/CD deployment lifecycle. Whenever a prompt, model version, or underlying tool definition is modified, the system automatically triggers the comprehensive test suite derived from accumulated human decisions and production traces. Rather than relying on slow manual reviews, these automated regression tests evaluate the agent's performance against historical benchmarks and newly discovered grey-zone cases. If consistency scores drop or alignment metrics fall below established thresholds, the deployment pipeline halts, preventing untested reasoning shifts or regressions from reaching production environments.
Agent-as-a-Judge Alignment: The continuous process of refining automated LLM evaluation rubrics using human-in-the-loop feedback to resolve judge disagreements and ensure reliable scoring.
Assessment
This session delivers practical value by solving the most persistent bottleneck in LLM engineering: reliable, scalable evaluation. Quality assurance cannot rely on static test suites; instead, teams must build self-improving evaluation flywheels that capture human expertise and encode it into automated regression gates.
Other Talks Attended
Five further sessions, summarised for completeness. All were strong; they are placed here rather than above because their contribution was narrower in scope, not weaker.
10. Delegated authorization at scale: ReBAC, Zanzibar and SpiceDB
Platform and DevOps teams now run hundreds of agents holding credentials that can touch everything, production included. Without deterministic permission checking that is a production incident waiting to happen. The talk argued that relationship-based access control is the right model and demonstrated it with a DevOps agent. Read alongside the Zalando session, it is the other half of the same answer: Zalando solves how a grant is obtained, brokered and revoked, while this session solves how the permission itself is expressed.
- ACL (access control list): an explicit list of who may do what on each resource. Precise, but it does not scale and encodes no structure.
- RBAC (role-based access control): permissions attached to roles and users receive roles. Scales better, but cannot express relationships between resources, which is precisely where agents need scoping.
- ReBAC (relationship-based access control): permission is derived from relationships between entities. A person is viewer of a folder, the folder contains a document, so the person can view the document. This is a graph over users, folders and documents, and an access check becomes a path-finding problem in that graph.
Relations are ordinary domain facts — member of a group, editor of a document, uploader of a video, owner of a record — written as tuples. For example, document:123#owner@user:3 states that user 3 is the owner of document 123.
Zanzibar is Google's global authorization system, described in a 2019 paper, with open-source implementations following, of which SpiceDB is the most prominent. It fits where you need low-latency, high-throughput authorization checks, hierarchical models, or ambient context in the decision. The mechanism has two halves: a schema declaring object types and how they relate, and the relationship tuples themselves. Schema plus relationships produce a working permission check, and permissions are computed rather than enumerated. A time window can be attached to a relationship, giving time-bound grants that expire without a cleanup job, and operators can be applied to relations so permissions compose.
The demonstration used a DevOps agent with access to staging and production, built on goose and SpiceDB. The agent goose_alice acts on behalf of user:alice, and permission is checked for every action rather than once at session start. The punchline: revoking the agent's staging access automatically suspended its production access, because production is defined in terms of staging in the schema. A role-based system cannot express that cleanly. The four agent-permission patterns the model handles well are scoped delegation, expiring grants, instant revocation and hierarchical permissions.
goose: an open-source AI agent framework, used here as the acting agent. SpiceDB: an open-source Zanzibar implementation, used as the authorization store.
11. Orchestrating a production multi-agent system on Google ADK
Most things called AI agents are a single LLM wrapped in a system prompt; they demonstrate well and collapse when real users arrive. The talk covered a tree of 12 specialist agents running in production for a developer community of more than 5,000 members, built on Google's Agent Development Kit and coordinating through sequential pipelines, parallel fan-out and LLM-driven dynamic routing. Onboarding chains three agents in sequence, external knowledge retrieval fans out across GitHub, Dev.to and StackOverflow in parallel, and the root agent decides delegation per message.
- Rule 1 — the orchestrator never owns tools. Tools belong to the agents being orchestrated. The orchestrator decides who acts; it does not act itself. This keeps routing decisions separable from execution, which matters when diagnosing whether a failure was a routing error or an execution error.
- Rule 2 — prefer plain tools over MCP, with three exceptions. Use MCP when the tool crosses an organisational boundary, when it lives in someone else's system, or when it will be consumed by someone else. Inside your own system, a direct tool call avoids a protocol hop you do not need. This is the most concrete guidance anyone offered on the integration question, and it anticipates the day-two keynote on the same topic.
- Rule 3 — callbacks are the immune system. They run deterministically at defined points in the flow: clearing the cache, checking for PII, or anything important enough that it must not depend on an instruction surviving in the context window. The speaker had to build this layer himself; none of it appears in the framework tutorials.
Fan-out means executing agent tasks in parallel. Each agent writes its output under its own state key and can expose that key for other agents to depend on. If an agent crashes or stops it must not take down the rest of the workflow. The speaker walked through the resulting synchronisation problems in detail, and the architecture patterns that emerge are the familiar ones from distributed systems. The failure stories reinforced this: a drain-loop cancelled a ParallelAgent mid-flight, and recency drift surfaced 2020 articles in 2026 results.
ADK (Agent Development Kit): Google's open-source framework for multi-agent systems, providing sequential, parallel and LLM-routed composition primitives rather than a single monolithic agent loop.
12. What is an agent's identity?
Enterprises understand human identity reasonably well and are passable at service accounts. An agent is neither. Agents are driven by intent, discover and invoke tools at runtime, make decisions, and interact with APIs, databases and other agents. The four questions an enterprise will ask are: who is this agent, what is it allowed to do, what has it done, and can we revoke its authority? The motivating example: an agent asked to investigate a problem may take a creative and unexpected path — instead of inspecting logs or reading deployment status, it could alter a database schema or restart the database.
Option 1 — use the user's identity. Sometimes this is enough: user to agent to Kubernetes, where Kubernetes sees the user. It breaks when several agents work on different parts of a task and you want to delegate part of your authority rather than all of it. It also maps poorly in practice: is the user the one currently in context, or the organisation?
Option 2 — use the workload identity. One service account per workload, for instance the one that deploys to production. This fails as soon as multiple agents share the same workload, because a Kubernetes service account cannot distinguish between them. Workload identity as commonly implemented is too coarse.
The proposed direction: an agent registers with an authority under a defined policy stating what it is supposed to do, and that policy authorises it at runtime. Identity is the anchor that policy binds to, and identity presentation should be cryptographic rather than a shared secret. The speaker summarised the flow as present, discover, authorize, and advocated AAuth as a portable identity usable at first contact. Note that Zalando reached a compatible conclusion by a different route: keep user identity and agent identity both in the authorization context, and require a live grant binding them.
AAuth in more detail
AAuth (pronounced "AY-awth") is an authentication and authorization protocol for AI agents, published as an IETF Internet-Draft, draft-hardt-oauth-aauth-protocol, by Dick Hardt, author of OAuth 2.0 and co-author of OAuth 2.1. It is explicitly not a replacement for OAuth or OIDC; it targets the cases those protocols handle badly. Traditional software knows at build time which services it will call, so registration, key provisioning and scope configuration happen before the first request. Agents discover resources at runtime, run long tasks spanning several services across trust domains, and need authorization decisions mid-task.
- Cryptographic identity per agent. an agent identifier of the form aauth:local@domain is bound to a signing key published at a well-known URL and verifiable by any party, with no pre-registration and no shared secret. This is what makes it portable at first contact.
- Signed requests instead of bearer tokens. it builds on HTTP Message Signatures and the HTTP Signature Keys specification. Every request is signed, so an intercepted token is not sufficient to impersonate the agent.
- Explicit delegation. a token can represent the agent and the person simultaneously, so delegation chains are visible and enforceable rather than bolted on through token exchange.
- Four resource-access modes. identity-based, resource-managed (two-party), person-server-asserted (three-party), and federated (four-party), with agent governance as an orthogonal layer. The federated mode separates authentication, which stays with the person server, from authorization, which is federated to a centralised access server implementing organisational policy — structurally the same split Zalando implements with an identity provider plus a broker.
- Graduated requirement levels. a resource can require pseudonym, identity, auth-token, interaction, or approval. This is how human approval enters the protocol as a first-class step rather than an application-level workaround.
Status: still a draft, with revision 10 published in August 2026, alongside a TypeScript implementation, a protocol explorer and working demonstrations using agentgateway as the policy enforcement point. Too early to say whether it becomes the standard, but it is the most complete proposal in this space.
For comparison, SPIFFE and SPIRE issue cryptographic identities to workloads and are the mature answer to the workload-identity question. The argument is that agents need an identity layer above the workload, because several agents can share one workload.
13. From a personal agent to an org chart: an AI chief of staff and a shared agent catalog
The speaker watched five teams present their agent setups at a meetup and saw the same patterns reinvented five times: duplicated prompts, duplicated tool configuration, no shared memory — the same observation Zalando made about repeated identity work, applied to agent definitions. The response was to build a personal system first, then extract the reusable parts into an organisational asset. The system, "Paul", is an AI chief of staff composed of 13 specialist agents covering job families such as analyst, writer, broadcaster, operator, architect and reviewer. One orchestrator delegates; a reviewer gates quality.
The most transferable idea was procedural. Before writing any prompt, the speaker draws an organisation chart. Paul sits at the top and decides what to handle himself, what to delegate and what to escalate; below him sit an analyst, a strategist, a writer and a reviewer. The chart forces decisions about responsibility and escalation before any prompt engineering starts.
A human in the loop is a permission boundary, not a credential. The role of the human is to limit what can proceed, not to lend identity to the agent.
Over-reporting is exponentially expensive. If every agent reports everything and asks permission for everything, the time to process a task grows disproportionately. The design goal is a model competent enough to delegate to a sub-agent without escalating or narrating each step.
"Please don't" is not an enforcement mechanism. A rule stated in a prompt can be lost as the context grows. Use hooks: deterministic logic embedded in a non-deterministic flow.
The organisational contribution is a platform AI catalog: agent definitions, flows and MCP server configurations become published assets, so one team publishes an agent and another reuses it. Skills are composable, and memory is scoped at three levels — user, project and organisation.
14. Call now, fetch later: MCP Tasks, durable state and event logs
MCP's Tasks primitive makes tool calls asynchronous. A request returns a durable handle immediately and the result arrives later. This is the correct model for work measured in minutes or hours — ETL jobs, deep research, batch reasoning — because holding a connection open that long is impractical. Note that in the 2026-07-28 specification covered in the day-two keynote, Tasks moved out of the core and became a formal extension, reshaped around the stateless model.
The speaker's point is that the specification defines the shape but leaves the hard parts to implementers: where in-flight work lives, how a task survives a server restart, and how to deliver a result exactly once while letting multiple clients subscribe to it.
The argument: a task is a state machine, and the history of a state machine is an ordered log of its transitions, which makes an append-only event log the natural backend. Creation, status changes and completion become events; recovery becomes replay; and the server itself can be stateless — which, in hindsight, anticipated exactly where the protocol went. The failure modes examined were orphaned tasks with no owner, duplicate side effects from retried work, and the gap between at-least-once and exactly-once delivery. The 2026 roadmap still has gaps around retry and expiry semantics.
Observation from the session: this is the same problem class as durable execution in distributed systems. The reference pattern is an event log such as Kafka, but standing up and operating Kafka for this is heavy. FastMCP is a lighter alternative handling durable state, restart and the surrounding lifecycle without that operational cost, and is worth evaluating before building the log layer by hand.
FastMCP: https://gofastmcp.com/getting-started/welcome
15. The Modern AI Stack: Agents, MCP and Skills (Adewale Abati, Independent)
As engineering teams rapidly adopt AI coding agents, they quickly run into two major failure modes: context window bloat (where feeding an entire repo exhausts the model's effective attention and degrades reasoning) and probabilistic drift (where agents violate internal coding standards because instructions are buried in vague system prompts). Adewale Abati’s modern AI stack addresses these hurdles through a structured architecture of specialized, decoupled components.
Agents.md : Global Context & Repository Guardrails
A concise, root-level markdown file that provides persistent documentation covering development standards, repository architecture, directory boundaries, and non-negotiable coding rules. It is loaded automatically into the agent's baseline context on startup and prevents the agent from making wild architectural assumptions. Instead of guessing how routing or error-handling is structured in your codebase, the agent consults Agents.md as its primary point of reference.
How to:
Keep it lean (under 100-200 lines).
- Include project structure overview, preferred libraries, build/test commands, and hard prohibitions (e.g., "Never modify database migrations directly").
Skills.md : On-Demand Domain Knowledge
Modular, task-specific documentation packages that encapsulate repeatable domain knowledge, complex workflows, or specific API integrations. Instead of stuffing your entire prompt or Agents.md with every edge case for every third-party service, the agent dynamically discovers and loads only the relevant skill when a specific task is requested. It eliminates context window bloat and reduces high rediscovery costs.
How to:
Identify the specific task the skill should solve.
Create a folder named after your skill (e.g., skills/unit-testing/SKILL.md).
Add SKILL.md with name, description, and instructions.
Add optional scripts or assets if needed.
Model Context Protocol: Standardized External Bridges
An open protocol that standardizes how isolated agent environments securely access external data sources, task histories, live browsing, and internal enterprise tools. It decouples the agent runtime from custom API wrappers, allowing teams to swap or scale tools cleanly without rewriting agent logic.
How to: Deploy standardized MCP servers alongside your internal services, exposing read/write endpoints through a unified protocol interface rather than ad-hoc script integrations.
Hook Events: Deterministic Safeguards
Programmatic validation gates triggered at critical execution checkpoints (e.g., pre-commit, pre-deploy, or post-generation). Probabilistic prompts cannot reliably enforce business rules. Hook events ensure that mandatory compliance and security checks are executed deterministically in code rather than trusted to model behavior.
- How to: Write deterministic code checks (linters, static analyzers, test runners) that intercept agent outputs before they touch production or merge requests.
The Harness: The Unified Operational Environment
The orchestration layer that binds the LLM model, MCP servers, modular skills, and deterministic hooks together into a single, cohesive runtime. Without a harness, an agent is just a loose script interacting with an API. The harness provides the secure, guided, and observable operational envelope required to transition agents from fragile experimental toys into reliable engineering partners.
How to: Build or adopt a local/cloud harness (such as containerized worker pods or developer desktops) that initializes the session, injects Agents.md, mounts dynamic Skills.md modules on demand, routes tool calls through MCP, and blocks execution if hook validations fail.
16. Legal Implications Under EU Law When Deploying AI Agents (Mirela Takacs, Law Office Takacs Mirela)
The EU AI Act enforces a risk-tiered regulatory framework where legal obligations scale directly with an agent's level of autonomy, function, and potential impact on individuals. Scoping an agent’s capabilities during initial design determines its legal tier. Restricting an agent to advisory or administrative roles keeps it in lower-risk categories, drastically reducing technical and compliance overhead.
How to classify agent architecture
While earlier drafts or general discussions often referred to four tiers (including "minimal" or "limited" risk), the final text of the EU AI Act establishes that "Minimal" and "Limited" risk categories no longer exist in the final text. Systems previously categorized under these tiers generally fall out of scope of the regulation unless they trigger specific transparency rules.
According to the presentation and official regulatory structure (AI act Consolidated V), AI systemsagent architectures fall into the following operational boundaries:
Out-of-Scope / Low-Risk AI (Voluntary Compliance): Systems such as internal utility agents, developer code-refactoring helpers, or spam filters are largely outside the mandatory scope of the EU AI Act. While not legally regulated under high-risk tiers, providers and deployers are encouraged to voluntarily adopt codes of conduct to ensure trustworthiness.
High-Risk Systems: Automated decision-making systems in sensitive domains (e.g., candidate ranking, credit scoring, or customer service agents using real-time emotion recognition) or safety components of regulated products. These require formal conformity assessments, continuous risk management, detailed audit logging, and mandatory human-in-the-loop oversight.
Prohibited Practices (Unacceptable Risk): Practices like social scoring, manipulative AI, or certain biometric categorizations are banned outright.
Systems such as internal utility agents, developer code-refactoring helpers, or spam filters are largely outside the mandatory scope of the EU AI Act. However, both providers and deployers have to formulate documentation disproving these qualifications and keep the documentation in their internal processesproceses. While not legally regulated under high-risk tiers, providers and deployers are encouraged to voluntarily adopt codes of conduct to ensure trustworthiness.
UUnder Article 50 of the EU AI Act, explicit cross-cutting transparency obligations apply regardless of an application's overall risk tier across four specific operational scenarios:
- when AI interacts directly with people (such as customer service bots or conversational agents),
- when AI generates synthetic content,
- when AI performs emotion recognition or biometric categorization,
- when AI creates deepfakes or text published on matters of public interest.
In all such cases, users must be clearly informed at the time of first interaction, and strict technical documentation and provenance standards must be maintained.
The "Post-Deployment Drift" Challenge
Unlike static software, autonomous agents dynamically select tools, execute open-ended plans, and evolve behaviors across production interactions. Unmonitored behavioral drift in production can legally invalidate an agent’s initial conformity assessment under EU law, forcing immediate operational suspension and mandatory re-auditing.
How to implement technical safeguards: Treat legal guardrails like automated CI/CD testing pipelines and static security analyzers:
- Deterministic Validation Gates: Enforce strict input/output schema validation in code rather than relying on probabilistic prompt instructions.
- Immutable Audit Logging: Capture structured telemetry of every reasoning step, tool call, and state transition.
- Human-in-the-Loop Escalation: Build programmatic review gates that intercept sensitive or boundary-crossing decisions before execution.
Key Takeaways
- Shift Legal Alignment Left: Treat regulatory rules as non-functional requirements during architecture planning rather than retrofitting them prior to deployment.
- Deterministic Boundaries > System Prompts: Enforce safety and compliance constraints strictly in code (via gateway proxies and schema checkers) because probabilistic LLMs cannot guarantee rule adherence.
- Auditability is a Core Feature: Structured, tamper-evident execution tracing is required not only for system debugging, but also as legal evidence to satisfy regulatory audits.
17. Sponsored Session: Breaking the Governance Bottleneck: Why Scaling Control Gets Agents to Production (Dr. Rania Khalaf, WSO2)
Enterprise adoption of AI agents frequently stalls in pilots due to non-deterministic behaviors and unmanaged agent sprawl, where teams deploy disparate agents across multiple frameworks without centralized control. To solve this governance bottleneck, WSO2 Agent Manager provides an open control plane that decouples agent governance from underlying application logic, establishing a clean separation of concerns (see repo here)
The platform enforces security, role-based access, and over 40 built-in guardrails—such as PII masking and prompt injection detection—across the agent, Model Context Protocol (MCP), and LLM layers. By combining zero-code OpenTelemetry auto-instrumentation for runtime observability with verifiable agent identity, token exchange, and instant revocation capabilities, organizations can safely manage the full agent lifecycle. This framework-agnostic architecture empowers engineering teams to scale production deployments securely, ensuring compliance with regulations like the EU AI Act while avoiding vendor lock-in.
18. Two Users, One App: Designing MCP Apps for Humans and Agents (Florian Bauer, OpenAI )
An MCP app is an interactive user interface extension built on the Model Context Protocol that enables AI agents to deliver rich, visual applications directly inside host environments, solving the limitations of text-only interactions by dynamically rendering components like charts, forms, and dashboards based on real-time data.
Designing these applications introduces unique state-management challenges stemming from concurrent multi-writer execution, where a human user interacting synchronously with the UI and an asynchronous AI agent running background loops simultaneously attempt to modify the underlying application state, such as JSON payloads, dashboard metrics, and form values. Because these independent actors update the application data model without waiting for one another, state collisions are inevitable, and naive systems risk silently overwriting user inputs or crashing the interface unless robust synchronization guards are established.To resolve this, MCP apps establish synchronized session semantics supported by two primary patterns: revision-based tracking for immediate state updates, and isolated draft states that hold agent-generated modifications until explicitly reviewed and accepted by the human user.
Revision-Based Tracking for Immediate State Updates
A concurrency control mechanism (similar to Optimistic Concurrency Control or vector clocks in distributed databases) that tags every state payload with an incremental revision number or hash. It ensures that asynchronous agent updates and synchronous user actions can be safely interleaved without silently clobbering each other's changes.
How it works:
- Version Tagging: Every state object or document node includes a strict revision_id integer (e.g., rev: 42).
- Conditional Writes: When the agent or user submits a state modification, the payload includes the base revision it was derived from (base_revision: 42).
- Conflict Detection & Rebase: The host runtime checks the current active revision. If the current revision matches base_revision, the write is committed and the revision increments (rev: 43). If a concurrent write has already occurred (e.g., current revision is now 43), the incoming update triggers a rebase or diff-merge pipeline rather than a destructive overwrite.
It provides real-time responsiveness for shared dashboards and collaborative views, allowing background agents to push live telemetry or calculated updates instantly without risking race conditions or wiping out user inputs.
Isolated Draft States for Safe Agent Modifications
A staging sandbox pattern (analogous to a Git branch or pull request workflow) that forces all agent-generated modifications into a segregated, non-destructive working memory space. Instead of directly mutating the live production state, the agent writes to an isolated shadow state or draft layer.
How it works:
- Ephemeral Shadow State: When an agent initiates a structural change (e.g., rewriting a dashboard layout or updating form fields), the runtime spins up an isolated JSON patch or draft branch bound to the session key.
- Diff Rendering: The UI renders both the live user state and the agent's proposed draft side-by-side or using visual diff overlays (additions highlighted in green, deletions in red).
- Explicit Human Gate: The agent's modifications remain dormant until the human user evaluates the diff and explicitly triggers an "Accept" or "Reject" action. Accepting merges the draft patch into the main state; rejecting purges the sandbox buffer cleanly.
It prevents unconstrained agent autonomy from corrupting production workflows or introducing silent errors. Users retain absolute sovereignty over the application state, transforming the agent from an unpredictable autonomous actor into a safe, collaborative co-pilot.
By decoupling rendering logic and enforcing strict boundary validation, developers can build secure, collaborative environments where agents seamlessly augment human workflows without overwriting active work, compromising control, or corrupting state.
19. Six Months of Proof: Independently-Verifiable Records for Agent Actions Under the EU AI Act (Steven Mih, Action State Group, Inc.)
The EU AI Act mandates automatic, long-term record-keeping for high-risk AI systems under Articles 12 and 19, requiring organizations to trace and prove what an autonomous agent actually did. However, standard application logs fall short because operators are not disinterested witnesses, and logs stored on internal infrastructure can be silently altered, rotated, or rewritten after the fact.
To solve this evidentiary gap, teams can implement cryptographic anchoring by committing each agent action record to a hash, chaining them together in an append-only structure, and registering them with an independent transparency log using open standards like SCITT (Supply Chain Integrity, Transparency, and Trust) and COSE (CBOR Object Signing and Encryption) . This approach ensures that any third party or auditor can independently verify the exact history of an agent's tool calls and decisions offline, providing verifiable proof that satisfies regulatory compliance without requiring blind trust in the system operator.
Implementation Architecture
Implementing verifiable audit records requires a three-tier pipeline that transitions traditional logging into provable cryptographic evidence:
- Structured Event Capture
Every meaningful agent action must record an immutable tuple containing the initiating identity (user ID, agent role), the exact request/prompt, the sources accessed (documents, APIs, versions), the policy decision, and the downstream action executed.
How : Wrap agent tool-call dispatchers in middleware that automatically extracts trace IDs, correlation IDs, and session keys into a standardized JSON schema before execution.
- Cryptographic Chaining
Instead of inserting raw log rows into a mutable database table, each event record is hashed and cryptographically linked to its predecessor.
How : Organize log entries into a binary Merkle tree structure. The root hash of the tree provides a single cryptographic commitment representing the entire history of agent actions up to that timestamp. Any modification to a past log entry alters its hash, immediately breaking the root commitment and rendering tampering mathematically detectable.
- External Anchoring & Transparency Logs (SCITT & COSE Receipts)
Publishing the cryptographic commitments to an external, append-only transparency log (using open IETF standards like SCITT—Supply Chain Integrity, Transparency, and Trust—and COSE envelopes).
How : Periodically anchor the Merkle tree root or signed tree heads (STH) to an immutable public or independent ledger (such as Azure Confidential Ledger). The transparency log returns a signed COSE receipt containing a Merkle inclusion proof. Auditors can use this receipt to independently verify offline that an agent action existed in the official log at a specific time without relying on the operator's infrastructure.
Assessment
By shifting compliance from subjective trust to mathematically verifiable auditability, this model adapts the proven principles of Certificate Transparency to AI governance. It requires careful trade-offs, including added operational demands, key management infrastructure, and continuous alignment with GDPR rules (such as maintaining cryptographic hashes on-chain while storing personal data in compliant, purgeable systems). While overkill for lightweight, low-stakes internal agents, cryptographic anchoring is the gold standard for provable accountability whenever agents execute high-risk automated decisions under the EU AI Act across regulated sectors like healthcare, finance, or HR.
20. Governance You Can Run: Checkable Properties for Production Agents (Seshu Tolety, Siemens)
Enterprise adoption of AI agents frequently stalls due to unmonitored risks like shadow AI (unauthorized, ungoverned agent deployments bypassing IT), compliance breaches, and operational chaos. To overcome this, organizations must replace static, wiki-based compliance manuals with programmatic, continuous governance centered around four core capabilities: Strategy, Engineering, Operations, and Trust
Four Governance Capabilities into Checkable Properties
To bridge the gap between high-level governance and engineering execution, each core capability must be translated into concrete, machine-testable invariants using industry standards and open-source tooling:
- Strategy
Ensuring that AI initiatives align with business objectives, risk tolerances, and authorized model portfolios. It eliminates unauthorized shadow AI by ensuring no agent can be built or deployed outside approved architectural and strategic boundaries.
How to turn it into a checkable property: Codify approved model versions, allowed data classifications, and business domain boundaries into repository-level manifest files (e.g., model catalog YAMLs). Use static analysis linters and CI/CD policy-as-code engines (such as Open Policy Agent (OPA)) to assert that any agent deployment references only approved foundational models and authorized data scopes.
- Engineering
Enforcing secure engineering practices across requirement translation, prompt/model management, and agent tool-use execution.Probabilistic prompts cannot guarantee rule adherence. Codifying policies into checkable properties ensures code-level boundary enforcement before agents reach production.
How to turn it into a checkable property: Define checkable behavioral invariants for agent tool calls (e.g., "The agent never executes unapproved database write APIs" or "The agent never exposes PII in tool parameters"). Implement these using runtime evaluation frameworks like NVIDIA NeMo Guardrails or prompt testing tools like Promptfoo to intercept reasoning steps and block invalid tool invocations.
- Operations
Managing production scaling, operational telemetry, cost control, and runtime behavioral monitoring.It protects the organization from runaway cloud costs, latency degradation, and unmonitored operational anomalies in production.
How to turn it into a checkable property: Establish automated threshold checks for Agent FinOps and latency overhead using OpenTelemetry standards and LLM tracing platforms (such as LangSmith or Arize Phoenix). Configure CI/CD or runtime circuit breakers that automatically flag or halt agent sessions exceeding token burn rate limits or exhibiting infinite-loop behavioral patterns.
- Trust
A cross-cutting capability ensuring compliance with external regulations (EU AI Act, GDPR) and management standards (ISO/IEC 42001, NIST AI RMF). It transforms compliance audits from subjective questionnaires into objective, reproducible evidence packages backed by continuous technical test results.
How to turn it into a checkable property: Translate legal clauses and AI management controls into automated integration tests and compliance assertions. For example, automatically verify human-in-the-loop escalation paths for high-risk agent actions and run data minimization assertions on inputs/outputs.
Key Takeaways
Governance is Code: Translate abstract compliance rules across Strategy, Engineering, Operations, and Trust into machine-testable invariants.
Continuous Validation: Pair runtime observability with automated property checks to catch drift before it impacts users.
Standardized Mapping: Tie technical telemetry directly to recognized frameworks to streamline enterprise audits.
21. Spotify’s Bet on MCP and Investment in Open Source (Reinoud Kruithof & Yannick Epstein, Spotify)
Integrating Model Context Protocol (MCP) servers into enterprise infrastructure presents severe architectural hurdles around authentication, credential safety, and scale, as shared by Reinoud Kruithof and Yannick Epstein from Spotify. When Spotify’s teams initially set out to integrate MCP servers, the open-source standard did not yet exist, forcing them to build a custom MCP gateway to interact with internal infrastructure.
The core technical bottleneck was authentication: the goal was to sign in once, generate a least-privilege token, and strictly keep credentials out of the AI agent's context while designing for continuous change. To solve this, they engineered an API that exchanges a valid user token for a task-dedicated token assigned to the agent via Role-Based Access Control (RBAC). However, a persistent challenge remained—engineers had to manually authenticate for every individual MCP server until the adoption of standardized authentication mechanisms like the Enterprise-Managed Authorization (EMA) extension.
EMA addresses this friction by establishing a delegated authorization flow where the enterprise Identity Provider (IdP) acts as an authoritative intermediary using an Identity Assertion JWT Authorization Grant (ID-JAG). This enables zero-touch OAuth, allowing employees to authenticate once through single sign-on so that all approved MCP servers are automatically connected without per-app consent screens or manual per-server token management.
Their key production lesson learned is that organizations must build upon established standards rather than custom solutions to handle the never-ending stream of changing user requests. By leveraging centralized gateways and scoped token exchanges, engineering teams can securely isolate agent permissions, prevent credential leakage, and scale MCP deployments across massive distributed systems without compromising security or architectural agility.
Favourite Tutorials
22. Agentgateway workshop: the TrendWatch agent (Solo.io)
A self-contained, hands-on workshop by Lin Sun and Eitan Suez, built around a sample agent called TrendWatch that reports what is currently hot in agentic AI. TrendWatch uses a set of MCP servers to read trending discussions, build and save digests on chosen topics, and publish them. The whole exercise runs on a laptop with a minimum of external dependencies.
Setup
Prerequisites are Python 3.11 or above, jq, docker, and a model. The default path uses Ollama with qwen3:8b running locally; a remote path targets Gemini's OpenAI-compatible endpoint on the free tier if you would rather not run a model locally. Jaeger runs in a container to collect distributed traces of the call flows between the agent, the LLM and the MCP servers — a good choice, because it makes the agentic loop visible as a trace rather than as log output. The agent itself is configured by three environment variables: the LLM endpoint, the model name, and the MCP server URL.
The exercises
- Introduction. Run the agentic loop locally: the agent calls the model, the model requests the trending_discussions tool, the tool fetches and filters AI-related discussions from HackerNews, and the agent presents the result. The logging shows the configuration, the tools visible to the model, and per-turn token consumption and tool calls.
- Proxy LLM traffic. Put agentgateway in the path between the agent and the model, so model traffic is routed and governed centrally.
- Proxy MCP traffic. Do the same for the MCP servers, which is where the control-plane argument from the day-one keynote becomes concrete.
- Prompt injection. Apply a prompt guard at the gateway rather than inside the agent.
- Expose OpenAPI as MCP tools. Take an existing OpenAPI specification and surface it to the agent as MCP tools without writing an MCP server, which is the most immediately reusable exercise for anyone with existing APIs.
- MCP authentication. Add authentication in front of the MCP servers — the same enforcement point Zalando's identity broker plugs into.
Workshop: https://solo-io.github.io/trendwatch/
Why it was worth doing
It is the day-one keynote made runnable. The control-plane argument is easy to agree with in the abstract and hard to evaluate without seeing what a gateway actually intercepts; this workshop puts a real agent, real MCP servers and a real trace viewer on your machine in under an hour. The OpenAPI-to-MCP exercise and the authentication exercise are the two most directly transferable to an existing enterprise estate.
APPENDIX A — CONCEPT GLOSSARY
MCP (Model Context Protocol). Open protocol through which agents discover and call external tools, resources and prompts. A host runs an MCP client; each MCP server fronts a system. Transports are STDIO (local child process) and Streamable HTTP (remote).
Stateless MCP. The 2026-07-28 revision. Removes the initialize handshake and the Mcp-Session-Id session so every request is self-contained and any instance can serve it. Adds server/discover, cacheable list results with TTL hints, HTTP-native routing headers, and multi-round-trip elicitation.
server/discover. Method a server must implement to advertise its supported protocol versions, capabilities and identity, replacing what the handshake used to carry.
MCP Tasks. Primitive for long-running tool calls: the server returns a durable handle and the client polls or subscribes and fetches the result later. Now a formal extension rather than part of the core.
WebMCP. W3C Web Machine Learning Community Group draft, developed by the Chrome and Edge teams, letting a web page register structured tools through a browser API so agents call functions instead of scraping the DOM. Tools run client-side with the user's session. In early Chrome preview.
ADK (Agent Development Kit). Google's open-source multi-agent framework, providing sequential, parallel and LLM-routed composition primitives.
Hooks and callbacks. Deterministic code executed at defined points in an otherwise non-deterministic agent flow. Used for anything that must always happen: PII checks, cache invalidation, permission checks.
Fan-out. Running agent tasks in parallel, each writing its output to its own state key for downstream agents to consume.
ACL / RBAC / ReBAC. Three access-control models: explicit per-resource lists, role-mediated permissions, and permissions derived from relationships between entities.
Zanzibar. Google's global authorization system, published in 2019, implementing ReBAC as a relationship graph with computed permissions. SpiceDB is the leading open-source implementation.
AAuth. IETF draft protocol giving each agent a cryptographic identity and replacing bearer tokens with signed HTTP requests, with explicit delegation and graduated human-approval requirements.
RFC 8693. OAuth2 Token Exchange. Lets a trusted party swap one token for another scoped to a specific resource. The mechanism behind Zalando's broker, and the reason the agent never holds a provider credential.
OPA (Open Policy Agent). CNCF general-purpose policy engine. Policies are written in Rego and evaluated outside application code, so tool-level authorization can change without redeploying the agent or the MCP server.
agentgateway. Open-source gateway sitting between agents and models or MCP servers, used as the enforcement point in both the Zalando architecture and the Solo.io workshop.
SPIFFE / SPIRE. Standard and implementation for issuing cryptographic identities to workloads. Mature, but scoped to the workload rather than the agent.
SAST / SCA. Static application security testing scans first-party source code; software composition analysis scans third-party dependencies.
SlopCodeBench. Benchmark measuring how code quality evolves when an agent repeatedly extends its own prior code under changing specifications. Tracks verbosity (redundant code) and structural erosion (complexity concentrated in few functions).
Claim ledger. Pattern where agents emit structured claims (entity, value, confidence, source) so that conflicts between agents can be detected and resolution correctness measured.
Utilisation and coverage. Handoff metrics. Utilisation is the share of passed context actually used; coverage is the share of required context that was present.
AI control plane. Central layer holding policy and configuration for agents, separate from the agents executing the work. Covers identity, authorization, policy, cost control and availability.