ARCHITECTURE SUMMARY
Kong AI Gateway — Unified Architecture
Single Ingress Point for All AI Model and MCP Traffic
Mission
Every AI interaction in the enterprise — app-to-model, agent-to-agent, agent-to-tool — flows through one gateway. No exceptions, no direct provider calls, no shadow credentials. This document ties together the four pillars of the architecture.
1. Single Ingress Point
All apps, agents, and MCP clients hit one endpoint family on Kong, never a provider directly. This is the architectural premise everything else depends on: if traffic can bypass the gateway, none of the governance below actually holds.
2. Credential Governance
- Every provider key (Anthropic, Bedrock, OpenAI, internal APIs, MCP servers) lives in a vault, never in app code or Kong config.
- Agents authenticate to Kong with their own Consumer credential — an identity Kong owns and can revoke independently of the underlying provider key.
- Kong resolves the real credential at request time and injects it server-side. A leaked client-side key exposes nothing about the provider account behind it.
- One Consumer model covers both LLM and MCP auth (OAuth2 for MCP), so access governance and credential governance are the same system.
3. MCP Registry & Gateway
- The MCP Registry (Konnect Catalog, tech preview) catalogs approved MCP servers — owner, purpose, status — so agents discover tools in one place instead of hardcoding server URLs.
- The AI MCP Proxy plugin fronts internal servers (passthrough), converts REST APIs into tools (conversion), or aggregates multiple servers into one endpoint (listener).
- Tool-level ACLs gate both discovery and invocation, so a Consumer’s permissions extend down to individual tools, not just “can call MCP or not.”
4. Multi-Provider Failover
- AI Proxy Advanced’s priority algorithm ranks providers into tiers (e.g. Anthropic → Bedrock → OpenAI). The top tier takes all traffic while healthy; the balancer only spills to the next tier on failure.
- Active and passive health checks on the Upstream detect a degraded provider and flip targets automatically — no manual intervention during an outage.
- Each tier still uses its own vault-backed credential, since Bedrock’s SigV4 auth, Anthropic’s API key, and OpenAI’s bearer token are structurally different.
5. Observability — The Thread Through All Three
- Metrics, audit logs, and cost data are captured centrally, keyed to Consumer identity, not to whichever provider or MCP server actually served the request.
- Failover events, tool invocations, and credential resolutions all land in the same logging/alerting pipeline, so “who called what, through which path, at what cost” is answerable from one place.
The Four Pillars at a Glance
| Credential Governance Vault-backed keys, Consumer-level auth, server-side injection. | MCP Registry & Gateway Catalog of approved servers, tool-level ACLs, proxy modes. | Multi-Provider Failover Priority groups, health checks, automatic tier fallback. | Observability Unified metrics, audit logs, cost tracking, failover alerts. |
Why This Design Holds Together
Take any pillar away and the “single ingress point” claim weakens:
- No credential governance → the gateway is a pass-through, not a security boundary.
- No MCP registry/gateway → agents discover and call tools outside Kong’s visibility.
- No failover → a single provider outage becomes a company-wide AI outage.
- No shared observability → incidents can’t be traced back to an owner or a root cause.
Together, they turn “Kong is where AI traffic goes” into “Kong is where AI traffic is governed” — which is the actual job of the architect role.