Kong’s AI capabilities ship as a set of plugins on top of the core Kong Gateway/Konnect data plane, rather than a separate product. The first six AI plugins shipped in Kong Gateway 3.6 (Feb 2024) as open source; more advanced governance plugins were added later and require Kong Enterprise or Konnect licensing.
Traffic proxying & routing
| Plugin | What it does |
|---|---|
| AI Proxy | The base plugin. Normalizes one request format across providers (OpenAI, Anthropic, Azure, Cohere, Bedrock, Gemini, etc.), translates it to the target’s native format, and transforms the response back. Open source. |
| AI Proxy Advanced | Multi-target version of AI Proxy — proxies to several providers/models at once and adds load balancing. This is what powers the failover setup we built earlier. |
| AI MCP Proxy | Fronts MCP servers (passthrough, conversion, or listener mode), covered in the registry/gateway duty above. |
| AI MCP OAuth2 | OAuth2 resource-server behavior specifically for MCP traffic. |
| AI A2A Proxy | Routes Agent2Agent (A2A) protocol traffic through Kong, with the same per-consumer auth and rate-limiting controls as other AI routes. |
Load balancing algorithms (inside AI Proxy Advanced)
Configured via the Upstream entity, not separate plugins:
- Round-robin (weighted) — proportional traffic split by weight.
- Consistent-hashing — sticky sessions via a hashed header (default
X-Kong-LLM-Request-ID). - Lowest-latency — routes to whichever target is currently fastest.
- Lowest-usage — routes based on token/cost usage, including a
costoption added to the usage strategy. - Priority — tiered failover (what we configured for Anthropic → Bedrock → OpenAI).
- Semantic — routes prompts to domain-specialized models based on embedding similarity, with
distance_metricandthresholdcontrolling match sensitivity. - 3.13+ added health checks and a circuit breaker for the load balancer itself, plus failover/circuit-breaking for semantic routing specifically.
Governance & safety
| Plugin | What it does |
|---|---|
| AI Prompt Guard | Blocks or allows prompts by pattern/keyword rules. |
| AI Semantic Prompt Guard | Same idea, but matches by embedding similarity rather than literal pattern — catches paraphrased attempts at a blocked topic. Enterprise. |
| AI Prompt Decorator | Injects or modifies system/user prompt content before it reaches the model (e.g. enforcing a system prompt). |
| AI Request/Response Transformer | General-purpose payload transformation for AI requests and responses. |
| AI Rate Limiting Advanced | Rate limits by token count and cost, not just request count — this is the plugin underpinning the per-team quota story from the credential-governance duty. Enterprise. |
Cost & performance
| Plugin | What it does |
|---|---|
| AI Semantic Cache | Caches responses by embedding similarity so repeated/similar prompts skip the model call entirely. Backed by a vector store (Redis and others). Enterprise. |
| AI RAG Injector | Pulls relevant context from a vector store and injects it into the prompt for retrieval-augmented generation. |
Observability
- Prometheus plugin was extended with AI-specific metrics (token counts, cost, latency, MCP call volume) — this is what feeds the shared observability layer from the architecture summary.
Licensing note
Per Kong’s own docs and third-party comparisons: AI Proxy, AI Proxy Advanced, AI Prompt Guard, and AI Prompt Decorator are open source. AI Rate Limiting Advanced, AI Semantic Cache, AI Semantic Prompt Guard, and AI MCP Proxy sit behind Kong Enterprise/Konnect licensing. Confirm current tier gating against Kong’s plugin hub for your specific version, since this has shifted release to release.
Good — pulled the current plugin hub listing directly. Here’s the accurate, up-to-date table, tagged by license tier and mapped to the four pillars from the architecture doc.
AI Plugin Tier Table (Kong Plugin Hub, current)
Free / open source (no AI license required)
| Plugin | What it does |
|---|---|
| AI Proxy | Base multi-provider proxy and normalization |
| AI Prompt Guard | Pattern-based allow/deny list for prompts |
| AI Prompt Decorator | Prepends/appends messages to chat history |
| AI Prompt Template | Fill-in-the-blank prompt templates |
| AI Request Transformer | LLM-based transform of the client request body |
| AI Response Transformer | LLM-based transform of the upstream response |
AI License Required (Enterprise/Konnect AI Gateway tier)
| Plugin | What it does |
|---|---|
| AI Proxy Advanced | Multi-target proxy + load balancing/failover |
| AI Rate Limiting Advanced | Token/cost-based rate limiting |
| AI Semantic Cache | Embedding-based response caching |
| AI Semantic Prompt Guard | Embedding-based topic allow/deny |
| AI Semantic Response Guard | Blocks/permits by similarity to known responses |
| AI PII Sanitizer | Redacts sensitive data in request/response bodies |
| AI Prompt Compressor | Shrinks prompts to cut cost/latency |
| AI RAG Injector | Injects vector-store content for RAG |
| AI MCP Proxy | MCP server proxying, conversion, aggregation |
| AI MCP OAuth2 | OAuth2 auth for MCP servers |
| AI A2A Proxy | Governance for Agent-to-Agent protocol traffic |
| AI LLM as Judge | Evaluates/optimizes model output quality |
| AI AWS Guardrails / AI Azure Content Safety / AI GCP Model Armor | Cloud-native guardrail integrations |
| AI Custom Guardrail | Bring-your-own third-party guardrails service |
| AI Lakera Guard | Lakera Guard safety-policy enforcement |
Third-party (separate vendor licensing)
CrowdStrike Falcon AIDR (Request/Response/MCP), Harness AI Security, Noma Runtime Protection, Prisma AIRS, Skyflow De-identify, Straiker (+ Coding Agent variants), Impart Security, VeriKnox.
Mapped to the four architecture pillars
| Pillar | Plugins that implement it |
|---|---|
| Credential Governance | AI Proxy / AI Proxy Advanced (vault-referenced auth blocks); AI PII Sanitizer for data leaving the gateway |
| MCP Registry & Gateway | AI MCP Proxy, AI MCP OAuth2 |
| Multi-Provider Failover | AI Proxy Advanced (priority/weighted/semantic/lowest-latency/lowest-usage load balancing, plus 3.13+ health checks and circuit breaker) |
| Observability | Prometheus (extended with AI metrics), OpenTelemetry, AI LLM as Judge for quality tracking |
| (Guardrails — a fifth pillar worth adding) | AI Prompt Guard, AI Semantic Prompt Guard, AI Semantic Response Guard, AI AWS/Azure/GCP Guardrails, AI Custom Guardrail, AI Lakera Guard |
One correction from earlier: I’d said AI Prompt Guard and AI Prompt Decorator might be gated — they’re actually free/open source. AI Proxy Advanced, AI Rate Limiting Advanced, AI Semantic Cache, AI Semantic Prompt Guard, AI MCP Proxy and AI MCP OAuth2 are the ones that need the AI license.
I’d also flag that Guardrails has grown into its own real pillar — with cloud-native integrations (AWS/Azure/GCP) plus third-party options (Lakera, CrowdStrike, Straiker) — separate from what I’d folded into “credential governance” earlier.