Skip to content
agentgateway has joined the Agentic AI Foundation — Learn more

For the complete documentation index, see llms.txt. Markdown versions of all docs pages are available by appending .md to any docs URL.

vLLM Semantic Router

Page as Markdown

Route LLM requests by prompt semantics with vLLM Semantic Router and agentgateway.

vLLM Semantic Router (vSR) classifies LLM requests and selects a model based on prompt content. With agentgateway, you can make this semantic decision before routing while continuing to apply gateway policies and record model, token, latency, and cost telemetry. See the vSR Router API reference for supported frontend and backend API types.

This integration is distinct from using vLLM as an inference provider. vSR provides the model-selection policy while your configured backend serves the selected model.

To connect to vLLM’s OpenAI-compatible API, see Custom providers.

How the integration works

A client sends a request to agentgateway. Agentgateway exchanges the request and model decision with vLLM Semantic Router through ExtProc, enforces model access, forwards to the selected backend, and records telemetry.
A client sends a request to agentgateway. Agentgateway exchanges the request and model decision with vLLM Semantic Router through ExtProc, enforces model access, forwards to the selected backend, and records telemetry.

  1. Process the request: A client sends an LLM request to agentgateway, which calls vSR as an external processor (ExtProc) before model selection.
  2. Select a model: vSR evaluates its configured signals, such as prompt content or caller tier, and returns a processing response that updates the request’s model field.
  3. Forward the request: Agentgateway uses the selected model to choose a configured backend, applies gateway policies such as model authorization, and forwards the request for inference. Agentgateway records usage and latency for observability.

When response caching is enabled, vSR can return a cached completion through ExtProc, allowing agentgateway to respond without calling the backend.

Choose an integration path

Use the following guides and examples to configure the integration.

Follow the example README to configure credentials, start the containers, send requests, and clean up.

To request an additional integration example, create an issue in the agentgateway repository and describe your use case.

Integration considerations

  • Client model selection: The examples use model: "auto" to opt in to semantic selection. This value is an example policy convention, not a reserved agentgateway model. If clients can request models explicitly, enforce model entitlements in agentgateway as well as in vSR decisions.

  • Backend choice: vSR can select models served by hosted providers or self-hosted inference workloads. Configure the corresponding LLM provider or routing backend in agentgateway.

  • Model names: Keep the names returned by vSR aligned with the models in your agentgateway routes, provider configuration, and cost catalog.

  • Trusted tier context: The single-runtime examples use caller-supplied user ID and tier headers to demonstrate routing, not authentication. Before exposing them to users, validate identity and bind those headers to trusted entitlements before ExtProc runs. API key authentication alone does not verify a caller-supplied tier.

  • Cost and observability: vSR makes the semantic decision. Agentgateway remains the source for completed-request telemetry and can calculate realized cost when you configure a model cost catalog. Use LLM metrics and logs to evaluate the result.

Before a broad rollout, compare routed traffic with a fixed higher-capability-model baseline. Confirm that the policy uses the intended model tiers, then evaluate task completion, user feedback, retries, and escalation rates alongside cost and latency.

Was this page helpful?
Agentgateway assistant

Ask me anything about agentgateway configuration, features, or usage.

Note: AI-generated content might contain errors; please verify and test all returned information.

Tip: one topic per conversation gives the best results. Use the + button in the chat header to start a new conversation.

Switching topics? Starting a new conversation improves accuracy.
↑↓ navigate ↵ select esc dismiss

What could be improved?

Your feedback helps us improve assistant answers and identify docs gaps we should fix.

Need more help? Join us on Discord: https://discord.gg/y9efgEmppm

Want to use your own agent? Add the Solo MCP server to query our docs directly. Get started here: https://search.solo.io/.