Benefits of A10 AI Gateway
Route and orchestrate LLM requests by users and teams, with speed and cost efficiency
Key Features
Complexity-aware Routing
Classifies each prompt as simple or complex and routes it automatically to the best model, without any manual model selection
Priority-chain Cascading
As budget is consumed, requests cascade to the next-best model in a chain you define, keeping traffic flowing without intervention
Virtual API Keys
Virtual keys map usernames to an abstraction layer that connects to any IAM system, each carrying its own budget and rate limits
Provider Credential Vault
Upstream credentials for OpenAI, Anthropic, Azure, and more live in an OpenBao-backed vault, never in config files
Real-time Cost Tracking
Spend is tracked live as requests route, giving teams immediate cost visibility instead of a delayed, end-of-month bill
Token Budgets
Cap spend at the individual model level, at user and team level, so you can experiment freely with budget-friendly models while premium spend stays controlled
Business-layer Rate Limiting
RPM and TPM caps enforced per key or team protect cost and shared provider capacity from a single noisy caller
Single-tenant Deployment
Deploys via Helm on Oracle Kubernetes Engine with full tenant isolation, no GPU required, though GPU-enhanced options are available
See How a Request Moves Through the A10 AI Gateway
Complexity and Budget-aware Smart Routing
Every request classified, checked, and routed before it reaches a model
A request enters with a virtual key, is classified for complexity, and checked against its budget. It’s then routed to the cheapest capable model in the best pool, falling back automatically if a budget is exhausted, with every step logged for cost tracking and audit
Frequently Asked Questions
Don’t see your question listed? Contact a product expert to get answers.
A lightweight classifier scores each prompt complexity in real time and routes it to the best model pool: simple tasks to smaller, faster models, complex tasks to more capable ones; without users picking a model by hand.
No. The AI Gateway orchestrates context across models, so even when different requests in the same conversation route to different models, each one still sees the full prior context needed to respond accurately.
Requests automatically cascade to the next-cheapest capable model in a priority chain you set, so traffic keeps flowing without manual intervention.
No. Classification and routing happen fast enough that the AI Gateway clears a typical request with low added latency.
It is deployed via Helm onto Oracle Kubernetes Engine with single-tenant options so each customer runs a fully dedicated, isolated deployment.
The A10 AI Gateway includes lightweight, pattern-based (regex) filtering that screens requests and responses against configurable rules with room to add stronger guardrails as your policies evolve.
Related Product
TrojAI by A10 Networks
TrojAI by A10 Networks secures every AI agent, application, model from build to runtime with agent-led automated red teaming, and the capability to detect, discover, and control MCP servers, tools, and agents.
