ControlAI gatewayx, held
AI gateway: one governed path for each model call.
Change an agent's base URL and key, and it runs under its own credential, against your policy, across 31 provider families.
One path, many models: support-bot to ColossalX; kyc-agent to ColossalX; unknown caller found calling ColossalX (flagged); ColossalX to provider A (allowed); ColossalX with data redacted, to self-hosted (PII redacted); ColossalX refused before retired (not offered).
In shortAI gatewayAn AI gateway is one path between applications and the models they call, where requests are routed, recorded and checked. A routing gateway mainly balances cost and availability; a security-focused gateway adds an identity for each caller, guardrails on prompts and answers, and the power to refuse a request before it reaches the model. In the glossary
The ColossalX AI gateway is one governed path for model calls. It is OpenAI-compatible, so an agent changes its base URL and key and runs under its own credential, across 31 provider families including self-hosted models. Your keys are sealed at rest, each team reaches only the models you allow, and each request keeps a decision record.
Teams call models with shared keys from scattered code, and nobody can say which controls applied.
ColossalX routes each call through one gateway, under the agent's own credential, and records what was decided.
How it works
From a changed base URL to a governed request.
One existing agent is pointed at the gateway with its own key. Followed through the checks it meets, the provider that answers and the record it leaves.
01 Pointed at the gateway
An existing agent changes its base URL and uses its own key.
02 Checked first
Switches, consent, model access and content checks run before routing.
03 Routed
A provider outage fails over; a policy decision is never retried.
04 Recorded
The request is stored with its decision record and its cost.
What you see
Traffic and blocks, from real requests.
The gateway overview shows requests against blocks each hour and what each policy matched and blocked, drawn from the requests it actually served.
- One endpoint, any agent
- Your keys, sealed
- Models by role
- Know what ran
Read the detail, step by step4
- One endpoint, any agent. OpenAI-compatible: change the base URL and key, and nothing else. Each agent gets its own credential, so a request is attributed to that agent whatever its body claims. Apps that cannot sit behind a proxy call the same content and injection checks directly.
- Your keys, sealed. Provider keys are sealed at rest and never shown again. A key that cannot be sealed is refused rather than stored in the clear. Each response records whose key served it.
- Models by role. Each team and use reaches only the models you allow. Model access policies set which roles and uses may reach which models, with daily request and spend allowances. Switch any model off in one place; a model its vendor retired is never offered.
- Know what ran. Each policy row says whether it is in force, and why not. A control written but not enforced reads "not enforced". Each request keeps a decision record, and requests where a control could not run are counted, so a quiet gate shows up.
An illustrative gateway decision log: each request crosses the gateway and ends allowed, redacted, held for a named person or refused, with the reason, and an unknown caller without a credential is found. A header confirms the controls are in force.
3notes
- Each control confirmed in force
- Held for a named person
- A caller with no credential, found
How it connectsx, held
Where a governed request goes next.
The gateway is where each other control takes effect. A request that passes through it is checked, routed and recorded.
Prompts and answers are checked for injection, personal data and unsafe content.
Tool calls are checked against default-deny rules before they reach a server.
Captured traffic shows which agent sent which personal data to which model.
Tokens and cost per request roll up by model, provider and conversation.
Honest by design
What it does, and what it does not.
When a check cannot run: kyc-agent to provider A, "Summarise the customer file for the review". Checks: Kill switch passed, Content check waiting, Gap recorded on request flagged. Verdict: allowed, Allowed, gap recorded.
What it does not do
x, not measured
Whole responses only: answers are checked complete, so replies are not sent piece by piece.
All 4 limits
- Most checks fail open if they cannot run, and the gap is recorded on the request.
- The per-request decision record is read through the API and session replay, not a gateway screen.
- Bring your own key means model provider keys, not customer-managed encryption keys.
How we know
- Provider health is scored from real traffic, and ColossalX's own blocks never count as a vendor's failure.
- A provider with no traffic scores nothing, never a perfect score.
- A key that cannot be sealed is refused, never stored in the clear.
- A request that names a model no enabled provider carries is recorded as a substitution.
Questions
Questions buyers ask
What is an AI gateway?
An AI gateway sits between your applications and agents and the model providers they call, so each call passes one point where identity, policy and content checks apply and a record is kept. In ColossalX it is OpenAI-compatible, so existing code changes a base URL and a key rather than its logic.
What is the difference between an AI gateway, an AI firewall and an API gateway?
An API gateway routes and meters API traffic without reading prompts. An AI firewall inspects prompts and answers for attacks and leaks. An AI gateway in ColossalX does both for model calls: it routes and meters by model, role and agent, and applies guardrails, tool rules, consent and kill switches on the same path.
How do we put an existing agent behind the gateway?
Give the agent its own credential and change its base URL to the ColossalX gateway endpoint. Nothing else in its code changes. From then on each request is attributed to that agent, checked against its policy and recorded. An app that cannot sit behind a proxy can call the same checks directly.
Which model providers and self-hosted runtimes are supported?
31 provider families: the major hosted providers, cloud AI platforms, and self-hosted runtimes such as Ollama, vLLM, LocalAI and LM Studio, plus any OpenAI-compatible endpoint. Models are discovered when a provider is added, and you choose which roles and uses may reach each one.
How do we know a policy is really in force?
The policies page says, row by row, whether each control is really in force for your workspace and, if not, why. Each request keeps a decision record, and requests where a control could not run are counted rather than hidden. That is what ColossalX means by know what ran.
Related
Where to look next.
-
Runtime guardrails
Injection, data leaks and approval holds
-
ColossalX MCP Firewall
Allow, monitor or block each tool
-
AI spend
Cost by model, with daily allowances
Next step
Know your x.
Point one of your own agents at the gateway and see what it calls, what was held and what ran.
- 01Tell us what you run
- 02See the four verbs on it
- 03Decide where to start