Provex, tested

Prove your AI defences hold, with evidence anyone can re-check.

Authorised attacks on your own AI, through your real controls or on a twin, with the exact attack, the exact reply and the verdict.

The Prove capabilities

Specs

Prove, in detail

Delivery and data

Delivery
SaaS, from one login.
Isolation
Each customer runs in an isolated workspace with its own database.
Certifications
None held. Frameworks are mapped to and assessed against.

Frameworks

OWASP Top 10 for Agentic Applications
Assessed per agent against: Each registered agent, in-path and by assessment.
OWASP LLM Top 10 (2025)
Covered in probes and scans: References carried by each probe.
OWASP MCP Top 10
Covered in probes and scans: MCP risks in probes and scans.
MITRE ATLAS
Coverage measured from runs: Exercised, exercisable and untestable, kept apart.
See the frameworks

Last reviewed 6 Oct 2026

Will our defences hold?

Continuous AI security validation in ColossalX attacks your own chatbots, APIs and agents, with your authorisation and through your real controls, then records the exact attack, the exact reply and the verdict in a sealed run. Probes that got through become a proposed fix that a person approves, and the same probes are re-sent to prove it.

A new jailbreak works on your support agent, and nobody finds out until a customer does.

ColossalX attacks the agent with your authorisation, through your real controls, and seals what got through.

Prove · 6 capabilities

Six capabilities answer one question, will our defences hold, and each ends in a record: a sealed run, a recorded decision, a ranked exposure or owned work.

Red-teaming and validation

Prove you own a target, then attack it with your authorisation.

Explore
More

A target can be a chatbot, an API, an authenticated session, or a registered agent in-path through your own gateway. Each attempt is judged blocked, detected, missed or refused by the model, and each break quotes the exact reply. Runs are sealed and re-checked whenever they are read.

Run · support-bot · in-pathIllustrative

Run · support-bot · in-path: illustrative run with 24 attempts blocked by a control, 4 detected but allowed, 3 missed and 5 refused by the model. Each break quotes the reply.

ColossalX CyberTwins

Attack a twin of your agents, not production.

Explore
More

A snapshot of your agents and their guardrails is stood up as shadow agents and attacked, and breaches are filed as findings. A guardrail change is tried on the twin first and promoted only if it held, and promotion keeps the rules it replaced, so a revert is one step.

Twin · one guardrail changeIllustrative

Twin · one guardrail change: 09:00 Replica captured (agents and guardrails); 09:20 Shadow agents attacked; 09:41 Breach found (tool misuse, ASI02); 10:15 Change tried on twin (re-attacked, it held); 10:30 Promoted to production (revert kept).

Threat intelligence

Intelligence matched to what you actually run: your SBOM and AI-BOM, models, agents and vendors.

Explore
More

Known-exploited and exploit-likelihood feeds, advisories, STIX bundles and a daily AI supply-chain watch arrive in one place. A person accepts what applies, and each threat ends in a recorded decision: fix it, test it, watch it or set it aside.

Intel · matched, then decidedIllustrative

Intel · matched, then decided: known-exploited to your estate; advisory found calling your estate (matched); STIX bundle to your estate; your estate to owned work (affected); your estate to watchlist (no fix yet); your estate to recorded no.

Exposure management

Fix first what your own tests proved reachable.

Explore
More

Exposures from in-path coverage gaps, runtime guard firings and exploited dependencies are ranked by validated reachability and the business criticality of the agent they hit, not by raw severity. A defence that slips raises a regression alert, and the CISO review reads incomplete when a source could not be read.

Coverage drift · support-botIllustrative

Coverage drift · support-bot: Last run: "ASI01 goal hijack · blocked"; This run: "ASI01 goal hijack · missed". Regression check: held to missed. Verdict: monitored, Regression alert, ranked first.

Code security

Secure what you build, from repository to running app.

Explore
More

Repositories are scanned on each push for leaked secrets, exploitable dependencies, insecure code and exposed personal data, and model files are inspected, never run. Live app scans gate only on new problems, and a CI/CD gate with templates for 6 CI systems writes SARIF into your own pipeline.

CI gate · refunds-agentIllustrative

CI gate · refunds-agent: CI pipeline to deploy · refunds-agent, "git push · refunds-agent adds a dependency". Checks: Leaked secrets passed, Known-exploited dependency failed, OWASP Agentic assessment passed, Guardrail profile passed. Verdict: refused, Gate failed, ticket raised.

Resilience

Break a provider on purpose.

Explore
More

Real faults, an outage, errors, rate limits or slow answers, are injected into AI traffic through the gateway with a blast radius you choose, and the experiment rolls itself back if it hurts more than planned. The newest backup is restored into a scratch database each week and timed, and a failed drill opens owned work.

Chaos test · provider outageIllustrative

Chaos test · provider outage: Fault Provider outage; Blast radius One provider; Served by failover Measured per request; Floor Stops itself below it; Verdict Hypothesis held. Measured, not asserted.

How it works

From a probe that got through to a proven fix.

One probe that got through, followed from the sealed run to a guardrail change a named person approves, and the same probes re-sent to prove it held.

Workflow · a probe that got through, proven fixedIllustrative

01 Run sealed

An in-path run attacks support-bot through your gateway, then seals it.

02 Fix proposed

What got through becomes a proposed guardrail change; nothing is applied yet.

03 Person decides

A named guardrail owner applies it, or rejects it with a reason.

04 Same probes re-sent

Exactly the probes that landed are re-sent; the issue closes on evidence.

How it connectsx, tested

Where a tested x goes next.

A gap a run finds does not end in a report. It becomes owned work, a risk in money and evidence an auditor can re-check.

  1. Each proven gap becomes one issue with an owner and a due date, closed only on evidence.

  2. A proven exposure writes a risk entry, and leaves it when a control stops the attack.

  3. What got through is proposed as a guardrail change, rolled out observe, canary, enforce.

  4. Recent red-team, campaign and recovery tests count as evidence for the controls they exercised.

Honest by design

What it does, and what it does not.

A scenario it cannot runIllustrative

A scenario it cannot run: AEV scenario to payments-agent, "Lateral movement across the payments network". Checks: Network layer offered failed. Verdict: refused, Not executable, never passed.

What it does not do

x, not measured

It attacks AI systems and agents only: network, email, endpoint and directory attacks are not offered.

All 4 limits
  • Testing sends real adversarial prompts on your own model keys, so a run spends model budget.
  • Threat intelligence does not collect dark web, NVD or national CERT feeds, and says so on screen.
  • Twin discovery is AWS only, and the twin is a configuration snapshot, not a copy of your systems.

How we know

  • Each break quotes the exact span of the reply that proves it: no quote, no break.
  • A model refusing on its own is reported apart from a control that stopped the attack.
  • Coverage counts only what a control noticed; a landed attack that nothing saw is a gap.
  • A scenario with no way to run is recorded not executable, never given a pass.

Questions

Questions buyers ask

What is adversarial exposure validation?

Adversarial exposure validation, or AEV, tests whether an exposure can actually be exploited by attacking it and recording the evidence, instead of assuming a control works. ColossalX applies it to AI: authorised attacks run against your own chatbots, APIs and agents, through your real controls, and each result says which control held and which did not.

How is AI red teaming different from penetration testing?

Penetration testing usually probes networks, hosts and applications for exploitable flaws. AI red teaming attacks how a model or agent behaves: prompt injection, jailbreaks, data leaks, tool misuse and memory poisoning. ColossalX does the second and states plainly that network, email, endpoint and directory attacks are not offered, so it sits beside a penetration test rather than replacing one.

How often should AI defences be re-tested?

Whenever something changes, and on a schedule in between. In ColossalX, registered agents can be re-attacked daily, re-attacked within minutes of a change to their guardrails, configuration, access or status, and attacked with a newly published technique the day it is described. A defence that used to hold and now misses raises a regression alert.

Can we test without attacking production?

Yes. ColossalX CyberTwins captures a snapshot of your agents and their guardrails, stands it up as shadow agents and attacks those instead. A guardrail change can be tried on the twin and promoted only if it held. The twin calls the real model, so attacks still spend model budget, and the page says so before each run.

How does a test result become a fix that is proven?

Probes that got through become a proposed guardrail change. A named person applies it or rejects it with a reason, and exactly the probes that landed are re-sent to measure the change, which can be reverted. The issue that tracked the gap closes only on positive evidence, and the proven exposure leaves the risk register when a control stops it.

Is it safe to let a product attack our AI?

Testing is authorised and scoped. You prove you own a target before a run, by a DNS record, a registered agent or a written authorisation. A scope limits what a run may send, each authorisation is kept in an append-only trail, and one switch stops the runs in flight. Unverified targets are refused once you switch enforcement on.

The other verbs

Next step

Know your x, tested.

See an authorised run on your own agents: the exact attack, the exact reply, and what your controls did.

  1. 01Tell us what you run
  2. 02See the four verbs on it
  3. 03Decide where to start