ProveResiliencex, tested
Resilience: real faults, proven recovery.
Break a model provider on purpose, measure the failover, and prove the newest backup restores, with each failure turned into owned work.
Provider outage · injected on purpose: research-agent to gateway (real traffic); gateway refused before provider A (injected outage); gateway to provider B (served by failover).
In short
Resilience in ColossalX tests whether your AI keeps working when something fails. Chaos experiments inject real faults, outages, errors, rate limits, slow or truncated answers, into AI traffic through the gateway, measure failover and roll back on their own. Weekly restore drills prove the newest backup restores, and failures become owned work.
A model provider goes down mid-morning, and nobody knows whether failover really works.
ColossalX causes that outage on purpose, measures what failover served, and rolls back if it hurts too much.
How it works
From a hypothesis to a measured, rolled-back fault.
One provider is taken down on purpose, inside a window you set. Failover is measured request by request, and the fault rolls back on its own.
01 Hypothesis set
You choose the fault, the blast radius and what holding means.
02 Fault injected
Real AI traffic meets the outage, request by request, through the gateway.
03 Failover measured
Served, failed over, failed and latency are measured, not estimated.
04 Rolled back
The window ends, the fault is removed and the verdict is recorded.
What you see
The traffic a chaos experiment runs through.
Faults are injected into the same gateway traffic you watch every day: requests per hour, blocks, provider errors and latency at the median and the tail.
- Choose the faultOutage, unreachable, errors, rate limits, slow or truncated answers.
- Inject into real trafficReal AI traffic through the gateway meets the fault, request by request.
- Rolls itself backIt stops itself below your floor and expires with its window.
- Prove the restoreThe newest backup is restored into a scratch database weekly and timed.
Read the detail, step by step4
- Choose the fault. Pick a blast radius: one agent, one provider, or a share of the workspace that is never more than half. Set the window, the share of requests hit and what counts as holding, and write a hypothesis; a sensible one is suggested.
- Inject into real traffic. Live and final figures show covered requests, faults fired, requests still served, served by failover, failed, and latency at the median and the tail. One experiment runs at a time, and Stop now ends it.
- Rolls itself back. The verdict reads hypothesis held, hypothesis failed, stopped, or not measured with the reason. Only measured experiments count toward the resilience pillar of the trust score.
- Prove the restore. The drill never restores over the live database. It measures recovery time and the recovery point, then drops the scratch copy. Exercises your team runs, from tabletop to full cutover, can be recorded with their targets; rows nobody measured read not measured.
An illustrative chaos experiment at the gateway: an outage injected at one provider, the request served by failover from another, the other agents still served within policy, and the experiment rolled back automatically.
1note
- Outage injected, then rolled back
How it connectsx, tested
Where a failed drill goes next.
A resilience result is a record other parts of ColossalX read.
A failed experiment or restore drill opens owned work with a due date; a later clean run closes it.
Only measured experiments and drills count toward the resilience pillar; seeded rows count for nothing.
Recent recovery tests count as evidence for the controls they exercise.
Failover routes between your providers by health scored from real traffic.
Honest by design
What it does, and what it does not.
Recovery tests · labelled: Restore drill Measured, scratch database; Team exercise Recorded by a person; Seeded row Not measured; Recovery time Took, against target; Recovery point Lost, against target. Counted only if measured.
What it does not do
x, not measured
Faults are injected into AI traffic only, not into servers, networks or cloud regions.
All 4 limits
- A restore drill proves a backup restores; it is not a live failover of the platform.
- Resilience drill scenarios in the attack catalogue are recorded as attestation, never executed.
- Only one chaos experiment runs at a time, and the platform asks you to confirm it first.
How we know
- Faults are injected into real AI traffic, and each covered request is recorded.
- An experiment without a measurement reads not measured, with the reason.
- Restore drills restore into a scratch database, never over the live one.
- Failover never retries a request a policy refused on another provider.
Questions
Questions buyers ask
What is chaos engineering for AI systems?
Chaos engineering is causing a failure on purpose, in a controlled way, to learn whether a system survives it. For AI, the failures that matter are model providers going down, erroring, throttling or answering slowly. ColossalX injects those faults into real AI traffic through its gateway, measures the result and rolls back automatically.
What happens when a model provider goes down?
Requests fail over to another provider you have connected, chosen by health scored from real traffic, and a policy decision is never retried on another provider. A chaos experiment lets you prove this before it happens for real: it takes one provider down and measures what failover served, request by request.
How are backups proven, not assumed?
Each week the newest backup is restored into a scratch database, never over the live one, checked, timed and dropped, and the recovery time and recovery point are recorded as a test. You can also record the exercises your team runs; rows nobody measured are labelled not measured and count for nothing.
What happens when an experiment or drill fails?
It opens owned work with a due date in the same queue as other findings, and a later clean run closes it. Only measured results count toward the resilience pillar of the ColossalX Trust Engine and toward control evidence, so a failure shows in the trust score and the compliance position.
Is there a business continuity and disaster recovery report?
Yes. The reports hub includes a business continuity and DR report among its 11 report types, produced as a branded PDF, on a schedule if you want, with second-person sign-off. It marks itself stale when the underlying results change, so a signed copy does not quietly stop matching the position it states.
Related
Where to look next.
-
AI gateway
One governed path to 31 provider families
-
Red-teaming and validation
Authorised attacks, sealed runs
-
ColossalX Trust Engine
A trust score that explains itself
Next step
Know your x, tested.
See a provider outage injected into your own AI traffic, and the failover measured request by request.
- 01Tell us what you run
- 02See the four verbs on it
- 03Decide where to start