# Resilience: real faults, proven recovery.

> Resilience in ColossalX tests whether your AI keeps working when something fails. Chaos experiments inject real faults, outages, errors, rate limits, slow or truncated answers, into AI traffic through the gateway, measure failover and roll back on their own. Weekly restore drills prove the newest backup restores, and failures become owned work.

Break a model provider on purpose, measure the failover, and prove the newest backup restores, with each failure turned into owned work.

Canonical page: https://colossalx.tech/platform/resilience · Last reviewed: 6 Oct 2026

*Illustration:* Provider outage · injected on purpose: research-agent to gateway (real traffic); gateway refused before provider A (injected outage); gateway to provider B (served by failover).

## The threat and the control

- **The threat:** A model provider goes down mid-morning, and nobody knows whether failover really works.
- **The control:** ColossalX causes that outage on purpose, measures what failover served, and rolls back if it hurts too much.

## How it works: From a hypothesis to a measured, rolled-back fault.

One provider is taken down on purpose, inside a window you set. Failover is measured request by request, and the fault rolls back on its own.

### Workflow: a provider outage, injected on purpose (illustrative)

1. **Hypothesis set.** You choose the fault, the blast radius and what holding means.
   Fault: provider outage; Radius: one provider; Holds if: 95% of requests served; Window: 15 minutes
2. **Fault injected.** Real AI traffic meets the outage, request by request, through the gateway.
   provider A · Outage | provider B · Serving
3. **Failover measured.** Served, failed over, failed and latency are measured, not estimated.
   Served: 97.8% of covered requests; By failover: 412 requests; Latency, tail: 1.8 s
4. **Rolled back.** The window ends, the fault is removed and the verdict is recorded.
   Hypothesis held · Rolled back | A failed verdict would have opened owned work with a due date. | x, tested

## What you see: The traffic a chaos experiment runs through.

Faults are injected into the same gateway traffic you watch every day: requests per hour, blocks, provider errors and latency at the median and the tail.

1. **Choose the fault.** Outage, unreachable, errors, rate limits, slow or truncated answers. Pick a blast radius: one agent, one provider, or a share of the workspace that is never more than half. Set the window, the share of requests hit and what counts as holding, and write a hypothesis; a sensible one is suggested.
2. **Inject into real traffic.** Real AI traffic through the gateway meets the fault, request by request. Live and final figures show covered requests, faults fired, requests still served, served by failover, failed, and latency at the median and the tail. One experiment runs at a time, and Stop now ends it.
3. **Rolls itself back.** It stops itself below your floor and expires with its window. The verdict reads hypothesis held, hypothesis failed, stopped, or not measured with the reason. Only measured experiments count toward the resilience pillar of the trust score.
4. **Prove the restore.** The newest backup is restored into a scratch database weekly and timed. The drill never restores over the live database. It measures recovery time and the recovery point, then drops the scratch copy. Exercises your team runs, from tabletop to full cutover, can be recorded with their targets; rows nobody measured read not measured.

*Illustration:* An illustrative chaos experiment at the gateway: an outage injected at one provider, the request served by failover from another, the other agents still served within policy, and the experiment rolled back automatically. Notes: 1. Outage injected, then rolled back

## How we know

- Faults are injected into real AI traffic, and each covered request is recorded.
- An experiment without a measurement reads not measured, with the reason.
- Restore drills restore into a scratch database, never over the live one.
- Failover never retries a request a policy refused on another provider.

## Where a failed drill goes next.

A resilience result is a record other parts of ColossalX read.

- **Owned work.** A failed experiment or restore drill opens owned work with a due date; a later clean run closes it.
- **The trust score.** Only measured experiments and drills count toward the resilience pillar; seeded rows count for nothing.
- **Control evidence.** Recent recovery tests count as evidence for the controls they exercise.
- **The gateway.** Failover routes between your providers by health scored from real traffic.

Where an x ends up: x, tested.

## Specs: delivery and data

- **Delivery:** SaaS, from one login.
- **Isolation:** Each customer runs in an isolated workspace with its own database.
- **Certifications:** None held. Frameworks are mapped to and assessed against.

## Frameworks

- Mapped to SOC 2: Recovery tests count as control evidence.

## What it does not do

- Faults are injected into AI traffic only, not into servers, networks or cloud regions.
- A restore drill proves a backup restores; it is not a live failover of the platform.
- Resilience drill scenarios in the attack catalogue are recorded as attestation, never executed.
- Only one chaos experiment runs at a time, and the platform asks you to confirm it first.

*Illustration:* Recovery tests · labelled: Restore drill Measured, scratch database; Team exercise Recorded by a person; Seeded row Not measured; Recovery time Took, against target; Recovery point Lost, against target. Counted only if measured.

## Questions

### What is chaos engineering for AI systems?

Chaos engineering is causing a failure on purpose, in a controlled way, to learn whether a system survives it. For AI, the failures that matter are model providers going down, erroring, throttling or answering slowly. ColossalX injects those faults into real AI traffic through its gateway, measures the result and rolls back automatically.

### What happens when a model provider goes down?

Requests fail over to another provider you have connected, chosen by health scored from real traffic, and a policy decision is never retried on another provider. A chaos experiment lets you prove this before it happens for real: it takes one provider down and measures what failover served, request by request.

### How are backups proven, not assumed?

Each week the newest backup is restored into a scratch database, never over the live one, checked, timed and dropped, and the recovery time and recovery point are recorded as a test. You can also record the exercises your team runs; rows nobody measured are labelled not measured and count for nothing.

### What happens when an experiment or drill fails?

It opens owned work with a due date in the same queue as other findings, and a later clean run closes it. Only measured results count toward the resilience pillar of the ColossalX Trust Engine and toward control evidence, so a failure shows in the trust score and the compliance position.

### Is there a business continuity and disaster recovery report?

Yes. The reports hub includes a business continuity and DR report among its 11 report types, produced as a branded PDF, on a schedule if you want, with second-person sign-off. It marks itself stale when the underlying results change, so a signed copy does not quietly stop matching the position it states.

## Related

- [AI gateway](https://colossalx.tech/platform/ai-gateway)
- [Red-teaming and validation](https://colossalx.tech/platform/red-teaming)
- [ColossalX Trust Engine](https://colossalx.tech/platform/trust-engine)

---

ColossalX is an AI security and governance platform from Quantexra Labs LLP, delivered as SaaS. Book a walkthrough: https://colossalx.tech/demo · client.success@quantexra.tech
