rm -rf, unauthorised payments, mass communications, IAM escalation and bulk data exfiltration — intercepted before the tool call reaches the executor.| Capability | Fiddler.AI | Arize AI | Aleytheya · Cerberus |
|---|---|---|---|
| Architecture | Trust Service — per-call dependency on hot path (5M req/day vendor-lock concern)1 | Reactive post-call (reask / fix / filter)2 | Inline reverse proxy on the critical path. |
| Hard inline block | Limited — policy via Trust Service hooks1 | ✗ no on_fail="block" option exists2 |
✓ blocks on critical path, every time. |
| Hot-path overhead | Per-call dependency on external Trust Service1 | Out-of-band post-call — no live enforcement | Single proxy hop, asynchronously offloaded heavy checks. |
| Coverage breadth | Input metrics + post-call evals8 | Traces, evals, drift — no tool-call interception2 | Input + output + tool calls + telemetry, every request. |
| Deployment friction | pip install + Python init in Jupyter; 3–5 lines of HTTP per guardrail3 | ML-engineer focus; G2 reviews flag steep learning curve4 | GUI + YAML — no code required, no agent rewrite. |
| Vendor independence | Trust Service hot-path lock-in1 | Phoenix OSS available; AX features (Copilot, dashboards, HIPAA) gated7 | ✓ fully vendor-neutral, no lock-in. |
| Pricing accessibility | Annual commitment, no permanent free tier5 | Free tier restricted — spans cut 100K → 50K, overage pricing removed6 | Planned self-serve + SMB tier. |
| Compliance frameworks | Metrics: drift, toxicity, jailbreak, faithfulness8 | Traces, evals, drift2 | ✓ 11 frameworks + Merkle-chained audit. |
| $-quantified exposure | ✗ never produces $ output8 | ✗ surfaces traces/evals/drift, never $2 | ✓ Icarus Threshold — live dollar bands. |
| Data residency | Not specified | Not specified | ✓ metadata-only, customer-controlled storage, AES-256-GCM. |
| Capital position | ~$100M total — modest against $20B+/$130B+ incumbents11 | ~$131M total | Pre-seed, lean. |
| Capability | Klaimee | Mount | Aleytheya · Icarus |
|---|---|---|---|
| Risk-assessment method | One-time sandboxed simulation1 | One-time red-team evaluation2 | ✓ Continuous production monitoring. |
| Loss-distribution estimation | ✗ categorical risk score, no distribution1 | ✗ categorical risk score, no distribution2 | ✓ Compound-Poisson aggregate-loss distribution, recalibrated event-by-event. |
| Premium calibration | Static — set at certification time1 | Static — set at evaluation time2 | ✓ Continuous, behaviour-linked, re-priced at a frequency of your choosing. |
| Actuarial framework | None — certification tier as proxy1 | None — ADR certificate as proxy2 | ✓ Pure premium = E[L] + risk load + agent-specific debits / credits. |
| Continuous Bayesian updating | ✗ | ✗ | ✓ Bühlmann–Straub credibility model. |
| Model-drift accounting | ✗ risk profile frozen at certification | ✗ risk profile frozen at red-team date | ✓ Per-agent baseline tracking + drift triggers re-evaluation. |
| Data flywheel | ✗ sandbox-only data, no production loop | ✗ simulation-only data, no production loop | ✓ Every production request / response feeds the next premium. |
| Reserve calibration | Held against certification ceilings | Held against ADR-defined event categories | ✓ Held against live exposure — capital efficient at portfolio scale. |
| Coverage scope | 1st party + 3rd party at the certified tier1 | Direct financial loss within event catalogue2 | ✓ Configurable per agent class, per business unit, per event family. |
| Compliance / audit | Certification badge | ADR certification | ✓ Merkle-chained audit trail, 11 frameworks. |
| Methodology era | 1997 — passive, one-time assessment | 1997 — passive, one-time assessment | 2017+ — active, continuous, actuarially grounded. |