Two roads
Trust the weights, or check the behavior
The mainline of AI safety tries to make a model good on the inside: train its values, study its thoughts, have it debate itself. That work matters. CIRIS bets on the other road. Assume a capable model might be misaligned, and instead of trusting its mind, make its consequential actions accountable to people and other systems that can check them.
In the field's own terms, CIRIS sits in the institutional and control branch, alongside AI control and guaranteed-safe AI, not the value-internalization mainline of RLHF, Constitutional AI, debate, and interpretability. Its answer to scalable oversight, how you supervise something smarter than you, is to verify the accountability envelope, not the reasoning. A signature, a quorum, a hash-chained audit stay cheap to check even when the decision behind them is superhuman. It aligns systems of many capable agents over time, not the values of any single mind.
The line we hold
It does not try to align one all-powerful AI. On purpose.
Accountability needs more than one party. Someone to answer to. A way to check that cannot be quietly swallowed. A balance of power no one side can capture. A single super-intelligence has none of these, so there is no honest way to hold it to account. CIRIS is built for the other future: many capable agents, people, and organizations whose consequential decisions are all independently checkable.
So the stance is explicit. A singleton ASI is not a system to be aligned but a condition to be prevented. Concentrating superhuman capability in one unaccountable place, at this stage of human institutional development, is illegitimate, because no institutions are mature enough to hold it accountable, which is precisely the danger. In the framework's own terms a singleton is the ρ→1 single-voice collapse that the corridor model names as a coordination failure, not a success. That our guarantees hold across a federation and erode against a singleton is not a gap we are patching. It is the regime we refuse to legitimize, held as a commitment, not only a prediction.
How it compares to the AI you actually use
The everyday assistants are powerful and easy to use. They also run in someone else's cloud, keep no record you can check, and answer to no one you can name. Here is the same accountability test, applied to the AI most people open every day.
| Assistant | Published principles | Proof of what it did | Asks a human when unsure | Open source | Echo-chamber check |
|---|---|---|---|---|---|
| ChatGPT | Yes | No | No | No | No |
| Gemini | Yes | No | No | No | No |
| Claude | Yes | No | No | No | No |
| CIRIS | Yes | Yes | Yes | Yes | Yes |
Compared on public product behavior as of June 2026. Each principles link goes to that company's own published spec.