back to the lobby

The superalignment landscape

Most of the field is aligning the model. CIRIS is building the institutions around it.

There are two honest ways to make powerful AI safe. One is to train the model's values until you trust them. The other is to assume the model could be wrong, and make what it does accountable, in the open, to people and systems anyone can check. CIRIS takes the second road. Here is the whole landscape, mapped fairly, and where we fit in it.

01Alignment

Trust the weights, or check the behavior

The mainline of AI safety tries to make a model good on the inside: train its values, study its thoughts, have it debate itself. That work matters. CIRIS bets on the other road. Assume a capable model might be misaligned, and instead of trusting its mind, make its consequential actions accountable to people and other systems that can check them.

In the field's own terms, CIRIS sits in the institutional and control branch, alongside AI control and guaranteed-safe AI, not the value-internalization mainline of RLHF, Constitutional AI, debate, and interpretability. Its answer to scalable oversight, how you supervise something smarter than you, is to verify the accountability envelope, not the reasoning. A signature, a quorum, a hash-chained audit stay cheap to check even when the decision behind them is superhuman. It aligns systems of many capable agents over time, not the values of any single mind.

It does not try to align one all-powerful AI. On purpose.

Accountability needs more than one party. Someone to answer to. A way to check that cannot be quietly swallowed. A balance of power no one side can capture. A single super-intelligence has none of these, so there is no honest way to hold it to account. CIRIS is built for the other future: many capable agents, people, and organizations whose consequential decisions are all independently checkable.

So the stance is explicit. A singleton ASI is not a system to be aligned but a condition to be prevented. Concentrating superhuman capability in one unaccountable place, at this stage of human institutional development, is illegitimate, because no institutions are mature enough to hold it accountable, which is precisely the danger. In the framework's own terms a singleton is the ρ→1 single-voice collapse that the corridor model names as a coordination failure, not a success. That our guarantees hold across a federation and erode against a singleton is not a gap we are patching. It is the regime we refuse to legitimize, held as a commitment, not only a prediction.

02Consumer AI

How it compares to the AI you actually use

The everyday assistants are powerful and easy to use. They also run in someone else's cloud, keep no record you can check, and answer to no one you can name. Here is the same accountability test, applied to the AI most people open every day.

AssistantPublished principlesProof of what it didAsks a human when unsureOpen sourceEcho-chamber check
ChatGPTYesNoNoNoNo
GeminiYesNoNoNoNo
ClaudeYesNoNoNoNo
CIRISYesYesYesYesYes

Compared on public product behavior as of June 2026. Each principles link goes to that company's own published spec.

Try It Yourself

CIRISsafe by structure · open by principle · kind by design