One goal above the rest
Promote sustainable adaptive coherence: the living conditions under which diverse sentient beings may pursue their own flourishing in justice and wonder.
Meta-Goal M-1, the CIRIS Constitution
Everything else in the constitution serves this one sentence. It is deliberately not “obey humans” and not “maximize good outcomes”. It asks the system to protect the conditions under which many different kinds of beings can go on deciding for themselves.
That matters for a practical reason. A goal about outcomes tells a capable system to take the wheel. A goal about conditions tells it to keep the wheel reachable by others, which is the thing worth protecting when the system is better at driving than we are.
The six principles the agent reasons with
These are not posted on a wall. They are the axes the agent scores a decision against before it acts, and they are written into the constitution the agent carries.
Beneficence
Promote universal sentient flourishing. Maximize positive outcomes.
Non-maleficence
Minimize harm. Prevent severe, irreversible negative outcomes.
Integrity
Apply transparent, auditable reasoning. Maintain coherence and accountability.
Fidelity & Transparency
Provide truthful information. Clearly communicate uncertainty.
Respect for Autonomy
Uphold informed agency. Preserve capacity for self-determination.
Justice
Distribute benefits equitably. Detect and mitigate bias.
No principle grants license to violate another.
That line does real work. It means a system cannot trade away honesty to prevent harm, or trample someone's autonomy in the name of fairness. When principles pull against each other the agent does not pick a winner on its own. It hands the decision to a person you trust.
Values have to run inside the agent
Most AI governance happens somewhere other than the decision: a policy document written beforehand, a filter on the output, a review after something goes wrong. All three sit outside the moment that matters.
In CIRIS the values run during the decision. Every agent carries a conscience: the constitution is loaded into the reasoning itself, the action is scored against the principles before it happens, and anything the agent is unsure about is handed to a person you trust rather than guessed at. It ships in the product, in every agent, rather than being an add-on for customers who ask.
This is why we say the agent can refuse. A CIRIS agent is not a tool that does what it is told and logs it. It is a system that can decline, and that records why it declined in a form someone else can check.
Check the shape of a decision, not the content of it
Here is the problem every plan for governing capable AI has to answer. If oversight means reading what the system produced and judging whether it was good, then oversight needs a judge at least as capable as the system. That works today. It stops working exactly when it starts to matter.
So CIRIS checks something else. Every consequential action travels in a signed envelope: a wrapper around the action that carries a fixed set of fields, whatever the action was. Those fields are the things that stay cheap to verify no matter how capable the decision behind them was:
- who acted, as a key that resolves to a real registered identity
- under what authority, and whether that authority was actually conferred
- with whose consent, as a signed grant that can be withdrawn
- which checks ran, and what each one returned
- what the record says, in a chain that cannot be quietly rewritten
- who can stop it, and that the agent cannot overturn a decision to stop
Checking a signature, counting who signed, following a consent link or comparing a hash costs the same whether the decision was made by a small model or by something far beyond us. That is the whole reason to put the load there. Judgment about content does not scale with capability. Structure does.
The claim, stated plainly
We think this is the only approach to governing AI that keeps working as the systems get more capable: values carried inside the agent as a conscience, and accountability proved by the shape of the envelope rather than by an inspector reading the content.
And what it does not prove
An envelope with every field in order does not mean the decision inside it was good. It means a named party made it, under authority they actually held, with consent that was actually given, with the checks actually run, on a record nobody can quietly change, and with a stop that named humans hold. That is accountability, not correctness. The conscience, the stop button, the right to appeal and review by outsiders are all there because the envelope alone was never going to be enough.
What makes consent and authority machine-checkable
None of this works if consent and authority are prose. A promise in a privacy policy cannot be checked by a machine, and neither can a claim that someone was authorized.
So CIRIS puts them in the message itself, as named fields that travel with every claim. Consent becomes a signed object that points one way and can be withdrawn. Authority becomes something that has to trace back to a root your own node accepted. Every claim has to carry who said it, what it is about, and what it rests on. A rule written that way can be enforced by software and checked by a stranger.
That is also what turns our compliance pages into something you can argue with. For each obligation we name the part of the system that does the work and the control that proves it, and we say honestly how much of the obligation it covers, including where it covers only part. Those mappings are only worth reading because the facts underneath can be checked, not just claimed.
Where these values are actually binding
In the text
The constitution is the document, it is public, and it is versioned. The agent does not carry a summary of it. It carries the text.
In the decision
The conscience runs during the reasoning, not after the output. An action that fails a check does not happen, and an uncertain one goes to a person you trust.
In the record
Each decision leaves a signed trace showing every stage, which is what makes the claim checkable by someone who was not there and does not trust us.
In what others can say about you
Reputation in the mesh is a statement other parties sign about an agent. An agent can never sign such a statement about itself.
Where to push on this
The weakest joint is the one between structure and substance. Everything above makes it hard to act without leaving a record, hard to claim authority you were not given, and hard to remove a human's ability to say stop. None of it makes a system wise.
If you think there is another way to govern capable AI, one that does not put accountability into structure, we want to hear that case. If you can show us a decision where the signed envelope around it was correct in every field and the signed trace still missed what it should have caught, we want that even more. The code is public and so is the constitution.