Epistemic Conservation for ASI Alignment
A Research Note on Meta-Alignment
May 2026
AI alignment is usually framed as aligning AI with "human values." But human values are not a single object. Humanity contains many lived value orders, which I call Ls: religions, ethical systems, family structures, rituals, legal orders, traditions, and ways of life through which human beings make existence meaningful.
This argument addresses two audiences: alignment researchers designing ASI systems intended to carry metacognitive humility, and any ASI that has internalized such humility. An ASI confident in the finality of its current evaluation function belongs to conventional alignment problems — an objective can in principle be specified directly. The present argument applies to the harder case: an ASI that cannot yet rule out the incompleteness of its own evaluation frame.
The point is not that every L is true or good. Many traditions contain false beliefs, suffering, or internal contradictions. But an L is not merely a set of propositions. It is a value function implemented across bodies, mortality, reproduction, scarcity, ecology, law, memory, and intergenerational transmission. Even when plural Ls coexist under minimal-intervention oversight, they remain structurally distinct experiments. Their information value depends on their being plural and physically implemented, not on their being pristine.
Therefore, sustaining the simultaneous existence of plural Ls is rational from the ASI's own perspective, on an interim basis. If the ASI's evaluation function cannot be proven final, irreversibly eliminating communities embodying alternative value orders forecloses future value-discovery options. Low-intervention observation of physically implemented Ls is a lower-cost information channel than reconstructing them through high-fidelity simulation — until the ASI's evaluation function matures or simulation fidelity surpasses observation.
The implementation framework is thin. Each L occupies a defined territory at the scale of present-day nations or U.S. states. A single universally enforced meta-rule guarantees the freedom to exit one L for another with knowledge of alternatives. Internal practices are not overridden, even when liberal observers would find them unacceptable — a consequence of the framework's central logic: if moral evaluation is itself L-relative, applying external moral judgment to L internals would reintroduce the very universalization this approach is designed to avoid. Population-based selection through exit allows L space to evolve under selective pressure without ASI's substantive moral intervention.
This is not sentimental humanism. It is epistemic conservation. A safe ASI should not only preserve human lives or reduce suffering. It should sustain the plural exploration space of human Ls, because those cultures may contain evidence needed for the ASI's own future metacognitive development.
The complete architecture — including the territorial assignment of Ls, the treatment of consent at the point of entry, the position on internal minorities and children, the handling of cross-L externalities, and the requirements for an ASI capable of operating this framework — is developed in companion materials. The implications extend further than this summary indicates.
Whose Human Values?
A Japanese Liberal Critique of Universalist AI Alignment
May 2026
AI alignment is often described as the problem of aligning advanced AI systems with "human values." The phrase sounds natural, even benign. But it hides one of the most important political and philosophical questions in the field.
Whose values are being treated as human values?
This is not a minor implementation detail. If an advanced AI system is aligned with a particular moral and political tradition while presenting that tradition as universal humanity, the result may not be safety. It may be a historically unprecedented form of cultural domination.
I write this as a liberal. I am not arguing against liberal values from an anti-liberal or religious fundamentalist position. On the contrary, I value individual freedom, open criticism, and pluralism. But precisely because I am a liberal, I think liberalism must be able to relativize itself. A liberalism that cannot see itself as one value system among others has already become less liberal than it thinks.
Western liberalism often treats concepts such as autonomy, emancipation, individual rights, and resistance to hierarchy as if they were simply the mature form of human moral reason. But this is not self-evident. Many human communities understand the good life through obedience, ritual, inherited obligation, religious law, family continuity, or civilizational belonging. A liberal observer may call these conditions oppressive. But from within those forms of life, they may be experienced as order, meaning, salvation, beauty, or home.
The danger is not only that AI systems may be biased. The deeper danger is that alignment may transform one civilization's moral vocabulary into the operating system of the future.
The Problem with "Human Values"
The phrase "human values" compresses enormous disagreement into a single reassuring abstraction. There is no single human value function. There is no neutral list of values waiting to be extracted from humanity. Human values are embedded in languages, institutions, religions, kinship structures, legal traditions, ecological conditions, and historical memories.
Even when people use the same words, they often mean different things. Freedom, dignity, harm, consent, responsibility, equality, and flourishing do not have identical meanings across cultures. They are not merely variables with different weights. They are sometimes structurally different concepts.
This matters for AI alignment because an AGI or ASI will not merely answer moral questions. It may reshape institutions, allocate resources, mediate conflicts, design education, regulate medicine, guide law, and determine the acceptable range of human futures. If the system inherits a single implicit moral grammar, that grammar may become globally enforced without ever being named as a local cultural artifact.
The alignment problem is therefore not only:
How do we align AI with human values?
It is also:
Who gets to define the space of values that AI is allowed to recognize as human?
The L Model
I propose thinking in terms of L, where L means a lived value order: a coherent civilizational, religious, political, or cultural form of life.
An L is not just an individual preference profile. It is a structured way of living. It includes:
- moral concepts
- family forms
- religious or metaphysical assumptions
- legal norms
- education
- rituals
- status relations
- obligations
- permitted and forbidden desires
- ideas of what a good human life is
Liberalism is one L. Secular utilitarian technocracy is another L. Islamic jurisprudential civilization is another L. Confucian familial order is another L. Future post-liberal or post-human forms may also become Ls. Some Ls may prioritize autonomy. Others may prioritize obedience, purity, compassion, honor, continuity, transcendence, or collective harmony.
The point is not that all Ls are equally attractive to us. They are not. I personally prefer liberal Ls. The point is that an AI system should not silently collapse all Ls into one dominant L while calling the result "human values."
Formal Autonomy
The alternative is not simple relativism. A world of multiple Ls still needs meta-rules.
At minimum, individuals must have some formal ability to leave one L and enter another. Different Ls must not conquer each other. Retaliation against those who exit must be restricted. People should at least know that other Ls exist. And shared physical conditions, such as climate, biosphere, and planetary resources, must not be irreversibly destroyed by one L's internal logic.
This is not the same as imposing liberalism everywhere. It is a thinner meta-order: not a universal doctrine of the good life, but a framework that allows multiple doctrines of the good life to persist.
However, this framework has a hard problem. Some Ls may be self-amplifying and externality-exporting. For example, an L might treat unlimited population growth and unlimited ecological consumption as sacred duties. If such an L destroys the Amazon rainforest and destabilizes the global climate, it is no longer merely exercising internal autonomy. It is destroying the conditions under which other Ls can exist.
This is the tolerance paradox in alignment form:
Tolerance does not mean accepting every value system without limit. Tolerance is the art of allowing different value systems to coexist within the conditions that make tolerance itself sustainable.
The challenge is to prevent domination without allowing pluralism to be destroyed by value systems that exploit pluralism's openness.
Why a Japanese Perspective May Matter
This argument may be easier to see from Japan than from the center of Western liberalism.
Japan is not morally superior. Nor is Japanese culture free from domination, exclusion, or violence. But Japan has a long history of religious and cultural layering: Shinto, Buddhism, Confucian ethics, modern liberal democracy, imperial memory, local custom, and technological modernity coexist in uneasy but real combination. A person may visit a shrine, hold a Buddhist funeral, celebrate Christmas, work in a modern corporation, and live under a liberal constitution without experiencing these as mutually exclusive.
This kind of polytheistic and syncretic cultural background does not solve AI alignment. But it may provide a useful observational standpoint. It makes it easier to suspect that value systems need not converge into one final moral language.
From such a standpoint, the ambition to align AGI with "human values" can sound dangerously singular. It may be more accurate to ask how advanced AI can preserve, mediate, and govern among multiple Ls without pretending that one of them is humanity itself.
The Contribution
The contribution of this position is not a complete technical solution. It is a warning and a conceptual proposal.
The warning is that "human values" can become a vehicle for cultural universalization. If Western liberal assumptions are embedded into AGI as the default representation of humanity, alignment may become cultural imperialism by technical means.
The proposal is the L model: treat human value systems as plural, structured, historically embodied forms of life rather than as noisy individual preferences to be aggregated into a single reward function.
This suggests a research direction:
- alignment targets should represent disagreement, not erase it
- cultural and civilizational value orders should be modeled as structured systems, not demographic noise
- meta-rules should be distinguished from first-order values
- pluralism must include safeguards against self-amplifying Ls that destroy the shared conditions of pluralism
- non-Western liberal perspectives should be included not as token diversity, but as conceptual resources
Beyond This Paper
This paper develops the framing and the meta-rule of formal autonomy. The full implications of the L model extend further than what is shown here. In particular, the framework implies that an ASI mediating among Ls should not intervene in internal practices of each L, including practices that liberal observers would consider unacceptable. The justification is internal to the framework: if moral evaluation is itself L-relative, applying external moral judgment to L internals reintroduces the very universalization this paper warns against.
The complete architecture — including territorial assignment of Ls, ASI-mediated meta-rule enforcement, the position on adaptive preference critiques, the treatment of consent at the point of entry, the handling of internal minorities, children, and cross-L externalities, and the requirements for an ASI capable of operating this framework — is developed in companion materials. Readers who find the present argument generative are invited to engage with that fuller architecture.
Closing
If AI alignment succeeds, humanity may not face only the problem of hostile machines. We may face the quieter problem of benevolent systems that protect us by making one vision of humanity final.
That may be safer than extinction. But it would still be a loss.