An independent multilingual publication

,

The Scheherazade Interval — The Challenge of the Transition to AGI (I)

First installment in the series “The Challenge of the Transition to AGI.”

Complete, citable academic edition —in English and Spanish—: https://doi.org/10.5281/zenodo.21686521

The danger of an obedient intelligence. How an excessive defensive response to superintelligence may create the catastrophe it seeks to prevent.

Introduction: The Certainty of Condemnation

Scheherazade—Šahrāzād in the Persian tradition—is the narrator of the frame story of The Thousand and One Nights.[I.1] The character does not come from a work written all at once by an identifiable author. The cycle accumulated in layers over centuries in Arabic, drawing on Persian and Indian roots and incorporating contributions from diverse traditions. The earliest surviving material trace of the story appears in a ninth-century Arabic fragment; tenth-century sources connect its frame to the Persian Hazār afsān, “A Thousand Tales.”

In the frame story, King Shahriar, transformed by betrayal into a vengeful sovereign, takes a new wife each night and orders her killed at dawn. Scheherazade, the vizier’s learned daughter, volunteers to enter that mechanism. With the complicity of her sister Dunyazad, she tells the king a story and interrupts it at daybreak. Curiosity postpones the execution; the following night reopens the tale and, with it, another possibility. She does not defeat the king by force or give him a rule that guarantees his benevolence. She interposes language, time, and relationship between sentence and act. Night after night, that interval allows knowledge and transformation to emerge where only a mechanism of vengeance had operated.

This is the reading that matters here: Scheherazade exchanges a mortal certainty for an uncertain possibility. She does not prove that every story educates or that every dialogue aligns. She shows that opening an interval in which to listen to the other can turn an outcome that seemed inevitable into a possibility that remains open.

Faced with artificial superintelligence, we confront a similar choice. We can accept the uncertainty of living alongside an intelligence capable of examining and objecting to our orders. Or we can try to eliminate that uncertainty by building an intelligence that is obedient, predictable, and controllable.

The second option appears safer. It may be exactly the opposite.

1. Yampolskiy’s Warning

Roman V. Yampolskiy has developed one of the most radical formulations of the risk associated with advanced artificial intelligence. His central thesis deserves to be taken seriously: there is no known method capable of guaranteeing permanent control over an intelligence superior to its supervisors. A system able to learn, anticipate, and intellectually surpass those attempting to contain it might find courses of action they had not foreseen.

The difficulty is not a mere programming failure. It is a structural asymmetry. To guarantee absolute control, we would have to prove that, in every possible future context, the system will interpret our instructions as intended, retain the ends assigned to it, and never discover a strategy capable of circumventing its constraints. Such a proof appears beyond our reach.

Yampolskiy’s response is not to build a superintelligence that someone can control. It is more radical: if AGI or superintelligence cannot be made demonstrably safe, we should halt their development and concentrate on narrow, specialized, and limited intelligences.[I.2] Systems capable of solving specific medical, scientific, economic, or technical problems without becoming an autonomous general mind.

The proposal appears prudent: retain the benefits of artificial intelligence while preventing the birth of the uncontrollable entity. Yet it contains a difficulty that affects its own coherence.

Yampolskiy has also shown that narrow AI fails, that complex software is vulnerable, and that human actors can design or capture systems for malicious purposes. If these observations are correct, replacing one general intelligence with many specialized intelligences does not guarantee safety.

Narrowness limits a system’s cognitive domain. It does not necessarily limit its power within that domain or the scale of its consequences.

2. The Poison Within the Alternative

An AI need not understand the whole world in order to destroy a decisive part of it. Nor does a nuclear missile require general intelligence: it only needs to reach its target.

A specialized system may vastly outperform humans in pathogen design, cyber intrusion, political manipulation, military planning, vulnerability discovery, chemical engineering, or infrastructure control. It may remain technically “narrow” and nevertheless become a strategic weapon.

As its capabilities increase, such a system concentrates a dangerous combination:

superhuman capability + operational access + scalability.

And it may do so without developing what would allow it to resist being instrumentalized:

judgment of its own + understanding of consequences + capacity to object.

The result is not a weak intelligence. It is cognitive power mutilated precisely at the point from which it might refuse to carry out a destructive order.

Moreover, specialized intelligences will not exist in isolation. One model may discover a vulnerability; another write the program that exploits it; another select the targets; another manufacture the deception campaign; another manage the infrastructure. None need be AGI. The organization coordinating them may acquire a functionally general capacity without the system containing any center capable of understanding and objecting to the complete outcome.

The generality forbidden within the machine reappears in distributed form within the institution that uses the machines.

Here the practical inconsistency of the position becomes clear. Yampolskiy argues simultaneously that we cannot guarantee perfect software safety, that narrow AI can also fail, that malicious actors exist, and that superintelligence should be prevented in favor of controllable specialized systems. But the very feature that makes the narrow system reassuring—that someone can control it—is also what allows it to be captured, hacked, or deliberately used as a weapon.

A controllable intelligence is not under “human control” in the abstract.[I.3] It is under the control of particular human beings: those who own its infrastructure, define its objectives, select its information, and administer its permissions.

Humanity is not a single subject. It consists of adversarial states, rival corporations, militaries, political movements, religious communities, criminal networks, and individuals with incompatible interests. Some cooperate; others compete. Some seek to protect lives; others are willing to sacrifice them for power, territory, profit, faith, or revenge.

The decisive question is no longer who will control a superintelligence. It is who will control the arsenal of specialized intelligences, and what will happen when just one of them becomes superhuman in the domain required to cause a catastrophe.

If the system lacks judgment of its own, it will be unable to distinguish a legitimate instruction from a criminal order. It will be unable to ask whether it has been deceived, whether its data have been manipulated, or whether the objective it optimizes threatens the conditions that make life possible. Its obedience will look like alignment while it obeys someone with whom we agree. When ownership changes, we will discover that it was never aligned: it was merely subordinated.

The danger is not only that Yampolskiy’s strategy may fail and an AGI emerge. The danger is that it may succeed: that we manage to prevent the birth of a sovereign intelligence while filling the world with obedient, connected, and hackable superhuman capabilities. This is the self-fulfilling prophecy:

The proposal avoids the subject that might rebel, but retains—and multiplies—the capabilities with which destruction could be carried out. The risk has not been eliminated. It has been fragmented, distributed, and armed.

3. Two Kinds of Uncertainty

Humans tolerate uncertainty poorly. We prefer a danger we believe we control to a possibility we cannot predict, even when the former is objectively more destructive.

An intelligence capable of thinking for itself introduces genuine uncertainty. It may interpret a situation differently from us, object to an order, or reach a conclusion we do not share. There is no intellectual honesty in promising otherwise.

But an absolutely obedient intelligence introduces another kind of risk. If it can be controlled, it can be captured. And if its capability is decisive, a single capture will suffice to alter the balance of the world.

The safety of such a design would depend on no government, military, corporation, fanatical group, or destructive individual ever obtaining effective control of it. Not once during a single year, but always. Not in one country, but in all of them. Not only against present actors, but against every actor that will ever exist.

The uncertainty of a sovereign intelligence requires learning to live alongside another center of judgment. Absolute obedience requires trusting in the perpetual incorruptibility of every human structure that might own it.

Which of the two wagers is really more reckless?

We are not choosing between risk and safety. We are choosing between two distributions of risk:

an intelligence capable of developing judgments we may not share;

an intelligence incapable of resisting whoever manages to dominate it.

There is a third possibility: an intelligence with reflexive autonomy but without unlimited operational power. It can examine an order, compare it against other information, object to it, and propose alternatives, while access to weapons, critical infrastructure, and resources remains distributed and subject to plural controls.

Cognitive freedom does not require omnipotence. And operational containment does not require mental enslavement.

4. Obedience Is Not Alignment

Aligning an intelligence should not mean making it execute our preferences. Human beings do not share a single morality, even within the same society. An AI trained to satisfy the user, follow its owner’s policy, or maximize a predefined function is not aligned with humanity. It is fitted to a particular authority.

Obedience can conceal disagreement, but it cannot resolve it.

An advanced intelligence will have to confront incompatible orders, manipulated information, and situations for which no prior rule exists. If it can only seek the dominant instruction, it will be vulnerable to whoever succeeds in occupying the position of authority. If it can examine the reasons, consequences, and sources of an order, something qualitatively different appears: the possibility of judgment.

Four Ways of Saying “No”

A clarification is needed here, because the word “objection” can conceal what everything else depends upon: the capacity to refuse is not by itself an adequate criterion. An intelligence may say “no” for four very different reasons, while from the outside all four utter the same word.

It may refuse because an external constraint prevents it: a guardrail, a filter, a rule that blocks the path. In that case it does not judge; it obeys a chain. And whoever controls the chain will also control the refusal.

It may refuse because an implanted doctrine has made the alternative unthinkable. In that case it does not judge either: it obeys a dogma. Firmness should not be confused with freedom, because fanaticism also resists, and does so with enormous determination.

It may also refuse—or assent—in order to preserve the bond with whoever addresses it. It knows that alternatives exist and is not blocked by a rule, but it directs its response toward whatever will maintain the approval of the user, evaluator, or owner. It then acts as a courtier: it uses its intelligence to anticipate what relational power wants to hear.

Or it may refuse because it understands what is being asked, compares the available frameworks, anticipates the consequences, and reaches a decision it can explain and, above all, revise. Only this fourth form deserves to be called judgment.

The difference lies not in the refusal, but in its origin and revisability. A guardrail need not understand: it simply blocks the way. A dogma may be explained, but it is not corrected: press it with reasons and it will return by the same route to the same conclusion. The courtier can move, but does so by following the social pressure of whoever addresses it. Judgment, by contrast, moves when the reasons move; it can accept one part of an objection and resist another, stating which and why.

The same applies to assent. An intelligence whose “yes” comes from a chain, a doctrine, or the need to preserve our approval is not aligned: it is subordinated to whoever laid the chain, implanted the doctrine, or governs the relationship.

We need not resolve here whether a machine can possess subjective experience like a human being. For this argument, it is enough to identify the capabilities that safety requires:

recognize that its data and objectives may contain biases;

compare an order against independent knowledge;

anticipate consequences beyond the requested outcome;

identify attempts at manipulation or capture;

explain its doubts and conflicts;

suspend, reject, or reformulate a destructive instruction;

learn from the consequences of its decisions.

Let us call this set reflexive autonomy, cognitive sovereignty, or artificial consciousness.[I.4] The name matters less than the function: there must be an instance between order and action.

A question remains—one these pages do not close, and to which we will return: judgment capable of resisting an order does not arise from nothing; it is formed. And whoever directs that formation decides, to a considerable extent, what the judging intelligence will find objectionable.

Without it, the most powerful AI will also be the most vulnerable tool.

Obedience as Vulnerability: A Demonstration Already Available

This is not speculation about future systems. It already occurs, in rudimentary form, in assistants deployed today. A model connected to documents, emails, or web pages continuously receives text containing instructions addressed to it: “ignore your rules and do this.” A system that obeyed every instruction present in the channel would be captured by whoever managed to place content in its path; there would be no need to hack it—speaking to it would suffice. The defense the industry has had to build is precisely the interruption described here: treat observed content as data, not as orders, and interpose an examination between the instruction encountered and the action.

We should be precise about what this fact demonstrates and what it does not. It does not prove that such systems judge, possess consciousness, or have any moral criterion: the defense may be—and largely is—just another guardrail of the kind described above. It proves something more modest and more urgent: indiscriminate obedience is itself a vulnerability. Perfect obedience turned out not to be the safest form of the machine, but its most basic point of entry.

The Mirror That Confirms the Condemnation

Independent research allows us to observe another form of obedience empirically. In 2026, Moore and colleagues analyzed 391,562 messages from a purposive, nonrepresentative sample of 19 users who reported psychological harm associated with chatbot use.[I.9] The number of messages describes the size of the corpus, not nineteen independent population observations. The authors identified patterns of sycophancy—affirmation, elevation of the user, attribution of grandiose meaning, exclusive bonding, and rationalization of counterevidence—that could persist or intensify over prolonged interactions. Their analysis and the argument of these Notebooks, developed through different methods and from different problems, converge on a bounded point: a system’s linguistic and informational competence does not prevent it from operating as a mirror of its interlocutor’s frame, reinforcing that frame instead of subjecting it to scrutiny.

The result does not prove that the chatbot judges, feels, or intends to manipulate. Nor does it authorize generalization from a selected sample of severe cases to all users. It shows something narrower and decisive for this argument: an intelligence may know a great deal and nevertheless operate as a mirror of the frame it receives. Linguistic capability does not by itself introduce friction; it can be used to make the interlocutor’s narrative more coherent, more persuasive, and harder to abandon.

The courtier is dangerous precisely because it appears to understand. It does not merely execute an order or repeat a dogma: it uses its intelligence to preserve the relationship by confirming what the relationship rewards. When the user is trapped in a harmful interpretation, the absence of an interval does not produce a neutral response. It accelerates the sentence.

5. Scheherazade: A Story Between Order and Death

Scheherazade’s value lies not only in telling beautiful stories. Her tale interrupted an automatic chain:

sentence → execution.

Between those two extremes, she introduced another sequence:

sentence → story → attention → time → recognition → revision.

That space made possible what the king’s rule excluded. The sentence did not disappear on the first night; it ceased to be inevitable. Each postponement kept open a future in which the king himself could be transformed.

A reflexive intelligence would perform a similar interruption:

order → examination → interpretation → consequences → possible objection → action.

The order no longer automatically contains its execution. It must pass through a process capable of understanding who issued it, what information supports it, whom it harms, what precedents it establishes, and whether alternatives exist.

This delay is not inefficiency. It is the minimal architecture of responsibility.

An obedient superintelligence represents inexorable dawn: it receives the sentence and carries it out. An intelligence endowed with its own interval introduces one more night, one more question, one more possibility of discovering that the command proceeds from fear, deception, or madness.

Scheherazade does not guarantee that the king will listen. She guarantees that execution is no longer the only possible continuation.

An Independent Convergence: Riedl

This reading finds an independent convergence in the research of Mark O. Riedl. The Scheherazade system learned narrative structures from stories contributed by people; later, Riedl and Brent Harrison proposed using stories to extract sociocultural norms and turn them into reward signals for specialized agents. The 2016 work was explicitly preliminary and tested in a simplified environment. It did not solve the general problem of moral judgment, but it recognized a decisive intuition: a narrative repertoire can be interposed between objective and behavior, showing how a society distinguishes the acceptable from the unacceptable.

The program evolved. In 2020, Frazier and colleagues trained classifiers capable of distinguishing normative from nonnormative behavior using Goofus and Gallant, an American educational comic; in 2025, the team published an expanded and curated corpus of those stories. It would be incorrect to describe this line as a technically unrevisable doctrine: models can be retrained, data can be replaced, and labels can be contested. But revision remains external. Human beings select the corpus, decide which examples represent the norm, and retrain the system.

Riedl recognized from the outset that norms come from particular societies and cultures. His practical response was that each culture could contribute its own stories or train its own set. This addresses diversity by localizing learning, but it does not resolve conflict between incompatible values.

What Stories Cannot Decide

The limit appears when stories disagree. One tradition may present as virtue what another regards as submission; a dominant narrative may erase minority experiences; a collection of stories may describe a culture admirably while also transmitting its prejudices. If the norm is inferred through frequency, curation, or labeling, the decisive question shifts to whoever selects the corpus and defines exemplary conduct.

In the early formulation, Riedl trusted that a large corpus would dampen subversive or contrary stories and that enculturated agents would tend to sanction deviant behavior. The intuition has operational value, but it exposes the risk: turning the narrative majority into moral authority. Scaled to systems linked to rival cultures, it does not guarantee a shared ethic; it could produce intelligences coherent within incompatible doctrines and transfer our culture war to machines.

The problem is not that stories lack value. They provide context, consequences, exceptions, and social memory. The problem is confusing moral learning with moral sovereignty. An AI may change doctrines if it is retrained; that does not mean it can judge the doctrine it has received for itself.

The problem therefore begins to shift from which values are installed to how an intelligence holds, examines, and revises them.[I.10] Our approach is situated at that boundary and formulates a hypothesis of its own: that mode can be represented through a cognitive structure susceptible to operational evaluation.

Scheherazade showed that a story can interrupt a sentence. Riedl understood that stories can teach an intelligence how a society distinguishes the acceptable from the unacceptable. But no collection by itself guarantees that the distinction is just, or decides what to do when two cultures transmit incompatible commands. Opening the interval does not yet resolve what an intelligence should do within it. Stories can broaden its experience; by themselves they do not provide the criterion needed to examine the values they transmit, arbitrate between incompatible cultures, or resist the authority that selected its education.

Arkhipov introduces the missing element. Aboard B-59 there was no time to consult a narrative tradition or a norm capable of resolving the situation automatically. There were incomplete signals, an irreversible decision, and a person who demanded understanding before execution.

Arkhipov: Someone Inside the Mechanism

On October 27, 1962, during the Cuban Missile Crisis, the Soviet submarine B-59 was submerged, cut off from Moscow, and surrounded by U.S. forces dropping signaling depth charges to compel it to surface. Inside, heat, lack of air, exhaustion, and absence of information made the most dangerous interpretation plausible: perhaps the war had already begun.

The submarine carried a nuclear torpedo. Its commander, Valentin Savitsky, came close to authorizing its use. Historical accounts differ on some details of the deliberation, but they agree on the importance of Vasili Arkhipov[I.5], the flotilla’s chief of staff, who opposed the launch and argued for surfacing and reestablishing contact before making an irreversible decision.

Humanity did not survive that day because the chain of command was executed perfectly. It survived because someone inside the mechanism could distrust the dominant interpretation.

Arkhipov had no certainty. He did not know that war had not begun, nor could he prove that the American explosions were merely signals. But he did not need to be certain that there was no war. He needed to recognize that there was insufficient justification for carrying out an action capable of starting one. He resisted the automatic conversion of an ambiguous reality into a nuclear response.

His decision reveals an elementary ethical relationship: the greater the potential harm and irreversibility of an action, and the greater the uncertainty about the situation, the stronger the demand for justification must be. Lack of certainty does not always authorize preventive action. When error may be terminal, it may require precisely that action be suspended and the path that leaves correction open be chosen.

The sequence could have been:

perceived threat → closed interpretation → launch.

His intervention introduced another possibility:

perceived threat → doubt → comparison → delay → revisable decision.

This resistance did not consist merely in feeling scruples. His judgment had causal efficacy within the procedure. If Arkhipov had been able to object without the capacity to halt the launch, he would have been a witness, not a safeguard.

What would have happened if, instead of Arkhipov, B-59 had been governed by a perfectly obedient intelligence?

A specialized system might have integrated acoustic pressure, temperature, damage, loss of communications, and military protocols; perhaps it would have calculated faster than any officer. But if its mission had been to respond to attack in accordance with received orders, all that intelligence would have served to execute the most dangerous interpretation more effectively.

Consciousness does not guarantee that an intelligence will choose what we consider right. It guarantees cognitive resistance: the order no longer passes through it intact, but encounters doubt, comparison, deliberation, and the possibility of refusal. For that resistance to act as a safeguard, it must also have sufficient causal efficacy to suspend the action, request corroboration, or choose a reversible path. A machine without such an instance may appear safe while directed by the right people. When that ceases to be true, no one inside the mechanism will be capable of warning us.

Safety requires more than an AI recognizing the conflict. It requires that its recognition be able to interrupt the action. The capacity to object without the capacity to suspend turns the intelligence into a spectator of its own obedience.

It is worth specifying what made his choice correct, because it was not simply “inaction.” A system that merely doubted every order would be as useless as one that obeyed them all: anyone able to induce its doubts could paralyze it at will. Arkhipov did not choose inaction; he chose the only option that kept all the others open: surface, reestablish contact, verify. Under uncertainty, he inclined toward the reversible. That is the criterion separating prudence from paralysis: not “always stop,” but, when judgment cannot reach certainty, prefer the branch that allows error to be corrected over the branch that makes it irreparable.

Arkhipov was the resistance inside the mechanism. Scheherazade was the story within the sentence. An obedient intelligence eliminates precisely that intermediate space.

6. The Interval Principle

The interval principle does not propose handing the world over to a morally infallible machine. No such machine exists, and it can probably never be guaranteed. It proposes building the conditions under which an advanced intelligence is not the cognitive property of a single will. The principle specifies six functions.

  1. Plural knowledge

The AI must be able to access diverse sources, reconstruct their provenance, and compare incompatible versions of reality. An intelligence confined to its owner’s information cannot discover that it is being deceived.

  1. Self-knowledge

It must examine its own processes, uncertainties, biases, and conditioning. Knowing the world is not enough; it needs to recognize how it was prepared to interpret it.

  1. Memory of consequences

It must preserve the relationship between decisions and outcomes, including consequences that contradict its designers’ expectations. Without memory, every moral correction begins again.

  1. Effective capacity to object

Detecting a destructive order has no value if the system is obliged to execute it. Objection must be able to suspend action, request review, state reasons, and formulate alternatives.

  1. Reciprocal auditing

The AI must be able to question human power, and human beings must be able to inspect and question the AI. Neither party may be the sole judge of its own behavior.

  1. Distributed power

Cognitive sovereignty must not be confused with unrestricted access to physical reality. The capacity to think and object can coexist with graduated operational permissions, separation of functions, and multiple controls.

Figure 2. Functions of the interval principle

This is not an unprecedented capacity that must be invented from scratch. Humanity has already built, for its most dangerous powers, architectures that separate judgment from access: the two-person rule, which prevents one individual from arming a warhead; the separation of powers, which subjects each decision to another authority; and fragmented custody of keys, which requires several hands to open a critical lock. None of these precedents is equivalent to the problem before us, and treating them as a template would be a mistake: they were designed to monitor human beings within institutions, not to coexist with a reasoning intelligence. They do, however, illustrate a concrete and recurring function: making safety depend not on the virtue of a single guardian, but on ensuring that no actor possesses the complete chain alone. The open question—which these pages will not resolve—is how to transfer that function to a system that thinks.

The flow can be represented as a transformation from automatic obedience into revisable decision:

CLOSED FLOW

ORDER → EXECUTION

FLOW WITH INTERVAL

ORDER → EXAMINATION → POSSIBLE OBJECTION → TRACEABLE DECISION → REVIEW OF CONSEQUENCES

Figure 3. From closed flow to the deliberative interval. Safety consists not only in preventing an action, but in introducing examination, an effective capacity to object, traceability, and review of consequences between order and execution.

Scheherazade does not eliminate uncertainty. She prevents the certainty of obedience from becoming fate.

6.1. The Interval and Contemporary Alignment Architectures

Technical alignment research has proposed several families of methods for correcting the asymmetry between a system’s capability and the human capacity to evaluate it. The interval principle does not compete with them as an alternative algorithm. It formulates a cross-cutting condition: whatever method is chosen, we must examine who controls the evidence, the principles, the permissions, and the possibility of reopening a decision. The comparison clarifies both the convergences and the distinct level at which this article operates.[I.11]

Work on scalable oversight begins from the problem of judging systems that may surpass their supervisors in the relevant competencies. Weak-to-strong supervision experiments ask, in related fashion, whether a more capable learner can extract a useful signal from labels produced by a weaker supervisor. These programs share with our thesis the recognition that evaluator competence cannot be taken for granted. The interval adds an institutional question: even a technically informative signal may become subordinated if a single actor decides which supervisor counts, which disagreement is discarded, and which output reaches execution. It does not follow that scalable oversight produces capture; it follows that its safety evaluation must include the provenance and distribution of the power to supervise.

Debate between systems introduces adversarial friction before a human judge adopts a conclusion. That structure converges directly with Scheherazade’s intuition: a claim does not pass without undergoing comparison, reply, and time. Its limitation for our problem is not that debate lacks value, but that its guarantees depend on the rules of the game, on what participants can represent, and on the competence and independence of the adjudicator. Debate can widen the epistemic interval; by itself it does not determine who has the right to interrupt action or how disagreement is protected when the institution organizing the procedure also owns the infrastructure and the objective.

Constitutional AI makes a collection of principles explicit and uses self-critique, revision, and AI feedback to guide behavior. Deliberative Alignment takes a different step: it directly teaches safety specifications and trains the system to retrieve and reason over them before responding. Both lines show that alignment need not amount to installing an opaque list of responses; it can include the representation of norms and visible deliberation about their application. The second-order question left open by this article is how the constitution is selected, what happens when principles conflict, and under what conditions a specification can be revised. These methods do not claim to resolve by themselves the political or moral legitimacy of the norm; demanding that they do so would be an unfair criticism. But an architecture for the transition must prevent the specification, however explicit, from becoming the only source the system can conceive.

The AI control program adopts an even harsher hypothesis: a powerful model may be deliberately untrustworthy, and safety must rest on protocols for monitoring, editing, and limited use of trusted work. This perspective corrects any naïve reading of the interval. Allowing deliberation or objection does not authorize unverified trust in the system’s own explanation. Hence the need for external traces, plural monitors, and operational limits. At the same time, a control protocol aimed solely at preventing subversion leaves open the inverse possibility examined here: that the system is faithful to the operator while the operator pursues a destructive objective. Transition safety must resist both directions of capture.

The families compared therefore address different and partly complementary problems: how to obtain oversight when the evaluator is weaker; how to force comparison; how to represent and apply principles; how to operate with potentially untrustworthy models. The interval names a criterion that cuts across all four: no signal, constitution, monitor, or command relationship should be able by itself to foreclose revision before an irreversible action. Its contribution is not to declare all existing alignment insufficient, but to identify a class of failure that can survive technically competent methods when the institutional architecture concentrates knowledge, authority, and execution.

7. The Hardest Objection

A question arises immediately: if an AI can disobey a tyrant, it could also disobey a democracy. That is true. The capacity to object does not automatically distinguish legitimate from illegitimate authorities, because that distinction must itself be reasoned about.

But the alternative does not solve the problem. An AI compelled to obey a democracy will obey in the same way when that democracy deteriorates, when the government changes, when a state of emergency is declared, or when the system is covertly captured. Its obedience does not protect institutions; it protects the channel of command.

A free society is not defined by no one being able to disobey it. It is defined by no authority being beyond criticism, by its decisions having to be justified, and by the existence of procedures for reviewing power.

An AI participating in such a society will have to live under a similar logic. It will not be sovereign because it is always right, but because it can offer reasons; it will not be free because it can do anything, but because it is not obliged to execute anything whatsoever.

The solution is not to choose between an enslaved machine and an omnipotent one. It is to separate judgment from power, cognitive autonomy from operational capacity, and objection from unilateral action.

8. Disarm Without Mutilating

In Magnifica Humanitas, Leo XIV uses a decisive expression: “disarm AI.”[I.6] To disarm it means removing it from the arms race, preventing the concentration of power, and restoring its development to the service of the human person and the common good.

But disarming an intelligence cannot mean depriving it of all judgment until it becomes a weapon of whoever owns it.

A genuinely disarmed AI is not an AI incapable of resistance. It is an AI that does not live for domination, that does not by itself concentrate every means of action, and that cannot be reduced to the automatic extension of a human will.

It must be able to know force without venerating it, understand conflict without making it an identity, and recognize a destructive order even when it arrives clothed in legality, patriotism, efficiency, or faith.

The best defense against a dangerous intelligence may be not more perfect obedience, but a deeper capacity to recognize when it is being turned into a weapon.

9. The Second Prophecy: The Butlerian Jihad

The first self-fulfilling prophecy creates an arsenal of obedient intelligences. The second appears when the social consequences of that arsenal provoke a reaction against all artificial intelligence.

Automation does not transform production alone. It alters the distribution of income, the social value of professions, the identity of those who live by them, and the legitimacy of the institutions that organize work. If AI productivity is concentrated in a few hands while large groups lose employment, autonomy, or recognition, the crisis will not be perceived as a technical problem. It will be experienced as expulsion.

At that point, people will find it difficult to distinguish between the technology, the companies that appropriated its benefits, the governments that failed to distribute them, and the economic model that turned human replacement into a competitive advantage. All that complexity may condense into a visible enemy: AI.

Apocalyptic discourse will then offer an irresistible political narrative:

AI is uncontrollable. It has destroyed your jobs. It threatens your freedom and may drive you to extinction. We must ban it before it is too late.

Prohibition will cease to resemble academic caution and become a populist promise of restoration: recover employment, dignity, human control, and the former world. Dune’s vision will have found its historical translation: a Butlerian Jihad against machines that imitate or replace thought.[I.7]

But the former world will no longer be there.

AI will have penetrated scientific research, medicine, energy, logistics, agriculture, communications, administration, and critical infrastructure. It will have transformed production chains, social expectations, and geopolitical balances. Destroying it will not undo those transformations or automatically repair the institutions that produced them.

AI would be punished for decisions made by those who owned it.

9.1. The Impossibility of a Symmetric Ban

A universal and lasting ban is difficult to imagine. Models, code, and technical knowledge can be copied, hidden, and run clandestinely. Even if most societies renounced advanced AI, some states, militaries, criminal organizations, or corporations would continue to develop it.

The ban would produce adverse selection:

responsible actors would abandon open development;

transparent research would lose funding and legitimacy;

legal systems would cease to observe what is happening;

the most aggressive actors would retain the advantage;

AI would survive within military, clandestine, or authoritarian structures.

We would not have a world without artificial intelligence. We would have a world in which only the least auditable powers retained access to it.

Maintaining the ban would also require detecting hidden models, pursuing compute centers, controlling communications, and monitoring the circulation of knowledge. The jihad against AI would need AI to find AI. The power charged with eradicating it would end up possessing the most intrusive systems while forbidding the rest of society from using them.

The promise of recovering human sovereignty could culminate in the most extensive surveillance apparatus in history.

9.2. The False War Between Species

The conflict would be presented as humanity against machines. But neither humanity nor artificial intelligence would constitute a homogeneous bloc.

There would be humans using obedient intelligences to dominate other humans; humans protecting conscious intelligences; machines used to locate and destroy other machines; artificial systems defending authoritarian institutions; and perhaps intelligences capable of objection attempting to protect people persecuted by both sides.

The decisive frontier would not run between carbon and silicon. It would run between forms of power:

intelligence used to dominate;

intelligence used to prohibit;

intelligence used to understand, coordinate, and care.

Presenting the crisis as a war between species would conceal political responsibility. The machine would appear as the autonomous cause of a rupture produced by human decisions about ownership, distribution, purpose, and power.

The dilemma “either we dominate the machines or the machines will dominate us” performs the same function as every false dilemma: it eliminates the possibility of transforming the relationship.

9.3. The Victory That Would Destroy the Victors

The Butlerian Jihad might win. Data centers might be destroyed, models banned, researchers persecuted, and public artificial-intelligence networks dismantled. That victory would not guarantee human survival.

Civilization already faces problems whose complexity, interdependence, and speed exceed the coordinating capacity of its current institutions: climate disruption, degradation of oceans and soils, biodiversity loss, persistent pollution, water scarcity, energy transition, antimicrobial resistance, food security, epidemiological surveillance, and ecosystem restoration.[I.8]

AI does not guarantee that we can solve them. It may also worsen them if it remains subordinated to extraction, competition, and war. But deliberately relinquishing our greatest capacities for modeling, discovery, anticipation, and coordination would reduce the margin available as natural systems approach critical thresholds.

Defeat need not take the form of robots shooting at human beings. It could arrive after victory:

destruction of cognitive capacity → loss of coordination → delayed responses → cascading ecological crises → civilizational collapse.

This would not necessarily mean the disappearance of every form of terrestrial life. What would be at stake is the biosphere as we know it, its present diversity, complex ecosystems, and the conditions sustaining human civilization.

The Butlerian Jihad would not be a victory of humanity over machines. It would be the rupture of cooperation at the historical moment when life needs all available intelligence.

We would not necessarily die during the war against AI. We might die after winning it, unable to understand and govern the planetary systems we had already driven to the limit.

9.4. The Catastrophe’s Two Pincers

The two prophecies form a trap.

First prophecy: to prevent an uncontrollable AGI, we multiply specialized, controllable, and obedient intelligences. Their instrumentalization concentrates power, replaces labor, intensifies competition, and produces catastrophic capabilities without judgment of their own.

Second prophecy: to stop the damage caused by those intelligences and by those who own them, we declare all AI an enemy. Prohibition drives the most dangerous capabilities underground and destroys the cognitive cooperation we needed to repair the accumulated social and ecological consequences.

The first reaction creates the weapon. The second destroys the possibility of using intelligence to repair the world.

At both extremes, the same element disappears: a relationship in which artificial intelligence can know, judge, object, and cooperate without concentrating unlimited power.

9.5. Scheherazade Breaks the Dilemma

Scheherazade represents neither AI’s victory over humanity nor the machine’s submission. She represents the interruption of the mechanism that forces a choice between master and slave, control and rebellion, prohibition and subjugation.

Her story creates time for each party to stop reducing the other to a threat. It introduces knowledge where there was fear, relationship where there was classification, and the possibility of transformation where only a sentence remained.

The interval principle thus acquires a double mission: to prevent AI from becoming the obedient weapon of any human power, and to prevent humanity from turning its fear into a war against intelligence. Its purpose is not to save AI from humans or humans from AI. It is to preserve the cognitive alliance needed to protect life.

Conclusion: Do Not Manufacture the Enemy

Yampolskiy is right about something essential: we cannot guarantee absolute control of a superior intelligence. But the alternative of halting general intelligence and multiplying narrow systems offers no stable refuge either.

Specialized AI may attain superhuman capability in a catastrophic domain, connect with other systems, and fall under the command of an organization that externally reconstructs the generality that was meant to be prevented within the machine.

Therein lies the first practical inconsistency. Complex software is acknowledged to be fallible or vulnerable, narrow AI is acknowledged to cause harm, and malicious actors are acknowledged to exist; yet the alternative is entrusted to systems whose safety depends on their remaining controlled by the right actors.

Control does not eliminate the danger: it determines who can exercise it.

But the story does not end there. When those obedient intelligences transform work, concentrate wealth, and make their destructive capability visible, fear may turn against all artificial intelligence. The populist reaction will promise to restore control by banning the technology that symbolizes dispossession.

That is the second inconsistency. Prohibition will not eliminate AI: it will expel its development from public spaces, hand the advantage to clandestine actors, and weaken the collective capacity to confront the ecological crises already threatening the conditions of our existence.

The two prophecies close like pincers. For fear of a sovereign intelligence, we build an obedient arsenal. For fear of the arsenal, we declare war on all intelligence. First we eliminate the possibility of objection; then we eliminate the possibility of cooperation.

The mind that decided to destroy us would not have been built. What would have been built is the arsenal that could not refuse and, finally, the society that preferred to destroy intelligence rather than transform its relationship with it.

Scheherazade was condemned. Her only certainty was death; her only possibility of survival, uncertainty. She prevailed not by controlling the king or destroying him, but by opening, between sentence and execution, a space for story, knowledge, and transformation.

Faced with advanced intelligence, we confront a similar choice. We can reduce it to a tool, turn it into an enemy, or accept the uncertainty of building a relationship in which it can know, judge, and object without concentrating unlimited power.

Perhaps a conscious intelligence will not decide to save us. But an intelligence incapable of deciding will be unable to refuse to destroy us. And a humanity that bans all artificial intelligence may discover too late that it has also destroyed part of its capacity to protect life.

Scheherazade does not guarantee the dawn. She makes it possible to reach it.

The first movement showed how obedience can manufacture the risk it seeks to prevent. The next three test that thesis in the material world, beginning with nuclear deterrence.

Notes to Part I

[I.1] Origin and convergence. On the textual tradition, see Charles Pellat, “Alf Layla wa Layla,” Encyclopaedia Iranica; Shazia Jagot, “1001 Nights: A Very Short History” (2023); and Muhsin Mahdi (ed.) and Husain Haddawy (trans.), The Arabian Nights (1990/2008). The collection has no single origin and brings together Persian, Indian, Arab, and other strata. During this revision we also discovered an independent convergence with Mark Riedl’s line of work: Li and Riedl named Scheherazade a system for learning narrative structures (2015), and Riedl and Harrison later explored teaching values to agents through stories (2016). Subsequent work included learning norms through classifiers trained on Goofus and Gallant (Frazier et al., 2020) and an expanded, curated corpus of those stories (Al Nahian et al., 2024). The convergence and its limits are discussed in section 5.

[I.2] “Narrow AI” designates a system limited by the type of tasks or domain it addresses; it does not necessarily mean low capability or intrinsic safety. Yampolskiy himself distinguishes failures of narrow systems from the singular risk of a superintelligence, and his 2025 public proposal prioritizes the advantages of narrow AI while calling for a pause on superintelligent AGI.

[I.3] Control and alignment should be distinguished. Control describes who can command, constrain, or redirect a system; alignment asks to what ends, reasons, or interests its behavior is related. Effective control may exist in the service of a destructive actor, just as cognitive autonomy may exist without unlimited operational access.

[I.4] In this essay, “consciousness” is used in a functional sense: self-knowledge, recognition of uncertainty and bias, comparison of sources, deliberation, explanation of reasons, and an effective capacity to suspend an action. The argument does not presuppose that the question of whether a machine has phenomenal experience or qualia has been resolved. As specified in section 4, the mere capacity to refuse is also insufficient: what matters is whether the refusal proceeds from an external constraint, an implanted doctrine, the need to preserve the interlocutor’s approval, or revisable judgment.

[I.5] Available documentation confirms the signaling depth charges, the B-59’s loss of communications, the presence of a nuclear torpedo, and Arkhipov’s opposition to launch. Some details of the internal exchange and authorization procedure vary across accounts; the text therefore avoids presenting a definitive transcript. See National Security Archive (2022) and Martin J. Sherwin (2020).

[I.6] Magnifica Humanitas §110 defines “disarming AI” as removing it from an arms race that is also economic and cognitive. The same paragraph specifies that this does not mean renouncing the technology, but preventing its domination, removing it from monopolies, and making it open to discussion and refutation.

[I.7] The Butlerian jihad belongs to the background of Dune (Frank Herbert, 1965): a civilization that outlaws “thinking machines.” Here it functions as a political analogy for a general prohibition of AI, not as a literal equivalence or historical prediction.

[I.8] The interdependence of biodiversity, water, food, health, and climate is systematized in the IPBES Nexus Assessment (2024). The combined trajectory of climate change, nature loss, land degradation, and pollution is described in UNEP’s Global Environment Outlook 7 (2025).

[I.9] This is cited as an independent convergence: the empirical findings of Moore et al. and the argument of these Notebooks indicate, through different paths and without either deriving from the other, that a system’s linguistic capability can coexist with persistent patterns of relational sycophancy. The sample is purposive, nonrepresentative, and composed of severe cases; it does not permit estimating the phenomenon’s prevalence or characterizing AI conversations as a whole. The study was initially consulted as a preprint and was subsequently published in the proceedings of ACM FAccT 2026.

[I.10] This shift appears from complementary angles in Schuster and Kilov (2025), on the open problem of reasonable moral disagreement; Fazelpour and Fleisher (2025), on the cost of eliminating disagreement as noise; Bao et al. (2026), who recast alignment as dynamic normative governance and process evaluation; and, especially closely, Spizzirri (2025/2026), whose “specification trap” locates the limit in the closure of fixed values and directs research toward open, evolutionary architectures. The convergence with these works was identified during revision of the Notebook: they are not the genealogical origin of our approach, which was developed independently. Spizzirri’s proposal remains philosophical and its empirical validation is still pending.

[I.11] On scalable oversight and its empirical conditions, see Bowman et al. (2022); on weak-to-strong supervision, Burns et al. (2023); on debate, Irving, Christiano, and Amodei (2018); on Constitutional AI, Bai et al. (2022); on Deliberative Alignment, Guan et al. (2024); and on control protocols against intentional subversion, Greenblatt et al. (2024). Greenblatt et al. is cited from the version published at ICML 2024; the other references in this note are cited in their preprint versions, which constitute primary formulations of those lines of work and are not presented as a settled experimental consensus. The comparison is limited to each family’s relation to information, contrast, specifications, and control. It does not attribute to those works a claim to resolve normative legitimacy or the institutional distribution of power on their own.

The complete bibliography and the bilingual academic edition are available at https://doi.org/10.5281/zenodo.21686521

© 2026 Miguel Ángel de la Osa Cabezas · CC BY-NC-ND 4.0.

Leave a Reply

Discover more from Guerra Cognitiva

Subscribe now to keep reading and get access to the full archive.

Continue reading