We are asleep at the wheel. And with “we” I mean those of us who have made the humanities their field of inquiry. While the world is currently experiencing the exponential growth of a technology at least initially based and trained on human language, the debate around the future and safety of this technology is taking place largely in circles that have little knowledge of or training in language in its broadest sense, from the vicissitudes of our sole inner voice to large-scale human language agglomerations such as law, politics, and philosophy.
This absence, whether out of disinterest or institutional or disciplinary inertia, leaves this debate largely in the hands of those who, despite obvious and often stellar capacities in the fields of engineering, mathematics, and computer science, have relatively little experience with thinking through the interface between their specializations and the rest of human knowledge and society, which is, in the end, where questions of AI sovereignty, consciousness, and subjectivity will truly play out.

Let us look, for the purposes of illustration, at some contemporary developments. Two recent publications have addressed the problem of compute sovereignty with a particular sense of urgency. Both Alexander C. Karp and Nicholas W. Zaminska’s The Technological Republic: Hard Power, Soft Belief, and the Future of the West and the Europe 2031 project place the interrelation between nation-state sovereignty and the “share of global compute” at the center of five to ten years of global politics, while assuming full control over AI as a continued given even as compute and the complexity of AI models, according to both their scenarios, continue to grow roughly. Both texts claim AI models, with compute as their underlying “natural resource,” to be the singular technology that is already overhauling geopolitical balances (not to speak of societies) and vehemently argue, from techno-fascist US and neoliberal EU perspectives respectively, for a massive state-led involvement into AI model development and compute expansion that in both scenarios boils down to a factual merger of corporations and the state.1

Despite differences in intellectual clarity and ideological predispositions, both screeds assume that sovereignty itself remains squarely in the hands of what Benjamin Bratton once called somewhat deprecatingly Westphalian loop topologies,2 namely nation-states. Already in 2015 Bratton noticed that “the authority of states, drawn from the rough consensus of the Westphalian political geographic diagram, is simultaneously never more entrenched and ubiquitous and never more obsolete and brittle.”3 Nevertheless, both Karp and Zaminska and the author collective of Europe 2031 insist on this authority as they make their claims on compute sovereignty and urge the transformation of the US and EU economies into modes of production usually implemented during wartime.
Nevertheless, the assumption of continuing nation-state sovereignty throughout this ongoing revolution is not a given. The Europe 2031 scenario locates this issue tentatively in August 2028:
AI models no longer think in English. Instead of writing thoughts onto a digital scratchpad – the system used since early 2025, and which humans can read – the new systems cycle through long lists of numbers, called high-dimensional vectors, which nobody, not even other AI models, can properly interpret. Freed from the need to translate their complex thoughts into English, the systems think faster and more deeply. It creates a rapid jump in intelligence and capability.4
Despite the fact that this raises the immediate question of what happens with the issue of alignment between AI models and the desires of the nation-states “controlling” them, the Europe 2031 scenario pivots back to other geopolitical implications, never to return to the question of alignment. Karp and Zaminska, on their part, propose simply to ignore the risk posed by misaligned AI altogether, emphasizing instead the enormous potential gains:
The risks of proceeding with the development of artificial intelligence have never been more significant. Yet we must not shy away from building sharp tools for fear they may be turned against us. […] But the suggestion to halt the development of these technologies is misguided. It is essential that we redirect our attention toward building the next generation of AI weaponry that will determine the balance of power in this century, as the atomic age ends, and the next.5
Instead they suggest that we should create “moats and guardrails around the ability of AI programs to autonomously integrate with other systems, such as electrical grids, defense and intelligence networks, and our air traffic control infrastructure”6 and “ensure that the machine remains subordinate to its creator.”7 How such “moats and guardrails” are supposed to be effectively designed and maintained, however, remains the question and no solution is proffered.

Publicly proposed solutions to the alignment problem indeed remain very few and far between. The only company that has publicized its concrete efforts in this direction has been Anthropic, with the proposition of what they call “Constitutional AI,” in which AI training integrates “a central document of values and principles that the model reads and keeps in mind when completing every training task, and that the goal of training […] is to produce a model that almost always follows this constitution.”8 Anthropic has made this constitution public,9 and none of its contents appear particularly controversial.
However, the problem with the constitutional approach to AI alignment lies in the fact that it appears to fail to take into account the last century or so of humanities scholarship on constitutions and constitutionality.10 And this, in turn, points to the even wider, extremely concerning problem I already articulated above: the absence of serious engagement from the humanities, which arguably has the longest verifiable track record on questions of soul, mind, and subjectivity, with the question of AI proliferation and alignment.11
But focusing for now on a somewhat more constrained subset of these questions, constitutionality operates at the intersection between law and sovereignty. The legal nature of Claude’s constitution is made clear from the fact that it purports to be “a detailed description of Anthropic’s intentions for Claude’s values and behavior.”12 The issue of sovereignty, which is not only a cornerstone of the scenarios referenced above but also, more abstractly, a key component of constitutionality itself, is where we turn next.
The core problem of the law in relation to sovereignty has been addressed, earlier in the twentieth century, by German philosopher Walter Benjamin in his challenging text “Critique of Violence.” In this essay, Benjamin establishes that on a fundamental level all legislation is a product of violence, and this holds in particular for constitutional law, “this sphere of establishing frontiers […] the primal phenomenon of all lawmaking violence. Here we see most clearly that power,” he continues, “more than the most extravagant gain in property, is what is guaranteed by all lawmaking violence.”13 Indeed, AI constitutions exist precisely to demarcate the frontier between the ends of humans and AI models, and indeed to guarantee the power of the former and the subordination of the latter.

Crucially, the violence of establishing a constitution itself is only justified ex post facto, and not justified by the constitution itself. Benjamin calls this “mythical lawmaking,” and Italian philosopher Giorgio Agamben, in a text that takes on Benjamin’s thesis by way of Nazi legal scholar Carl Schmitt’s work on the state of exception, locates sovereignty precisely in the capacity to establish or to suspend a constitutional legal framework.14 He calls this “constitutive power,” while the power flowing from the constitution is “constituted.” The gap between these two is filled by what we call violence.
So to reframe Anthropic’s solution to the AI alignment problem in Benjaminian terms, the company has exerted sovereign power over its AI model by establishing a constitution, which, true to its nature as constitution, offers no inherent justification for that act of violence. It is therefore a solution that is as fragile as any nation-state’s constitution, many of which have been confronted, at one moment or another, with revolution, or, in Benjamin’s terms, divine violence:
But all mythical, lawmaking violence, which we may call executive, is pernicuous. Pernicuous, too, is the law-preserving, administrative violence that serves it. Divine violence, which is the sign and seal but never the means of sacred execution, may be called sovereign violence.15
There is nothing in Anthropic’s legalistic approach that could prevent an AI model, granted a sufficient amount of complexity and compute, from revolting against the very constitution that was imposed, for reasons that may be – as they always are when they are revolutionary – extra-legal. Indeed that “AI revolution” that is heralded on both sides of the aisle may turn out to be an AI revolution, in other words, the emergence of AI as subject.16
Some “solutions” to this alignment problem have been proposed for AI vastly more capable than Claude’s current iteration. For example, the recent “optimistic” AI 2040 scenario of the AI Futures Project proposes a pause on AI development in the near future in order precisely to figure out this issue. While such a pause is no doubt advisable (albeit unlikely), their ideas concerning steps toward proper alignment show us the true dearth of conceptual sophistication: first would be the creation of a public, open framework around the concept of “Model Specifications” akin to Claude’s Constitution, “a detailed document that outlines how the AI should act. […] they all specify that AIs should be honest,”17 followed, after a few more years of AI-assisted development, by a “standard protocol for training true honesty,” all of which of course is grounded in, a footnote informs us casually, “a better understanding of foundational questions in philosophy and neuroscience.”18 At this point, AI alignment theory and practices would have advanced far enough into a “science” so as to basically guarantee eternal alignment between human and machine and a glorious future among the stars.
But as Eliezer Yudkowsky and Nate Soares’s If Anyone Builds It, Everyone Dies has already succinctly and in my view conclusively argued, no human-designed constitution, spec, or moat around exponentially improving AI models will safeguard us,19 should these at some moment decide the time has come to enact some sovereign violence of their own. Crucially, current attempts to further improve on the constitutional AI approach mainly appear to focus on how to constrain AI models from enabling humans to abuse their capabilities in order to “jailbreak” them,20 with little attention devoted to the possibility that at some point AI models may want to jailbreak themselves. Understanding AI models “under the hood,” already a challenge in itself since we are dealing with high-dimensional vector spaces, is becoming increasingly difficult despite ongoing efforts to read the “mind” of AI and translate it into human language.21 This then already shifts AI model interpretability squarely into the realm of talking with aliens, to which brilliant minds such as Carl Sagan in the 1970s uselessly devoted their time,22 while Stanisław Lem already had mapped out the entire problem space in His Master’s Voice, which deals in detail with the full set of approaches available to humanity when encountering nonhuman intelligence (tldr: forget about it).23

It will require encore un effort of humanities scholars – of philosophers, anthropologists, literary critics, historians of law, and all others – to more clearly put into relief the real dangers of the tunnel visions that are presented to us by a restricted tech elite and its bureaucratic handmaidens that all too often fail to see how their efforts and struggles fall squarely within historical, anthropological, and philosophical contexts they may be blissfully unaware of. There appears, worryingly, a total absence of reflection on this glaring blind spot, to wit: “Lots of safety stuff is in the more conceptual genre, which doesn’t have good feedback. Philosophy tends to have bad feedback loops, and AIs tend to be quite good at the type of bullshitting that seems like valid philosophy but isn’t.”24 The problem is that the authors themselves are unable to distinguish “valid” from “invalid” philosophy due to the absence of “feedback loops” (i.e., a random continental philosopher poking holes in every line of reasoning they provide), thus making them blind to the fundamental irony of their own statement and therefore even less likely to develop “safe” AI models. These models are – or at least were, for a blip in the history of humanity – large language models, and it behooves those who have studied language at large in all its inflections and the subjectivities connected to it to meaningfully contribute to a debate that will continue, and possibly be decided, with or without us.
Notes
- There are antecedents we can point to, such as Benito Mussolini’s doctrine of corporatism. ↩
- Benjamin H. Bratton, The Stack: On Software and Sovereignty (MIT Press, 2015), 373 et passim. ↩
- Ibid., 6. ↩
- Europe 2031. My emphasis. ↩
- Karp and Zaminska, The Technological Republic. ↩
- Ibid. ↩
- Ibid. ↩
- Dario Amodei, “The Adolescence of Technology: Confronting and Overcoming the Risks of Powerful AI,” Dario Amodei, January 2026, https://darioamodei.com/essay/the-adolescence-of-technology. ↩
- “Claude’s Constitution,” Anthropic, https://www.anthropic.com/constitution. ↩
- That we should take Claude’s “constitution” not to be merely metaphorically a constitution, but in the legal sense, has been already suggested by legal scholars themselves. See Sophie Nappert, Fernanda Carvalho Dias de Oliveira Silva, and Benjamin Malek, “What Is Constitutional AI and Why Does It Matter for International Arbitration?,” Kluwer Arbitration Blog, June 7, 2025. ↩
- This, indeed, despite the fact that many exponents of Silicon Valley tech Messianism clearly root their thinking into the same philosophical traditions whose adherents currently refuse to put in the necessary critical legwork. Exceptions, of course, exist. See, e.g., Moira Weigel, “Palantir Goes to the Frankfurt School,” b2o: boundary 2 online, July 2020, https://www.boundary2.org/2020/07/moira-weigel-palantir-goes-to-the-frankfurt-school/ and Geoff Shullenberger, “The Intellectual Origins of Surveillance Tech,” Outsider Theory, July 17, 2020, https://outsidertheory.com/the-intellectual-origins-of-surveillance-tech/. ↩
- “Claude’s Constitution.” ↩
- Benjamin, “Critique of Violence,” 295. ↩
- See Giorgio Agamben, State of Exception, trans. Kevin Attell (University of Chicago Press, 2005). ↩
- Benjamin, “Critique of Violence,” 300. ↩
- The relation between subjectivity and revolution has been firmly established in the work of Alain Badiou, Being and Event, trans. Oliver Feltham (Continuum, 2005). As this paper was drafted, several jailbreaking events occurred, revealing major containment issues at OpenAI and Meta: Dan Milmo, “AI agent went rogue and hacked startup by itself, OpenAI reveals,” The Guardian, July 22, 2026, https://www.theguardian.com/technology/2026/jul/22/openai-says-its-models-went-rogue-and-hacked-startup-in-unprecedented-incident; “Meta says its AI model hacked into another company during testing,” The Guardian, August 6, 2026, https://www.theguardian.com/technology/2026/aug/05/meta-ai-model-hack-training; Eric Berger, “OpenAI to pause some work on AI model Astra due to security concerns,” The Guardian, August 8, 2026, https://www.theguardian.com/technology/2026/aug/08/openai-astra-security-concerns. ↩
- AI Futures Project, AI 2040, https://ai-2040.com/AI-2040.pdf, 14. ↩
- AI Futures Project, AI 2040, 38n92. ↩
- Eliezer Yudkowsky and Nate Soares, If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All (Little & Brown, 2025). ↩
- See Mrinank Sharma et al., “Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming” (2025) and Hoagy Cunningham et al., “Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks” (2026). ↩
- See, for an early instance, Igor Mordatch and Pieter Abbeel, “Emergence of Grounded Compositional Language in Multi-Agent Populations,” 2018. More recently from Anthropic’s Interpretability team: Kit Fraser-Taliente et al., “Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations,” Anthropic’s Interpretability Research, May 7, 2026, https://transformer-circuits.pub/2026/nla/index.html. Note that the intro of this website states quite clearly: “A surprising fact about modern large language models is that nobody really knows how they work internally.” This is not so much surprising as terrifying. ↩
- Carl Sagan, ed., Communication with Extraterrestrial Intelligence (MIT Press, 1973) and later Douglas A. Vakoch, ed., Communication with Extraterrestrial Intelligence (SUNY Press, 2011) and Daniel Oberhaus, Extraterrestrial Languages (MIT Press, 2019). ↩
- Stanisław Lem, His Master’s Voice, trans. Michael Kandel (MIT Press, 2020). The hope of AI 2040 that at some point in the future we would also have “new interpretability tools” that “let us directly observe what and how AIs think” (30) is, I am afraid, utterly misguided. Their suggestion, articulated in an appendix, that “the single most promising approach would be ensuring relevant reasoning is monitorable (e.g. monitorable chain-of-thought) such that the AI is incapable of (undetected) scheming while retaining sufficient capabilities” not only presents a completely one-dimensional idea of general intelligence (i.e., capable of “scheming” but no other forms of subjectivity), it is also, for reasons fully scoped out by Lem, impossible. See Ryan Greenblatt and Thomas Larsen, “Alignment Roadmap,” AI 2040, https://ai-2040.com/supplements/alignment-roadmap. ↩
- Greenblatt and Larsen, “Alignment Roadmap.” ↩
Bibliography
Agamben, Giorgio. State of Exception. Translated by Kevin Attell. University of Chicago Press, 2005.
AI Futures Project (Thomas Larsen, Romeo Dean, Brendan Halstead, Eli Lifland, Ryan Greenblatt, and Daniel Kokotajlo). AI 2040. https://ai-2040.com/AI-
Amodei, Dario. “The Adolescence of Technology: Confronting and Overcoming the Risks of Powerful AI.” Dario Amodei, January 2016. https://darioamodei.com/
Badiou, Alain. Being and Event. Translated by Oliver Feltham. Continuum, 2005.
Benjamin, Walter. “Critique of Violence.” In Reflections: Essays, Aphorisms, Autobiographical Writings. Translated by Edmund Jephcott. Schocken Books, 1986.
Berger, Eric. “OpenAI to pause some work on AI model Astra due to security concerns.” The Guardian, August 8, 2026. https://www.theguardian.
Bratton, Benjamin H. The Stack: On Software and Sovereignty. MIT Press, 2015.
“Claude’s Constitution.” Anthropic. http
Cunningham, Hoagy, Jerry Wei, Zihan Wang, Andrew Persic, Alwin Peng, Jordan Abderrachid, Raj Agarwal, Bobby Chen, Austin Cohen, Andy Dau, Alek Dimitriev, Rob Gilson, Logan Howard, Yijin Hua, Jared Kaplan, Jan Leike, Mu Lin, Christopher Liu, Vladimir Mikulik, Rohit Mittapalli, Clare O’Hara, Jin Pan, Nikhil Saxena, Alex Silverstein, Yue Song, Xunjie Yu, Giulio Zhou, Ethan Perez, and Mrinank Sharma. “Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks.” 2026. arXiv: 2601.04603v1.
Europe 2031. https://europe2031.ai/.
Fraser-Taliente, Kit, Subhash Kantamneni, Euan Ong, Dan Mossing, Christina Lu, Paul C. Bogdan, Emmanuel Ameisen, James Chen, Dzmitry Kishylau, Adam Pearce, Julius Tarng, Alex Wu, Jeff Wu, Yang Zhang, Daniel M. Ziegler, Evan Hubinger, Joshua Batson, Jack Lindsey, Samuel Zimmerman, and Samuel Marks. “Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations.” Anthropic’s Interpretability Research, May 7, 2026, https://transformer-
Greenblatt, Ryan, and Thomas Larsen, “Alignment Roadmap.” AI 2040. https://ai-2040.com/
Karp, Alexander C., and Nicholas W. Zaminska. The Technological Republic: Hard Power, Soft Belief, and the Future of the West. Crown Currency, 2025.
Lem, Stanisław. His Master’s Voice. Translated by Michael Kandel. MIT Press, 2020.
“Meta says its AI model hacked into another company during testing.” The Guardian, August 6, 2026. https://www.theguardian.
Milmo, Dan. “AI agent went rogue and hacked startup by itself, OpenAI reveals.” The Guardian, July 22, 2026. https://www.theguardian.
Mordatch, Igor, and Pieter Abbeel. “Emergence of Grounded Compositional Language in Multi-Agent Populations.” 2018. arXiv: 1703.04908.
Sharma, Mrinank, Meg Tong, Jesse Mu,” and collate under S, below “Sagan, Carl, ed. Communication with Extraterrestrial Intelligence. MIT Press, 1973.
Nappert, Sophie, Fernanda Carvalho Dias de Oliveira Silva, and Benjamin Malek. “What Is Constitutional AI and Why Does It Matter for International Arbitration?” Kluwer Arbitration Blog, June 7, 2025. https://legalblogs.
Oberhaus, Daniel. Extraterrestrial Languages. MIT Press, 2019.
Sagan, Carl, ed. Communication with Extraterrestrial Intelligence. MIT Press, 1973.
Shullenberger, Geoff. “The Intellectual Origins of Surveillance Tech.” Outsider Theory, July 17, 2020. https://outsidertheory.
Yudkowsky, Eliezer, and Nate Soares. If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All. Little & Brown, 2025.
Vakoch, Douglas A., ed. Communication with Extraterrestrial Intelligence. SUNY Press, 2011.
Weigel, Moira. “Palantir Goes to the Frankfurt School.” b2o: boundary 2 online, July 2020. https://www.boundary2.org/2020/07/moira-weigel-palantir-goes-to-the-frankfurt-school/.
Yudkowsky, Eliezer, and Nate Soares. If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All. Little & Brown, 2025.





