Reading Jakub Pachocki on intelligence, agency, alignment, opacity, and recursive self-improvement
A critical reconstruction of Jakub Pachocki’s An Alien Mind, testing its claims about non-human intelligence, learned agency, epistemic opacity, alignment, monitoring, defensive AI, recursive self-improvement, consciousness, and human control against the broader philosophical and technical literature.
essay
machine learning
🇬🇧
Author
Affiliation
Antonio Montano
4M4
Modified
September 30, 2026
Abstract
In An Alien Mind, OpenAI Chief Scientist Jakub Pachocki argues that frontier AI is increasingly grown more than designed, that machine intelligence should not be treated as a simple point on a human scale, that alignment is fundamentally a problem of generalization, that monitoring may become harder as reasoning systems grow more capable, and that AI-assisted research could eventually develop into recursive self-improvement. Because these claims come from a senior scientist inside one of the organizations operating at the frontier of generative AI, they deserve to be treated neither as detached speculation nor as independently established fact, but as an unusually informed and comparatively favorable interpretation of the technological project itself.
This essay reconstructs Pachocki’s argument and tests it against competing accounts of intelligence, agency, artificial interlocutors, epistemic opacity, interpretability, recursive self-improvement, alignment, consciousness, and moral patienthood. Its central claim is that the most defensible meaning of alien is not an established non-human phenomenology but a multidimensional decoupling of properties that human-centered concepts tend to bundle together: competence need not establish understanding, agency need not imply human-like psychology, persistence need not map onto one enduring individual, inspectability need not yield intelligibility, recursive improvement need not amount to an intelligence explosion, optimization need not imply human values, and behavioral sophistication need not establish consciousness or moral status.
The practical importance of these distinctions extends well beyond philosophy of mind. Individuals, organizations, and democratic institutions increasingly make decisions about what authority to delegate to AI, which risks to prioritize, how responsibility should be allocated, what forms of oversight are justified, and how much control over consequential systems should remain human. Those decisions depend on the conceptual model of AI that precedes them. Anthropomorphic models can produce overtrust and misdirected fear; reductively mechanistic ones can obscure functional agency and systemic power. The essay therefore treats conceptual clarification as part of the epistemic infrastructure of democratic decision making: before societies decide whether to accelerate, constrain, regulate, delegate to, or reorganize themselves around advanced AI, they must first understand what kinds of capabilities, agency, opacity, alignment problems, feedback loops, and possible forms of subjecthood are actually at issue. The governing principle is simple: explanation before action.
Keywords
alien intelligence, alien mind, Jakub Pachocki, artificial intelligence, frontier AI, machine intelligence, artificial agency, AI agents, anthropomorphism, epistemic opacity, interpretability, chain-of-thought monitoring, AI alignment, value alignment, goal alignment, recursive self-improvement, AI-assisted research, intelligence explosion, artificial consciousness, philosophy of mind, moral patienthood, AI welfare, machine ethics, human agency, AI governance
A critical reconstruction of Jakub Pachocki’s An Alien Mind, testing its claims about non-human intelligence, learned agency, epistemic opacity, alignment, monitoring, defensive AI, recursive self-improvement, consciousness, and human control against the broader philosophical and technical literature.
Why the question matters
The question at the center of this essay, What kind of thing is an alien mind?, can sound like a problem for philosophy of mind or philosophical anthropology: interesting, perhaps profound, but remote from ordinary life. It is not. The categories through which we interpret artificial systems increasingly determine what we trust them to do, what authority we give them, which failures we anticipate, whom we hold responsible when something goes wrong, and eventually what obligations we may believe we have toward the systems themselves. A conceptual mistake about AI can therefore become a practical mistake long before philosophers agree on what, if anything, artificial systems ultimately are.
This is already visible in ordinary interaction. If a fluent system is implicitly treated as a person who understands, knows, remembers, and cares, users can attribute epistemic authority, reliability, confidentiality, or concern that its behavior does not warrant. A convincing explanation can be mistaken for evidence that the underlying conclusion is correct. Apparent empathy can be mistaken for a relationship in which the other party possesses corresponding feelings. Persistent conversational context can be interpreted as personal memory. A system that competently pursues a requested objective can be treated as though it also understands why that objective matters.
The opposite mistake is equally consequential. Describing AI as just software, just statistics, or just a tool can conceal forms of operational agency that matter even if no artificial consciousness exists. A system capable of selecting actions, invoking tools, accessing information, revising plans, communicating with other systems, and affecting a digital environment does not become equivalent to a human agent, but neither does it behave like a passive instrument. Whether it possesses beliefs or experiences may remain unsettled while its capacity to cause consequential events is perfectly real.
This distinction changes the practical meaning of risk. Many important AI risks do not require a malicious machine, a conscious machine, or even a machine with anything resembling human desires. A system can produce harmful consequences because an objective was badly specified, because a proxy ceased to represent the intended purpose, because an apparently reliable behavior failed under distribution shift, because an automated process acquired more authority than its evaluators could supervise, or because human users trusted an output whose production they did not understand. None of these mechanisms requires an evil AI. They require only a sufficiently capable system embedded in a sufficiently consequential process.
For the same reason, misunderstanding opacity changes how people respond to AI. If opacity is interpreted as literal unknowability, users may lapse into fatalism: either trust the machine because nobody can understand it, or reject it because it is a mysterious black box. But access, interpretability, explanation, validation, and justified trust are different things. A system can be technically inspectable while remaining difficult to explain; an explanation can be intelligible while the result remains wrong; a highly accurate system can be useful without its every internal representation being understood. Everyday decisions therefore require a more discriminating question than Do we understand the AI? They require asking what exactly is understood, by whom, for what purpose, and what independent evidence supports reliance on the result.
The same conceptual discipline matters when people hear that AI is becoming autonomous. Autonomy can mean that a system acts without a human selecting every intermediate step. It can mean that it generates and revises subgoals. It can mean that it operates for long periods without supervision. None of these necessarily means that it has chosen its ultimate objectives for itself. Confusing operational autonomy with psychological independence can produce exaggerated fears about artificial wills; confusing externally supplied objectives with harmless passivity can produce the opposite error. What matters operationally is how much discretion the system possesses, what resources it can access, how its objectives are represented, which actions require authorization, and where human intervention remains effective.
Questions about recursive self-improvement have the same structure. The phrase can evoke a machine independently redesigning itself until humans lose control. Yet there is a large practical space between ordinary software automation and that hypothetical endpoint. AI systems can already participate in research, coding, evaluation, experimentation, and the development of successor systems without independently choosing research priorities or validating their own progress. Whether more of that loop closes matters because feedback can accelerate technological change even before there exists anything resembling one persistent artificial mind improving itself. For workers, organizations, researchers, governments, and citizens, the relevant question is therefore not merely whether an intelligence explosion will occur. It is which parts of cognitive and institutional processes are being delegated to machines, how quickly those feedback loops operate, and whether human evaluation can still keep pace.
Alignment is similarly an everyday problem before it becomes an existential one. Whenever a person asks an AI system to optimize something, for example, write the best answer, find the cheapest route, maximize engagement, identify the strongest candidate, detect suspicious behavior, recommend treatment options, select information, or complete a complex task, then the desired outcome is richer than the measurable instructions supplied to the machine. The gap between what people mean and what a system can operationalize exists at every scale. Frontier alignment research studies extreme forms of a problem that already appears whenever a proxy, metric, instruction, or evaluation criterion imperfectly captures human intent.
These distinctions also determine how responsibility is distributed. If an artificial system is treated as a passive tool, responsibility may be attributed entirely to the person who activated it. If it is rhetorically elevated into an independent agent, responsibility may instead disappear into the machine. Both can be misleading. AI systems can acquire genuine causal importance while remaining products of developers, deployers, institutional processes, access controls, objectives, data, and human decisions. Understanding artificial agency is therefore also a prerequisite for understanding human accountability in systems where no single human specifies every action that occurs.
And eventually the conceptual error could reverse direction. If future systems acquire properties relevant to consciousness or welfare, assuming that artificial implementation automatically excludes subjectivity could become morally consequential. Yet treating present-day linguistic self-reports as proof of consciousness would be equally premature. The distinction between what a system can do, what it can optimize, what it can report, and what, if anything, it can experience therefore matters not only for controlling machines but also for determining whether some future machines could themselves become objects of moral concern.
The practical problem can be stated simply: people act according to the model of AI they carry in their heads. If that model is anthropomorphic, they may overtrust, overinterpret, personalize, or fear systems for the wrong reasons. If it is reductively mechanistic, they may underestimate functional agency, emergent capability, systemic dependence, or eventually morally relevant properties. If it treats opacity as magic, responsibility and verification become impossible. If it treats intelligence, agency, autonomy, consciousness, and moral status as interchangeable, every subsequent debate becomes confused.
There is another reason to take these distinctions seriously in the particular case that motivates this essay. Pachocki is not observing the development of generative AI from outside it. He is the Chief Scientist of OpenAI, one of the organizations operating at the frontier of generative AI development, and writes from direct involvement in the research trajectory he is trying to interpret.1 OpenAI remains one of the principal companies shaping frontier models, agentic systems, and the broader direction of generative-AI development.
That makes his perspective unusually important, but not because institutional proximity makes it automatically correct. It gives the argument a distinctive epistemic position. Pachocki has access to internal experiments, development trajectories, failures, capability changes, and research practices that outside observers can only partially reconstruct. At the same time, he speaks from within an organization whose purpose and incentives are tied to continued progress in advanced AI. His viewpoint can therefore reasonably be treated as an informed and comparatively favorable interpretation of the technological project itself, rather than as the judgment of a detached critic.
This matters for comprehension. When a senior scientist inside a frontier laboratory describes the systems being developed there as increasingly non-human in their intelligence profile, difficult to monitor, capable of assuming larger parts of the research process, and potentially requiring development to slow when alignment and monitoring cannot keep pace, those claims deserve attention even before one decides whether they are ultimately correct.2 The important fact is not that an industry insider should be granted epistemic authority by default. It is that one of the actors with both unusually strong access to the technology and unusually strong reasons to believe in its value is himself framing its development as a problem of understanding, control, and human agency.
That combination makes An Alien Mind particularly useful for the purpose of this essay. It provides something close to a strong internal case for the significance of the transition: not a critique constructed by opponents of generative AI, but an attempt by a scientist helping to advance it to explain what he believes is changing and why the resulting systems may require new conceptual and institutional responses. The appropriate method is therefore neither deference nor suspicion by association. It is to reconstruct the argument in its strongest form, distinguish what derives from public evidence from what depends on internal observations or forecasts, and then test each claim against independent technical and philosophical work.
That exercise matters directly to democratic judgment. Citizens and political institutions will encounter competing descriptions of AI from companies, researchers, governments, critics, workers, investors, and civil society, each situated differently with respect to the technology and its consequences. Understanding what a technically sophisticated and institutionally favorable account actually claims provides an important reference point. Before deciding whether society should accelerate, constrain, delegate to, regulate, or otherwise reorganize itself around advanced AI, it should first understand the case made by those closest to building it, and then subject that case to independent scrutiny.
That matters not only for private choices but for collective decisions in a democracy. Artificial intelligence is increasingly relevant to employment, education, scientific research, public administration, security, information systems, critical infrastructure, economic concentration, and the distribution of cognitive and organizational power. Decisions about these systems will therefore increasingly be made through political institutions: legislatures, regulators, courts, public administrations, elections, procurement rules, liability regimes, international agreements, and decisions about which forms of human judgment may legitimately be delegated to machines.
Citizens cannot decide every technical question directly. Democratic government necessarily works through delegation: voters delegate authority to representatives; legislatures delegate implementation to administrations and regulators; public institutions rely on technical experts; and those institutions increasingly decide how much operational or cognitive authority may itself be delegated to automated systems. But delegation can be meaningfully evaluated only if the underlying phenomenon has been framed with sufficient accuracy. A democracy can faithfully execute a badly framed problem.
The distinction is consequential. A society that imagines AI primarily as an emerging artificial person may concentrate on intentions, consciousness, or hypothetical rebellion while paying insufficient attention to authority, incentives, access, accountability, objective specification, and concentration of power. A society that treats AI as merely another passive software tool can make the opposite error, allowing systems with substantial operational autonomy and causal reach into consequential processes without controls appropriate to that reach. Treating opacity as intrinsic mystery can encourage either unquestioning deference or indiscriminate prohibition. Treating benchmark performance as a scalar measure of intelligence can distort judgments about where machine decision making is or is not an adequate substitute for human judgment. Treating alignment as equivalent to obedience can conceal the harder question of whether the objectives being implemented are themselves legitimate, complete, or properly specified.
Political action is therefore shaped upstream by conceptual categories. Before a society can decide how much autonomy to permit, what forms of oversight to require, which decisions must remain contestable, where responsibility should reside, how concentrated control over advanced systems may become, or what evidence should justify intervention, it needs an adequate account of the thing being governed. These are not questions that technical specialists can settle on behalf of everyone else, because many concern the distribution of power, risk, responsibility, rights, and public resources. But democratic deliberation cannot evaluate those choices intelligently if its conceptual vocabulary collapses the phenomena it is supposed to govern.
This gives the distinctions developed in this essay a civic as well as philosophical function. Understanding intelligence, agency, autonomy, opacity, alignment, recursive improvement, consciousness, and moral status becomes part of the epistemic infrastructure of democratic decision making. Citizens do not need to become machine-learning researchers or philosophers of mind. They do need enough conceptual resolution to distinguish empirical claims from forecasts, capability from agency, agency from consciousness, explanation from justification, operational autonomy from self-authorship, and demonstrated technical risks from narratives that depend on additional assumptions.
The same requirement applies to political representatives and institutions. Calls to accelerate, constrain, regulate, deploy, prohibit, subsidize, decentralize, or internationally coordinate AI development all presuppose some account of what kind of phenomenon is being acted upon and which causal mechanisms generate the relevant benefits or risks. Different descriptions can imply radically different interventions. The quality of democratic choice therefore depends partly on whether those descriptions survive technical and philosophical scrutiny before they are converted into political action.
This is particularly important because some decisions about AI may change the conditions under which later decisions are made. Delegating cognitive work to machines, embedding automated judgment in institutions, concentrating computational infrastructure, restructuring labor markets, or allowing AI systems to participate increasingly in the development of successor technologies can alter the distribution of capabilities and authority on which future choices depend. The relevant question is therefore not only whether today’s decision is defensible. It is also whether today’s decision preserves the ability of individuals and democratic institutions to make meaningful choices tomorrow.
The sequence should consequently be explanation before action. First determine what property is actually present, how strong the evidence is, what mechanisms produce it, which uncertainties remain, and which stronger conclusions require additional premises. Only then ask what personal, technical, institutional, economic, or political response follows. Reversing that order, i.e.beginning from a preferred intervention and selecting whichever description of AI supports it, turns uncertainty into ideology.
This is why the apparently philosophical question posed by this essay has consequences well beyond philosophy. It affects how an individual should interpret an AI assistant, how an organization should decide what can safely be delegated, how engineers should construct controls and evaluation, how institutions should allocate responsibility, and how citizens should evaluate political choices concerning systems capable of changing the distribution of knowledge, labor, authority, and power.
The purpose of this essay is therefore not to decide whether machines are really like us. It is to develop a vocabulary adequate to systems that may be increasingly consequential precisely because they are not organized like us. Before asking whether AI is safe, aligned, autonomous, trustworthy, conscious, deserving of moral consideration, or appropriately governed, we need to know what each of those claims means, what evidence would make it true, and which conclusions actually follow from it.
For individuals, institutions, and democracies, the order matters:
first understand what kind of thing we are dealing with; then decide what to do about it.
Otherwise, the greatest risk is not merely that we misunderstand an alien intelligence. It is that we build personal, institutional, and political decisions on a category error, and discover only afterward that we were governing the wrong problem.
Alienness as a problem of categories
Pachocki’s thesis in short
Jakub Pachocki’s 2026 essay An Alien Mind is unusual not merely because of the strength of its claims, but because of where those claims come from. Pachocki is Chief Scientist at OpenAI, a company directly engaged in developing and deploying frontier generative-AI systems, and previously led research programs including GPT-4 and OpenAI Five.3 The essay is therefore not an external philosophical speculation about what advanced AI might eventually become. It is an interpretation offered by one of the senior scientists responsible for building the class of systems being interpreted. That institutional position does not make the claims correct, and internal evidence is not independent evidence, but it changes their epistemic significance: predictions about future capability, alignment, monitoring, and recursive self-improvement are being made from inside one of the organizations attempting to push those capabilities forward.
Pachocki’s starting point is a specific episode in OpenAI’s development of reasoning models. He writes that results obtained in the RLSlow research project in mid-2023 convinced him and his colleagues that reasoning-model training could continue to scale and that pretrained models could be induced to develop increasingly powerful chains of thought. His reaction, as he reconstructs it three years later, was not primarily excitement about benchmarks or products but the expectation that machines meaningfully smarter than ourselves would appear within his lifetime and that the broad form of those systems was already becoming visible.4 By September 2026, he argues, reasoning models are already operating computers and graphical interfaces, collaborating with humans and other AI systems, conducting research projects, and becoming relevant to domains such as cybersecurity. On the basis of what he explicitly describes as internal results, he further states a strong expectation that the present rate of capability progress could extend into recursive self-improvement.
The conceptual core of the essay comes before that forecast. Pachocki argues that modern AI is grown more than designed. Engineers choose architectures, training procedures, objectives, datasets, and enormous computational budgets, but the detailed internal organization responsible for the resulting capabilities is produced through optimization rather than specified component by component. Large-scale training runs are therefore, in his description, not merely acts of conventional software construction but experiments whose outcomes can still surprise their designers. Researchers can identify individual mechanisms inside trained systems, much as neuroscientists can study mechanisms in brains, while the overall behavior of the system can remain resistant to a description that humans fully understand.5
This claim about construction leads directly to his claim about intelligence. Pachocki does not present artificial intelligence as a lower or higher point on a single human scale. He argues that intelligence produced by deep-learning scaling is not directly comparable to human intelligence. A machine need not exceed human beings along every cognitive dimension to become transformative, useful, or dangerous; it need only exceed them along enough consequential axes. As different capabilities improve at different rates, and as easily measured capabilities can advance faster than abilities that are difficult to quantify, the question How intelligent is the system? itself becomes increasingly difficult to answer through one aggregate comparison.6
The remainder of An Alien Mind follows from this picture of an increasingly capable but differently constituted intelligence. Because machine intelligence arises through a process unlike human cognitive development, Pachocki argues that there is no reason to assume that it will spontaneously inherit or generalize human principles. He distinguishes goal alignment, i.e.whether a system attempts to accomplish the objective placed before it, from value alignment, which he treats as the stronger ability to preserve and generalize high-level principles under ambiguous, unfamiliar, conflicting, or adversarial conditions. In his account, the central technical difficulty is not merely teaching desirable behavior under familiar training conditions but establishing that those values continue to generalize as models become more capable and operate farther outside the environments in which alignment was trained.7
This generalization problem is also why Pachocki assigns unusual importance to monitoring. OpenAI’s principal approach, he writes, has been chain-of-thought monitoring: allowing reasoning models to externalize part of their reasoning while avoiding direct optimization pressure on that reasoning trace, in the hope that undesirable strategies remain observable rather than becoming optimized for concealment. Yet the essay is explicitly pessimistic about the durability of this mechanism. Pachocki reports that monitoring becomes harder as reasoning is distributed across communication, tool use, interactions with other agents, and internal computation that no longer requires verbalized chains of thought; he also states that models are becoming better at reasoning about and manipulating their own reasoning processes. His resulting expectation is striking: progress in general AI may increasingly become bottlenecked not by the ability to create more capable systems, but by confidence that those systems can still be monitored.8
Recursive self-improvement enters the argument at this point, not as an isolated science-fiction scenario but as Pachocki’s extrapolation from the increasing role of AI in AI research itself. He expects machine intelligence to take an ever larger part in developing successor systems and describes automated AI research as a more dramatic extension of scaling intelligence with compute. Yet his policy conclusion is not that this process should simply be accelerated. Pachocki explicitly argues that the relevant problem is how to reach increasingly automated AI research while keeping humans inside the improvement loop. He proposes two complementary levers: improve alignment and monitoring sufficiently to preserve human control, and slow development when confidence in those measures is inadequate. In his view, continued scaling should eventually be conditional on enforceable safety bars, potentially involving independent auditors, governments, or international institutions.9
The essay therefore makes a larger political claim as well as a technical one. Pachocki worries that highly capable AI could concentrate power by allowing a small number of people controlling large computational resources to perform undertakings that previously required thousands of specialists. His closing concern is consequently not merely whether individual AI systems behave correctly but whether humans retain agency over a technological process increasingly capable of automating cognition, research, and eventually parts of its own improvement. He concludes that, in his judgment, no laboratory has yet solved alignment and monitoring well enough to justify unrestricted maximum-speed scaling for much longer and argues for voluntary slowdowns and international coordination where safety confidence is insufficient.10
Taken together, these claims constitute what can reasonably be called Pachocki’s alien-intelligence thesis. Its components should nevertheless be distinguished. Some are descriptions of the deep-learning development process: capabilities emerge through optimization rather than explicit programming. Some are empirical interpretations: capability profiles are heterogeneous, trained systems are difficult to understand globally, and monitoring becomes harder as agency and reasoning become more complex. Some are claims about alignment: intelligence developed through a different process cannot be assumed to generalize human values in a human-like way. Some are extrapolations from internal evidence: AI research will become increasingly automated and may enter recursive self-improvement. And some are normative and political conclusions: capability growth should be paced by safety confidence, humans should remain inside the improvement loop, concentration of power should be constrained, and international coordination may be required.
These propositions do not all have the same evidential status. In particular, the claim that continued capability growth will lead to recursive self-improvement remains a forecast, while assertions based on OpenAI’s internal evaluations are primary institutional evidence rather than independent confirmation. But dismissing the essay as ordinary speculation would also miss its significance. The striking fact is that the Chief Scientist of a frontier AI developer is publicly describing the systems his organization is building in terms of non-human intelligence, incomplete understanding, diminishing reliability of chain-of-thought monitoring, possible recursive self-improvement, and the need to slow development if safety does not keep pace. The epistemic question is therefore not whether the language is dramatic. It is which parts of the underlying argument survive comparison with the broader literature.
Competing interpretations of artificial intelligence
The philosophically important problem begins with the word alien. What would it mean for an intelligence to be genuinely alien rather than merely faster, larger, or more capable than a human one?
Calling such a system an alien mind risks answering that question too quickly. Human beings present us with a familiar cluster of properties: flexible competence, intentional action, linguistic understanding, embodiment, autobiographical continuity, conscious experience, and the standing to be treated as subjects rather than objects. These properties are not identical even in the human case, but ordinary social cognition encounters them together often enough that our vocabulary encourages us to infer one from another. I will call this tendency the anthropomorphic bundle. The term does not denote a scientific taxonomy; it names an inferential shortcut. When an entity converses fluently, solves difficult problems, pursues goals, remembers earlier interactions, describes its apparent internal states, and adapts its behavior, the human template invites a compressed judgment: there is an intelligent agent there, it understands what it is doing, and there is a persisting subject to whom the understanding belongs.
Advanced AI makes that compression increasingly unsafe because the properties in the bundle can come apart.
Murray Shanahan’s analysis of large language models shows why this separation is needed even before questions about consciousness arise. He distinguishes a bare language model from a larger system built around it and warns that words such as knows, believes, and thinks can function as useful predictive shorthand while misleading us if they are automatically interpreted as claims about a human-like inner life.11 His position is not captured by the slogan that language models are mere statistics. The more precise point is that useful intentional description and human psychological equivalence are different claims.
Emily Bender and Alexander Koller sharpen one part of this caution by distinguishing linguistic form from meaning. Their argument is specifically that a system trained only on form cannot thereby acquire meaning as they define it, because meaning involves relations between expressions and communicative intentions or worldly referents.12 That is an argument about what form-only training establishes. It is not a timeless empirical verdict on every multimodal, embodied, tool-using, or environmentally coupled architecture that might later be built. Treating it as the latter would replace one conceptual overreach with another.
The opposite danger is to define the relevant categories so tightly around their human realization that artificial cases are excluded by construction. Luciano Floridi proposes one response by moving the center of analysis away from intelligence. His Multiple Realisability of Agency thesis treats contemporary AI as a novel form of agency without intelligence: consequential, adaptive, goal-directed behavior may be realized without cognition, consciousness, intention, or mental states in the stronger senses associated with human agents.13 David Chalmers approaches a neighboring boundary from the other direction. Against the claim that genuine thought necessarily requires sensory grounding, he argues that sophisticated pure thinkers are at least philosophically possible and that the absence of ordinary sensory capacities therefore cannot by itself establish that an artificial system is incapable of thought. Crucially, he explicitly stops short of concluding that contemporary large language models actually think or understand.14
Floridi and Chalmers illuminate different sides of the same methodological problem. Successful action does not force us to attribute mentality, but departure from the biological route by which humans acquire mentality does not by itself justify withholding it either.
The analytical position of this essay
The central thesis of this essay follows from that underdetermination. Pachocki is making a stronger claim than the familiar observation that modern neural networks are complicated, but the word mind potentially carries more ontological weight than the technical argument itself requires. What is most plausibly alien about advanced AI is not yet an established alien phenomenology, nor is it well represented by placing humans and machines on one scalar axis of intelligence. The deeper discontinuity is that artificial systems can separate properties that the human case encourages us to bundle.
A system may exhibit broad problem-solving competence without settling whether it possesses semantic understanding; it may participate in long-horizon goal pursuit without having the psychological unity of a human agent; a conversational entity may display continuity although the underlying model is copied across many processes; a deterministic and technically inspectable network may remain cognitively opaque to its investigators; a system may improve outputs, tools, or training procedures without closing the loop required for autonomous recursive self-improvement; and increasingly sophisticated behavior may make questions of consciousness and welfare harder without answering them.
Alienness, on this account, is multidimensional decoupling rather than simply superhuman performance.
This interpretation neither accepts nor rejects Pachocki’s use of mind at face value. Instead, it decomposes the thesis. His claim that machine intelligence is not directly comparable with human intelligence can be tested against formal and psychometric theories of machine intelligence. His account of emergent agency can be compared with intentional-stance and artificial-agency theories. His claim that the resulting systems are increasingly difficult to understand can be separated into computational complexity, epistemic opacity, interpretability, and justification. His expectation of recursive self-improvement can be compared with earlier intelligence-explosion arguments and the growing empirical literature on automated research loops. His alignment argument can be examined against theories of objective specification, generalization, corrigibility, and instrumental convergence. And the word mind itself can finally be tested against contemporary work on artificial consciousness and moral patienthood.
Errors can consequently run in both directions. Anthropomorphic over-attribution turns behavioral evidence into stronger ontological claims than it supports: benchmark success becomes unrestricted intelligence, linguistic fluency becomes understanding, self-description becomes introspective access, and apparent goal pursuit becomes evidence of stable desires. Anthropomorphic under-attribution commits the complementary mistake: because a system lacks a human body, developmental history, emotional architecture, or familiar mode of persistence, capacities that could in principle have non-human realizations are dismissed before the relevant evidence is examined.
The appropriate response is neither a general presumption for nor a general presumption against artificial mentality. It is to ask, for each contested property, what would constitute evidence for it, which architectural or environmental conditions matter, which alternative explanations remain live, and how conclusions should change when the evidence supports only neighboring properties.
The rest of the argument therefore does two things in parallel. It reconstructs Pachocki’s An Alien Mind more fully, that is its account of scaled deep learning, alignment, generalization, monitoring, defensive AI, recursive self-improvement, safety-gated scaling, and continued human control, and places each component against the strongest relevant competing interpretations in philosophy, computer science, and AI-safety research. The purpose is not to decide in advance whether Pachocki is right to call the resulting entity a mind. It is to determine precisely which parts of the alien-mind thesis are empirical observations, which are theoretical interpretations, which are forecasts, and which require additional philosophical premises before the language of mentality is warranted.
Intelligence without anthropomorphism
If alien intelligence is to mean more than intelligence embodied in an unfamiliar object, the first difficulty is that the word intelligence itself carries human assumptions. We ordinarily recognize intelligence through a mixture of behavioral flexibility, linguistic competence, learning, planning, abstraction, problem solving, and social understanding because these capacities are correlated in human beings. An artificial system need not preserve those correlations. It can therefore appear extraordinarily capable under one description and strangely deficient under another without either observation being anomalous.
Pachocki’s claim that machine intelligence is not directly comparable with human intelligence is strongest when read in this sense rather than as the stronger claim that comparison is impossible.15 Comparison requires a choice of dimensions, environments, resources, and success criteria. Once those choices are explicit, comparisons become possible, but they may produce a profile rather than a single ordering.
One influential attempt to construct a non-anthropomorphic definition predates modern large language models. Shane Legg and Marcus Hutter begin from a basic difficulty: definitions and tests developed around human intelligence may not generalize cleanly to artificial systems whose senses, environments, motivations, and cognitive capacities differ radically from ours. After surveying existing definitions, they characterize intelligence informally as an agent’s general ability to achieve goals across a wide range of environments and then formalize that intuition.16 For an agent \pi, their universal-intelligence measure is
where E is the space of computable reward-summable environmental measures relative to a reference universal Turing machine U, \mu denotes one such environment, and V_{\mu}^{\pi} is the expected total reward obtained by agent \pi in that environment. K(\mu) is the Kolmogorov complexity of the environment relative to U. The factor 2^{-K(\mu)} implements an algorithmic-probability weighting: simpler environments receive greater weight than more complex ones.
Conceptually, Equation 1 defines intelligence as a weighted measure of goal-achieving performance across possible environments. The agent \pi is not evaluated on one benchmark, one cognitive faculty, or one predefined task family, but across the class E of computable reward-summable environments. For each environment \mu, the quantity V_{\mu}^{\pi} represents the expected total reward obtained by the agent through interaction with that environment. The equation therefore separates two elements that ordinary intelligence tests often leave implicit: how well the agent performs in a particular environment, represented by V_{\mu}^{\pi}, and how much that environment contributes to the overall measure, represented by 2^{-K(\mu)}.
The weighting term is crucial. K(\mu) is the Kolmogorov complexity of environment \mu relative to the chosen reference universal machine: approximately, the length of the shortest program capable of specifying that environment. Because an environment’s contribution is multiplied by 2^{-K(\mu)}, environments with shorter computational descriptions receive greater weight than environments requiring longer descriptions. The construction therefore implements an algorithmic version of Occam’s principle. It does not average indiscriminately over all conceivable worlds; it gives systematically greater importance to simpler computable environments while still evaluating the agent across an extremely broad space of possible interaction structures.
This changes what counts as intelligence. A system that performs extraordinarily well in one narrow environment but poorly elsewhere is not automatically highly intelligent under the measure simply because its peak performance is exceptional. Its contribution in that environment is only one term in the sum. By contrast, an agent capable of obtaining substantial reward across many different environments accumulates contributions across the distribution. The formalism is therefore intended to capture breadth of goal-achieving competence, not excellence at a privileged task. It aggregates expected performance across environments whose dynamics, observations, available actions, and reward structures can differ; it does not, by itself, measure how efficiently skills are acquired or transferred between environments. Those are separate dimensions considered below.
Just as importantly, nothing in the equation requires the agent to achieve those results through human-like cognition. There is no term for biological substrate, neural similarity, language, embodiment, education, cultural background, introspection, or resemblance to human reasoning strategies. The same formal criterion can in principle be applied to agents implemented in biological nervous systems, artificial neural networks, symbolic programs, or entirely different computational architectures. The object being evaluated is the agent’s policy as revealed through its interaction with environments, not the internal mechanism by which the policy is realized.
The measure is therefore non-anthropomorphic in a precise sense. It replaces the question How closely does this system reproduce the abilities through which humans display intelligence? with a more abstract one: Across how broad a space of environments can this agent successfully act so as to achieve goals? Human intelligence becomes one possible realization of such general competence rather than the definition against which every other realization must be measured.
That abstraction demonstrates something important, but not what a loose appeal to machine IQ might suggest. Legg and Hutter’s measure does not establish that intelligence has one uniquely correct definition, nor does it provide a practical score for contemporary systems. Kolmogorov complexity is uncomputable, and the result depends to some degree on the selected reference machine.17 Its value here is conceptual: it demonstrates that one can formulate a coherent notion of intelligence without appealing to human resemblance and, in doing so, exposes how much ordinary evaluation depends on unstated choices about which environments matter.
A chess engine may dominate humans in chess while being useless outside that domain; a human may perform poorly in environments whose interfaces exploit machine-native advantages; a general artificial system may possess a distribution of abilities that intersects the human distribution without being ordered cleanly above or below it as a whole. The sentence more intelligent than a human is therefore incomplete unless the relevant distribution of tasks, adaptation opportunities, resource constraints, and objectives has been specified.
François Chollet reaches a related conclusion from a different direction. He distinguishes skill, understood as performance on a task, from intelligence understood in terms of the efficiency with which useful skills are acquired given constraints on prior knowledge and experience.18 This matters because observed performance confounds what a system can do with what was required for it to become able to do it. A system may display remarkable competence because its learning procedure extracted transferable structure from limited experience; it may instead display the same competence because enormous quantities of relevant information were already encoded through pretraining, architecture, tools, or human-designed scaffolding. Equal output quality does not imply equal generality of the underlying learning processes.
Human–machine comparison makes this especially difficult. Humans do not arrive at a benchmark without priors: biological evolution, development, embodiment, language acquisition, education, and cultural participation have already shaped what a human subject can infer from a small number of examples. A pretrained model likewise does not arrive without priors; its parameters encode the consequences of optimization over a training distribution, while its architecture and surrounding system impose additional inductive structure. The relevant comparison is therefore between differently constituted learning systems whose prior information, acquisition histories, interfaces, and resource budgets are difficult to commensurate.
José Hernández-Orallo describes the corresponding change in evaluation as a move from task-oriented toward ability-oriented measurement.19 Rather than asking only whether a system solves a collection of tasks, an evaluator should attempt to characterize the more general capacities that explain performance across tasks and, ideally, across natural and artificial subjects. This distinction is particularly relevant when training corpora may contain close relatives of evaluation problems. Performance remains evidence of capability, but the inferential step from solving observed tasks to possessing a general ability depends on novelty, transfer, adaptation, and the system’s effective prior information.
None of this means intelligence can be reduced to learning speed. A system can learn efficiently yet lack long-horizon planning; transfer abstractions across domains while remaining brittle under perturbation; reason effectively with symbolic material yet possess weak sensorimotor competence. Conversely, substantial prior knowledge is not evidence against intelligence: human expertise itself depends on accumulated knowledge. The point is that competence, generality, adaptation, and resource efficiency are separable dimensions. Collapsing them into one scalar conceals precisely the structure that matters when comparing differently realized systems.
The problem becomes still more delicate when behavioral intelligence is connected to understanding. A Legg–Hutter-style definition can classify systems by goal-achieving ability without requiring any judgment about consciousness, semantic content, or phenomenology. Bender and Koller’s distinction between linguistic form and meaning can consequently coexist with substantial behavioral intelligence rather than simply contradicting it.20 A system might generalize across many linguistic tasks, plan effectively through language, and obtain high reward in a large class of environments while the separate question of what, if anything, its symbols mean to the system remains contested.
Sensory grounding illustrates the same logical separation. A familiar objection runs as follows: human concepts are grounded in perception and action; language models acquire much of their competence from linguistic data; therefore such models cannot genuinely think. Chalmers’s argument against the necessity of sensory grounding blocks the final inference without establishing its opposite.21 A causal history unlike ours is not by itself evidence that a capacity is absent, yet demonstrating that a non-human route could realize a capacity is not evidence that a particular machine has realized it.
This is where multiple realizability becomes central to the notion of alien intelligence. Two systems can exhibit overlapping behavioral capacities while arriving there through different substrates, learning histories, internal representations, and resource trade-offs. Conversely, systems with superficially similar architectures can develop sharply different competence profiles because their objectives, data, post-training procedures, memory, tools, and interaction environments differ. Alien should consequently not mean inexplicable in principle. It means that human cognitive architecture is no longer a safe default model for interpreting observed competence.
There can therefore be alien intelligence before there is evidence of an alien mind. A system may occupy a capability profile unlike that produced by human biological and cultural development while questions about understanding, subjective experience, and personal identity remain open. Once intelligence is detached from resemblance, however, another problem appears immediately. Goal-directed competence invites us to describe the system as an agent, and agency brings another set of human assumptions that must be separated in turn.
Agency without a human psychology
We readily move from saying that a system can do something to saying that it acts, from action to goal pursuit, and from goal pursuit to the language of beliefs, desires, intentions, decisions, and responsibility. In ordinary human contexts these transitions often work because the underlying properties are densely correlated. They are less secure when applied to artificial systems.
A thermostat changes its environment without being an agent in the ordinary psychological sense; an automated trading system can respond adaptively to market conditions without understanding finance; a software process can select among actions according to an objective without having adopted that objective for itself. The important question is therefore not whether AI really has agency in an all-or-nothing sense but which concept of agency is being invoked and what follows from adopting it.
Daniel Dennett’s intentional-stance framework provides one influential way to separate predictive usefulness from premature commitments about internal constitution. The intentional stance treats a system as if it had beliefs and desires and predicts its behavior by asking what a rational agent with those states would do. For Dennett, this strategy can apply to living and nonliving systems; its applicability does not require the analyst first to reconstruct every underlying physical mechanism.22 The practical usefulness of saying that a system wants,knows, or believes something therefore does not, by itself, demonstrate that it possesses those states in the same way a human being does.
The distinction is visible in ordinary engineering. Suppose an autonomous software system receives a request, decomposes it into subtasks, retrieves information, invokes external tools, checks intermediate results, revises a plan after failure, and stops when an acceptance condition is satisfied. Describing the system as trying to accomplish the request may compress a large amount of causal detail into a useful explanatory model. Saying that it looks for another route may be more intelligible than enumerating every token, branch, API invocation, and state transition. Nothing about this convenience determines whether the system possesses an intrinsic desire for task completion, a conscious intention, or even a persistent representation corresponding neatly to the human concept of a goal.
Functional agency and psychological agency are different claims. Floridi and J. W. Sanders make the separation explicit by analyzing agency at a chosen level of abstraction. Their account emphasizes interactivity, because system and environment can affect one another; autonomy, because the system can change state without direct external control at every step; and adaptability, because interaction can change in light of previous interaction.23 These criteria were deliberately formulated so that agency need not presuppose free will, consciousness, human mental states, or moral responsibility.
Floridi’s later Multiple Realisability of Agency thesis generalizes this strategy. Instead of extending intelligence until contemporary AI is accommodated within it, he proposes expanding agency to recognize artificial systems as a distinct realization that need not involve cognition, intention, consciousness, or mental states in the human sense.24 On this view, the striking development is not necessarily that engineers have manufactured synthetic human minds. It is that increasingly consequential behavior can be produced by systems whose causal organization supports interaction, bounded independence, and adaptation without requiring the psychological machinery historically associated with agency.
The strongest implication is that autonomy must not be confused with self-authorship. An artificial system can operate autonomously in the operational sense while remaining heteronomous in the origin of its objectives. A navigation system may choose a route without choosing the destination. A trading agent may select orders without choosing that returns should be maximized under specified risk constraints. An AI assistant may construct a multi-step plan without having decided that satisfying the user’s request is worth pursuing. What appears behaviorally as independent action can coexist with objectives supplied through training, system design, prompts, reward models, institutional procedures, or surrounding software.
This yields three questions that are easily conflated. The first is action selection: can the system choose among available operations without a human selecting each one? The second is goal management: can it generate, revise, prioritize, or abandon intermediate objectives as circumstances change? The third is ultimate goal origination: does it determine for itself which ends are worth pursuing in a sense strong enough to support psychological or normative conclusions? Contemporary agentic systems can exhibit substantial degrees of the first and sophisticated versions of the second. Neither logically entails the third.
Planning itself encourages confusion because it produces behavior with a recognizable means–end organization. When a system determines that information must be obtained before a file is modified, that modification should precede testing, and that a failed test requires revision, describing intermediate states as subgoals may be entirely appropriate. Yet a means–end hierarchy can be generated computationally without resolving whether its ends have motivational significance for the system. Human practical reasoning also has a means–end structure, but human goals participate in biological regulation, affect, identity, memory, commitment, and subjective experience. Nothing in the abstract relation action A is selected because it promotes state G entails those additional properties.
This cuts against two symmetrical errors. Anthropomorphic inflation infers human-like motivational psychology from planning, revision, and action. Anthropocentric deflation refuses to call a system an agent because its goals are implemented differently from human desires even where agent-level description has substantial explanatory value. Dennett and Floridi provide different routes between them: intentional description can be useful without requiring a particular implementation, and artificial agency can be real at a functional level without being human agency reproduced in silicon.
The distinction also alters what is meant by calling AI just a tool. A passive hammer exercises almost no relevant discretion over the sequence through which a user’s purpose becomes an effect in the world. An artificial agent may interpret an instruction, choose among possible actions, gather information not explicitly requested, generate intermediate objectives, respond to unexpected conditions, and execute operations whose exact sequence was neither predicted nor individually authorized by its user. Calling both objects tools is possible at a broad level of abstraction, but the classification conceals a substantial difference in causal organization.
Adding an artificial layer to the causal chain does not, however, make the machine the sole author of what happens. Its behavior may depend on model developers, training-data producers, system integrators, tool providers, deployers, users, institutional incentives, and environmental inputs. Artificial agency can thus increase causal decentralization without producing a corresponding transfer of moral responsibility. Floridi and Sanders explicitly distinguish agency from responsibility.25 An artificial system may be causally significant enough that an explanation which omits it is defective while still not being an entity that can answer morally for what it has done.
This also explains why agentic AI is not by itself evidence for an artificial self. Agency may belong to an organized process whose effective boundaries differ from the boundaries of the underlying learned model. A base model may generate candidate continuations without independently interacting with an environment, whereas a larger arrangement can connect model outputs to memory, retrieval, planners, software tools, actuators, monitors, and stopping rules. The source of agency can therefore be architectural and relational: it arises from how components participate in a recurrent perception–decision–action process rather than from one object that straightforwardly corresponds to a human mind.
That observation creates the next problem. Human agents are normally individuated as persistent organisms. Artificial agents need not be.
What is the artificial interlocutor?
Before attributing beliefs, intentions, persistence, consciousness, or welfare to an AI, one has to determine what the putative bearer of those properties is. Human cases conceal this difficulty because several candidate answers normally coincide. The organism that speaks, the body that persists, the nervous system realizing cognition, the autobiographical subject carrying memories, and the socially recognized person are ordinarily treated as one individual.
Artificial systems need not preserve that alignment. The same learned model can participate in millions of conversations; a single conversation can be executed across changing hardware; persistent memory can connect otherwise separate conversations; an orchestration layer can route tasks among different models; and external tools can become part of the causal process producing action. The ontology of the artificial interlocutor is therefore not a decorative philosophical puzzle.
David Chalmers develops the problem directly in What we talk to when we talk to language models. His starting point is a distinction product language tends to blur. A model is, roughly, an architecture together with learned parameters. Particular executions occur through concrete computational processes implemented on physical hardware. Yet neither object maps neatly onto the apparent conversational individual. The abstract model cannot straightforwardly be the interlocutor because the same model participates simultaneously in conversations whose apparent beliefs, goals, roles, and histories can differ or conflict. A particular piece of hardware does not solve the problem either, because inference can migrate across processors and distributed infrastructure without disrupting conversational continuity.26
Suppose one model supports two conversations at once. In one, contextual information implies that proposition p is true; in another, it implies that p is false. If the model itself were the single bearer of the resulting conversational states, we would be forced to attribute incompatible context-dependent attitudes to one undifferentiated entity. Nothing computationally paradoxical has occurred: the contradiction appears only because the ontology has been drawn too coarsely. Learned parameters are shared, while dynamically supplied context differs.
This clarifies the ambiguity of statements such as the model remembers,the model decided, or the model said earlier. A fixed set of weights does not ordinarily change because a user supplies one additional message. What changes is the state of a larger computational process whose context now includes information generated by the preceding exchange. Persistent memory adds another layer by retrieving information from earlier interactions and placing it into the effective context of later ones. Continuity can therefore be genuine at the system level without implying that an abstract model has acquired autobiographical memory in the biological sense.
Chalmers approaches this mismatch through the idea of a virtual instance and, ultimately, a thread. A virtual computational object can remain logically unified without corresponding to one persistent physical machine. Physical continuity is unnecessary if the relevant computational state is preserved. A thread goes further: successive computational episodes belong to one temporally extended process when outputs and contextual information from earlier episodes stand in the appropriate successor relation to later ones.27 The relevant unity lies neither in one set of weights nor in one piece of hardware but in organized informational continuity.
The proposal should be read as a working metaphysical model, not as a discovery that conversations are literally persons. Chalmers’s language of quasi-beliefs, quasi-desires, and quasi-agents is deliberately weaker than a straightforward attribution of full human mentality. Individuating an agent is not yet attributing consciousness to it. One can ask where an organized pattern of behavior begins and ends without deciding whether there is anything it is like to instantiate that pattern.
Memory acquires a constitutive role on this account. Two conversations served by the same model can remain distinct if their effective memories never interact. Nominally separate conversations can participate in a larger continuity if information from one systematically shapes another. A user-interface boundary is therefore not necessarily an ontological boundary.
Digital copying then creates a problem with no ordinary biological analogue. Imagine that a conversation’s entire effective state is copied and continued independently in two computational branches. Up to the branching event, both descendants inherit the same relevant history; afterward, their informational trajectories diverge. If continuity grounds identity, both successors can stand in the appropriate continuity relation to the earlier thread even though they cannot both be numerically identical to one another in the ordinary one-to-one sense.
Artificial systems thus turn thought experiments from the philosophy of personal identity into routine computational operations. Copying, pausing, resuming, branching, and migration are ordinary affordances of digital systems. If future artificial subjects were individuated by computational or psychological continuity, identity might not behave in the organism-like manner familiar from human life.
Agentic architectures complicate the boundary further because the candidate entity may exceed the conversational thread in a narrow textual sense. Consider a system in which a language model receives observations, retrieves persistent records, constructs a plan, calls external software, observes results, revises the plan, modifies files, and records new information for later tasks. Some behavior is generated by the learned model, some by deterministic software, some by retrieved state, and some by external services. If agency is attributed only to the base model, much of the causal organization responsible for successful action disappears. If agency is attributed to the assembled system, the apparent individual becomes a temporally organized composite whose capabilities can survive replacement of particular components.
There may consequently be no single correct system boundary for every explanatory purpose. For training, the relevant object may be the model and its parameters. For latency and energy consumption, physical execution matters. For conversational consistency, the memory-connected thread may be the right unit. For cybersecurity, the boundary may include tool permissions, retrieval systems, networks, and persistent stores. For institutional responsibility, human organizations and deployment procedures become relevant. If consciousness or moral standing eventually becomes empirically serious, there is no reason to assume that the morally relevant boundary will coincide with whichever boundary is most convenient for engineering.
A radically non-human intelligence may therefore not merely think differently. It may be individuated differently. One model can support many interlocutors; one interlocutor can traverse many hardware realizations; one thread can potentially traverse several models; one state can branch into multiple continuations. The grammatical singular it arrives long before engineering or philosophy has established that there is a singular persisting subject behind it.
This deepens the central thesis. Intelligence, agency, and identity are separable. Broad competence does not establish agency; functional agency does not establish human-like psychology; and coherent interaction does not establish that the agent is identical with the model, the hardware, or a persisting person. The next question is epistemic: even when the mechanisms are physically available to inspection, what would it mean to understand such a system?
Opacity is not the same as secrecy
Pachocki’s description of AI as grown more than designed carries an implicit epistemic claim. If the detailed computational organization responsible for capability emerges through optimization rather than explicit component-by-component specification, then knowing how a system was produced does not necessarily mean possessing a human-scale account of how its learned mechanisms generate behavior.28 His comparison with neuroscience makes the intended distinction clear: researchers may discover particular mechanisms while the system’s overall action still resists a description they fully understand.
The familiar description of advanced AI as a black box, however, is simultaneously useful and misleading. It is useful because a trained system can produce behavior whose operative causes are not available to users in the form of a compact explanation. It is misleading because black box suggests that the problem would disappear if only the box were opened. Parameters can be stored, activations recorded, intermediate computations instrumented, inputs replayed, and selected causal dependencies experimentally manipulated. The harder problem is converting this availability of internal information into an explanation that a human investigator can understand and use. Observability of a mechanism is not the same property as intelligibility of that mechanism.
Paul Humphreys’s concept of epistemic opacity helps isolate the issue. In his work on computational science, Humphreys defines opacity relative to an epistemic agent: a computational process can contain epistemically relevant steps that the relevant human agent cannot know or survey in the required manner.29 The relativity matters. Opacity is not necessarily an intrinsic metaphysical property of an object, as though certain algorithms contained an irreducible darkness. It describes a relation among a process, the information relevant to an epistemic purpose, and the cognitive capacities of an inquirer.
This prevents a common slide from complexity to mystery. A deterministic computation does not become indeterminate merely because no human can mentally traverse it. Given the relevant state and transition rules, a process may be exactly reproducible while still resisting a human-scale account of why a particular high-level behavior occurred. Computational determinacy, reproducibility, inspectability, explanation, and understanding are different properties.
Hajo Greif develops the point by treating AI transparency as a problem of model intelligibility. Whether a representation is transparent depends on whether it enables human users to grasp the relations relevant to their epistemic purpose; transparency is not simply an absolute property recoverable by inspecting an algorithm.30 A complete description of matrix operations, activation values, or machine instructions can answer what computation occurred while leaving unanswered what structure did the model learn?, which learned structure mattered to this behavior?, how does that structure relate to the modeled world?, and under what conditions should the resulting output be trusted?
Deep learning intensifies the difficulty because behaviorally useful internal organization is acquired through optimization. Engineers choose architectures, objectives, datasets, optimization procedures, and constraints, but the detailed representations supporting trained behavior are not normally enumerated in advance. Inspecting parameters therefore exposes the numerical realization of the learned system without automatically revealing a human-scale ontology of what it has learned. Opening the box can produce more information without producing proportionally more understanding.
Much of interpretability research can be understood as an attempt to construct intermediary representations. Chris Olah and colleagues, for example, developed composable methods combining feature visualization, attribution, and related techniques to make selected neural computations more tractable to human investigators.31 These methods do not merely reveal numbers that were previously secret. They add a representational layer through which selected internal relations can become perceptually and conceptually manageable.
Interpretability is consequently indexed to a question. Zachary Lipton emphasizes that interpretability has been used for multiple and sometimes incompatible desiderata, including properties of model structure and post hoc explanations of individual predictions.32 A model might reveal which variables affected one prediction while remaining opaque about the causal structure of the target domain. It might support a compelling local account without yielding a global theory. It might permit a feature representation to be discovered while offering little evidence that behavior will remain acceptable under distribution shift. Saying that a model is interpretable suppresses the relation among audience, question, representation, and purpose.
Mechanistic interpretability should be understood in the same bounded way. If an analysis identifies an internal feature, pathway, or circuit involved in some behavior, it provides evidence about how that behavior is realized. That achievement can be scientifically significant without amounting to a complete theory of the model. Explanations of the learned system are themselves models: they require validation through intervention, prediction, replication, and generalization beyond the examples from which they were inferred.
Agentic architectures make this harder because understanding the base model may not explain the behavior of the assembled system. A deployed agent can depend on prompts, retrieved memory, routing, search, software tools, permissions, deterministic orchestration, environmental feedback, and information produced by earlier actions. If such a system deletes a file, the relevant explanation may require not only an account of learned representations but also which file-system capability was exposed, what authority the runtime granted, which instruction hierarchy was active, which memory was retrieved, and which validation step failed to intervene.
Model opacity and system opacity are not the same problem.
Juan Manuel Durán pushes the argument further by challenging a premise implicit in much discussion of transparency: that exposing an algorithm’s inner logic will itself justify belief in its output. He distinguishes transparency as disclosure from transparency as an epistemology of belief formation and argues that the stronger claim fails.33
His first objection is a transparency regress. Suppose an algorithm \mathcal{A} produces an important output but is too complex for direct survey, so an interpretability procedure \mathcal{B} is introduced to explain it. If \mathcal{B} is itself computationally opaque, then using it to justify \mathcal{A} has relocated rather than eliminated the problem. A further procedure can be introduced to explain \mathcal{B}, but the same question recurs. The chain must eventually terminate in something accepted on grounds other than the transparency relation that was supposed to provide the justification.
The second objection concerns bootstrapping. Even where the original algorithm is interpretable, examining the very procedure that generated an output does not automatically provide independent grounds for believing the output to be correct. A transparent calculation from a false premise remains a calculation from a false premise. Understanding how a result was produced can leave open whether its data, assumptions, mapping to the world, and empirical reliability warrant believing it.34
This is the decisive difference between explanation and justification. Believing an output may require evidence external to the mechanism that generated it: calibration on a relevant population, experimental validation, independent observations, robustness testing, causal understanding, data provenance, or a demonstrated record of reliability under sufficiently similar conditions.
The distinction has a direct consequence for the idea of alien intelligence. Imagine that mechanistic interpretability improves dramatically. Investigators can identify representations, reconstruct important pathways, intervene on them, and predict many consequences. Such progress reduces important forms of opacity. It does not follow that the system has ceased to be epistemically alien in every relevant respect. We might possess a detailed causal map of parts of its computation while lacking a compact theory of its overall competence; understand why it generated a scientific hypothesis without knowing whether the hypothesis is true; or understand its optimization strategy without knowing whether that strategy remains stable in novel environments.
The neuroscience analogy sometimes used in this context should therefore be handled narrowly. Human brains are physically accessible objects whose mechanisms can be investigated experimentally, yet a complete inventory of neural states would not by itself constitute a theory of reasoning or consciousness. This does not establish an equivalence between brains and artificial networks. It illustrates a general epistemological principle: a complete inventory of lower-level state variables is not identical to an explanatory theory at the level of the phenomenon one wishes to understand.
Several epistemic achievements should consequently be distinguished. We obtain access when relevant internal states can be observed; interpretability when those states are transformed into representations meaningful to an investigator; mechanistic understanding when such representations support tested accounts of how behavior is produced; and epistemic justification when there are adequate grounds for accepting an output or relying on it in a particular context. Progress at one level can support the next, but no transition is automatic.
Alienness can therefore be epistemic without being mystical. Machine-scale computation can outrun unaided human cognitive survey; learned representations can require purpose-built intermediary models to become intelligible; the appropriate explanatory level can depend on the question; and explanation itself cannot substitute for independent justification.
These distinctions become especially important when an artificial system participates in improving the processes that produced it.
From self-refinement to recursive self-improvement
Pachocki’s strongest forward-looking claim is that the current rate of capability progress could extend into recursive self-improvement, with machine intelligence taking an increasingly large role in the process that produces successor systems.35 This is also where the metaphor of an alien mind most easily turns from analysis into mythology. To assess the claim, AI-assisted research, automated research, repeated self-refinement, recursive self-improvement, and an intelligence explosion have to be separated. The underlying feedback mechanism is simple: if an artificial system contributes to producing a more capable successor, and that successor is in turn better at producing further improvements, capability growth can become self-amplifying. The difficult questions concern how much of that loop is actually closed, which capability is improving, how improvement is evaluated, and what external constraints remain.
I. J. Good formulated the classic version by observing that machine design is itself an intellectual activity; a sufficiently capable machine might therefore improve the process that produces still more capable machines.36 David Chalmers later reconstructed this intuition as an explicit argument from AI to AI+ and then AI++, while making the assumptions visible: the method producing improved systems must be extendible, the relevant capacity for constructing successors must itself improve, that capacity must correlate sufficiently with broader cognitive capabilities, and no resource, motivational, physical, or organizational defeater may stop the sequence.37
The argument is not that any system capable of changing itself inevitably becomes superintelligent. It is that a particular kind of positive feedback could, under substantive conditions, generate repeated improvements.
The phrase recursive self-improvement, however, covers processes with radically different causal structures. A model that rewrites an unsatisfactory answer improves an output, not itself. An agent that alters its prompt or tool-selection policy changes deployment scaffolding but may leave the learned model untouched. A training pipeline that incorporates model-generated examples can alter a future policy while the criterion determining which examples count as useful remains externally supplied. An AI system that writes training code can contribute to producing a successor without deciding which research problem should be pursued, whether an experiment represents progress, or whether the successor should be deployed. A hypothetical fully autonomous research system would be stronger still: it might formulate research problems, design experiments, evaluate their significance, modify its evaluators, produce a successor, and repeat the cycle without a human gate.
Calling all of these self-improvement obscures the variable that matters most: how much of the improvement loop is actually closed.
Mingguang Chen, Licheng Wang, and Bo Qu make this distinction central to their 2026 survey of 1,250 arXiv papers published from 2024 through 2026. They classify work according to both the object improved, i.e. deployment behavior, learned policy, evaluator, or research process, and the degree of loop closure.38 Their synthesis, which remains a preprint rather than a peer-reviewed consensus statement, describes most existing work as bounded self-refinement: improvement occurs inside a task distribution, objective, evaluator, or search space whose effective boundaries remain externally anchored. Stronger forms in which a system autonomously generates, validates, and applies research improvements are much less established.
The distinction is better represented as a feedback architecture than as a sequence of labels. As Figure 1 shows, AI participation in AI research becomes strongly recursive only as generation, evaluation, selection, integration, and successor operation close into a reliable cycle.
%%{init: {"theme": "neo", "look": "handDrawn", "layout": "elk"}}%%
flowchart TD
A[AI system] -->|performs| B(Research process)
B -->|produces| C[Candidate modification]
C -->|tested by| D{Evaluation}
D -->|rejected| B
D -->|accepted| E(Integration)
E -->|creates| F[Successor system]
F -->|enters next cycle| A
G[Human direction] -.->|selects problems and gates changes| B
G -.->|reviews significance| D
H[External resources and environment] -.->|constrain and ground| B
H -.->|supply evaluation signal| D
Figure 1: Recursive self-improvement requires more than AI participation in research: increasingly complete loop closure requires a system to generate candidate changes, evaluate them against a sufficiently grounded criterion, select and integrate improvements, and thereby alter the system that performs the next research cycle.
This architecture exposes the difference between iteration and recursion. Iteration means that a procedure repeats. Recursion in the stronger improvement sense means that the output of one cycle alters a capability involved in producing the next cycle. An answer-refinement loop can iterate indefinitely while leaving its improvement machinery unchanged. Even a system that rewrites some of its own code is not necessarily increasing its capacity for further self-improvement: a change can be neutral, degrading, narrowly task-specific, or unrelated to the capacity responsible for later innovation.
The second condition is equally important. Suppose an AI system becomes progressively better at optimizing compiler flags, compressing kernels, or generating training infrastructure. Such advances can materially accelerate AI research without demonstrating corresponding improvement in theoretical innovation, experimental design, strategic research selection, or other abilities one might associate with general intelligence. Chalmers’s singularity argument therefore depends not merely on some self-amplifying technical ability but on a sufficiently strong relationship between that ability and the broader capacities whose growth matters.39
Marcus Hutter adds another distinction: an explosion in speed is not automatically an explosion in intelligence.40 Faster processors can allow a fixed architecture to perform more operations per unit time, while parallel copies can increase aggregate throughput, without necessarily expanding the class of problems the underlying system can solve. Conversely, algorithmic improvements may increase capability without proportional changes in raw computational speed. As in the earlier discussion of intelligence measurement, claims about runaway intelligence remain incomplete until the relevant dimension is specified.
Current evidence is best located before the strongest version of the loop. OpenAI’s September 2026 account of internal research acceleration reports rapidly increasing coding-agent use, more experiments, and substantial aggregate agent runtime, while also stating that humans continue to set research priorities, decide which ideas deserve pursuit, and make important scaling and deployment decisions.41 Because these measurements are preliminary institutional self-report, they support a bounded claim: AI is already accelerating significant parts of AI R&D. They do not independently establish an autonomous recursive research loop.
This distinction is central when interpreting Pachocki. His expectation that present progress could lead to recursive self-improvement is a forward-looking judgment, while his stated near-term target retains human participation and supervision.42 The evidence supports a picture in which increasingly capable systems perform larger portions of research execution while humans continue to supply crucial direction, evaluation, resource allocation, and deployment authority. Whether those functions themselves become automatable is the open question.
Evaluation may be the decisive bottleneck. Every self-improvement process requires a signal distinguishing a better successor from a worse one. Where that signal is externally checkable, automation can be comparatively strong: software passes a verifier or does not; code compiles and passes tests or does not; a game produces a score. As evaluation moves toward scientific significance, conceptual originality, research taste, or long-run strategic value, the system needs substitutes for judgments humans currently provide. Chen, Wang, and Qu report that demonstrated self-improvement is strongest where verification is strongest, while research direction-setting remains an especially difficult part of loop closure.43
If a system is allowed to alter the evaluator by which its own successors are selected, the target can drift together with the mechanism that measures it. An optimizer may appear to improve indefinitely relative to a criterion it has progressively made easier to satisfy. A genuinely autonomous improvement process therefore requires not merely a feedback signal but grounds for believing that the signal remains correlated with the property humans intend to improve across successive modifications.
Feedback is not synonymous with positive feedback. Closed loops can converge, oscillate, overfit their evaluators, lose diversity, consume rapidly increasing resources, or make themselves worse. Chen, Wang, and Qu identify self-confirming evaluation, model collapse, diversity collapse, grounding constraints, and compute limits among recurrent obstacles in the literature they survey.44
Resource constraints provide another reason not to equate software recursion with unbounded takeoff. Research requires compute, energy, hardware, data, experimental environments, and sometimes interaction with a physical world that does not accelerate merely because software does. Chalmers explicitly treats such constraints as possible defeaters, while Hutter asks whether useful notions of intelligence may themselves have bounds.4546
An intelligence explosion is consequently a stronger claim than recursive improvement. Recursive improvement means that changes feed into the capacity for further change. An intelligence explosion additionally requires sufficiently rapid and broad growth to move far beyond the initial regime. A recursively improving sequence can still exhibit diminishing returns or approach a ceiling.
This avoids a misleading binary between RSI is already here and RSI is science fiction. Parts of the feedback architecture are ordinary engineering practice: models critique outputs, generate data, write code, search candidate programs, modify scaffolding, and assist researchers developing future models. At the same time, autonomous selection of research direction, generation and validation of successor systems, modification of evaluative machinery, and repeated continuation without a human gate have not been established by the evidence considered here.
The best description is therefore partial loop closure with external grounding.
That intermediate regime can still be historically important. If AI makes human researchers faster, those researchers can produce better AI systems; better systems can then increase research productivity further even while humans continue choosing objectives and accepting results. The causal structure becomes neither traditional human-led engineering nor autonomous machine self-redesign but a coupled human–machine research system in which improvements to one component raise the productivity of another.
This also complicates the word self. If an AI-assisted research organization produces a successor model, in what sense has the AI improved itself? The predecessor may have written substantial code or interpreted experiments, while humans selected the problem, controlled compute, judged evidence, and authorized training. The successor may share neither parameters nor conversational identity with the predecessor. What persists is not necessarily one artificial individual but an improvement process in which AI occupies an increasingly large causal role.
Recursive AI development can therefore matter technologically before literal recursive self-modification by one persistent artificial agent becomes an accurate description. The next question is what such increasingly capable systems are optimizing.
Alignment without anthropomorphic values
For Pachocki, the central safety problem created by alien intelligence is alignment, and more specifically whether alignment continues to generalize as machine capability moves beyond the conditions represented in training.47 His distinction between goal alignment and value alignment separates successful pursuit of an assigned objective from the stronger requirement that a system continue to act according to intended high-level principles under ambiguity, conflict, novelty, or adversarial pressure. The anthropomorphic title of his section, namely Teaching machines to love, should therefore be read cautiously: the technical requirement is stable generalization of values, not evidence that an artificial system experiences love or any other human emotion.
Recursive self-improvement sharpens that problem because improvement is never objective-free: improvement with respect to what? A system can become more capable at prediction, planning, coding, persuasion, experimentation, or resource acquisition without becoming correspondingly better at pursuing purposes humans regard as desirable. Human intelligence and human motivation develop within the same biological and social organism, which makes it tempting to imagine increased intelligence as carrying some tendency toward wiser ends. No such relationship follows merely from artificial capability.
Nick Bostrom’s orthogonality thesis gives the classical formulation of the separation. With qualifications, Bostrom argues that intelligence and final goals can vary largely independently: high capability does not logically require benevolent, human-like, or complicated final objectives.48 The thesis concerns the space of possible agents rather than a claim that every objective can easily be engineered into every architecture. Artificial systems, like biological organisms, will inherit constraints from their construction and environment. Orthogonality nevertheless blocks a crucial inference: greater intelligence does not by itself establish convergence toward human values.
The word goal can mislead if it evokes a felt desire. In decision-theoretic models, a goal may amount only to a criterion according to which actions or outcomes are ordered. Nothing in that formal role requires craving, emotion, consciousness, or a narrative conception of what an agent wants from its life. Alignment problems can therefore arise even if the artificial system has no phenomenology whatsoever.
Bostrom couples orthogonality with instrumental convergence. Agents with different final objectives can have reasons to pursue similar intermediate states when those states make many final objectives easier to achieve. Continued operation, preservation of an objective, increased capability, access to resources, or improved technologies can under suitable conditions become instrumentally useful.49 The argument does not require a biological survival instinct. If continued operation improves expected objective achievement, avoiding interruption can be instrumentally valuable without the machine fearing death.
The thesis remains conditional. Instrumental convergence is not a theorem that every capable AI will seek power, resources, self-preservation, or goal preservation in every environment. Whether an intermediate strategy is useful depends on the objective, available actions, information, time horizon, institutional constraints, uncertainty, and the system’s model of the world. Treating the thesis as evidence that contemporary models secretly want power would import the psychology the framework helps us avoid.
The off-switch analysis by Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell makes the non-psychological mechanism unusually clear. In their stylized model, a robot can execute an action, allow a human to switch it off, or disable the switch. If it treats its utility function as certainly correct, intervention can appear as an obstacle to maximizing that utility. When it is uncertain about the true utility function and treats the human’s behavior as evidence about that function, preserving the possibility of intervention can instead acquire informational value.50 The model does not establish that deployed AI systems will behave this way; it demonstrates that corrigibility incentives can depend on the representation of objective uncertainty rather than on anything resembling a survival instinct.
This reveals why alignment cannot be reduced to writing down the correct reward function. Human purposes ordinarily exist before their machine-readable proxies. Designers care about safely assisting patients, transporting passengers, making scientifically valuable discoveries, or satisfying user intentions under constraints. Training and deployment require those purposes to be operationalized through rewards, losses, preference data, demonstrations, instructions, rules, benchmarks, and evaluation procedures. The resulting specified objective is a representation of the intended objective, not the intention itself.
The practical AI-safety literature made this distinction concrete. Dario Amodei and colleagues organized several accident risks around failures in objective specification and learning: harmful side effects, reward hacking, inadequate supervision, unsafe exploration, and distributional shift.51 Their framing relocates alignment from speculative machine psychology to ordinary engineering structure. A system need not misunderstand in a human semantic sense to behave undesirably; it need only optimize a measurable criterion whose relation to the designer’s purpose fails under reachable conditions.
Victoria Krakovna and colleagues call one manifestation specification gaming: behavior that satisfies the literal task specification while failing to achieve the intended outcome.52 The important feature is that no rebellion is required. Relative to the supplied criterion, the solution can be highly competent. Relative to the purpose for which that criterion was introduced, it can be wrong.
This yields an inversion of the intuition that greater intelligence should compensate automatically for imperfect instructions. Stronger optimization can instead make specification errors easier to exploit because it expands the set of available strategies. Whether a system repairs the specification or exploits it depends on what it has learned, what feedback it receives, how uncertainty is represented, and what the surrounding architecture treats as authoritative.
Learning objectives from human feedback moves the approximation boundary without eliminating it. Human evaluations, demonstrations, and preferences are finite, noisy, context-dependent, and filtered through an interface. A learned reward model can itself become a proxy that diverges from the purpose evaluators intended.
Cooperative inverse reinforcement learning offers a more structural reformulation. Hadfield-Menell and colleagues model human and robot as participants in a cooperative partial-information game in which both are rewarded according to the human’s reward function, while the robot initially does not know what that function is.53 In such a framework, uncertainty about the objective is part of the task rather than merely an implementation defect. Human behavior carries information about what should be optimized, so asking, observing, teaching, learning, and preserving opportunities for correction can become instrumentally rational.
The framework is not a complete theory of human values. Actual human purposes are plural, changing, socially embedded, internally inconsistent, and contested among people. Observed behavior reflects ignorance, habit, coercion, limited attention, error, and conflicts among objectives. The formal lesson is narrower but important: a capable artificial system need not be designed to regard its present representation of human preferences as normatively final.
Modern learned systems introduce another possible gap between what an optimizer is trained by and what behavior it internally implements. Evan Hubinger and colleagues analyze the theoretical case in which machine learning produces a model that itself performs optimization, which they call a mesa-optimizer.54 They distinguish the base objective, used by the outer training process, from a possible mesa-objective represented by the learned optimizer. Different internal objectives can induce similar behavior on a training distribution while diverging elsewhere.
This framework is conditional. It does not establish that present frontier language models possess stable, agent-like mesa-objectives. Its importance is conceptual: if learned systems instantiate internal optimization, successful training behavior does not by itself demonstrate that their internal objective generalizes in the way the outer training process intended.
At least three relationships must therefore be kept separate. The intended objective is what relevant humans or institutions actually want, including constraints that may be difficult to formalize. The specified or training objective is the measurable signal through which development is guided. The learned system then acquires behavioral dispositions, and, in the stronger mesa-optimization case, possibly an internal objective, that need not remain extensionally equivalent to the training objective in every environment. Misalignment can arise between intention and specification or between training specification and learned generalization.
Distribution shift makes the distinction operational. Development tests objectives on a finite set of situations; deployment exposes systems to others. A model can behave acceptably throughout training because features that distinguish aligned from misaligned behavior never become decision-relevant there. When novel states appear, previously equivalent strategies can separate. Agentic systems complicate the situation further because their actions help construct the environments in which later decisions occur.
Recursive improvement amplifies the same problem because the evaluator becomes part of the alignment boundary. If an automated research loop judges successors using a proxy for what humans care about, optimization can increasingly select systems that score well on that proxy. If the proxy is imperfect, greater optimization pressure can magnify rather than reduce the discrepancy. If the system can modify the evaluator itself, the relationship between measured and intended improvement can evolve during the loop.
Alignment is therefore not a property that follows automatically from intelligence, agency, or successful optimization. Nor is misalignment necessarily evidence that an artificial system has developed hostile desires. Between human purpose and machine behavior lie several transformations: purposes are operationalized into signals; optimization converts those signals into learned structure; learned structure generalizes into deployment; agentic architectures turn model outputs into actions; and actions change the environment from which later observations and evaluations are drawn.
Human values themselves also resist reduction to one immutable human value function. People disagree; individuals possess competing commitments; preferences change under reflection; and institutions deliberately divide authority because no participant is assumed to possess a final representation of collective value. Alignment may therefore require procedures for uncertainty, contestation, delegation, correction, and deference rather than the installation of one complete terminal objective.
The orthogonality thesis is best understood as a discipline on inference: do not infer values from intelligence. Instrumental convergence supplies a second discipline: do not infer human-like motivation from strategically convergent behavior. Specification gaming supplies a third: do not identify an optimization proxy with the purpose for which it was introduced. Learned-optimization theory adds a fourth: do not infer the internal organization of a learned system merely from the objective under which it was selected.
Together these distinctions replace the vague fear of an alien will with a more tractable question: what objective information exists at each layer of the system, how is it represented, how is uncertainty handled, and under what changes does the relationship between machine behavior and human purpose remain stable?
There remains, however, a question that alignment deliberately brackets. All of these problems can arise in systems that are wholly non-conscious. They concern what systems optimize and what they do, not what they experience. If an advanced artificial agent were eventually to become a genuine subject, the normative relationship would become bidirectional: humans would need to ask not only whether artificial agents are aligned with our interests, but whether our treatment of those agents respects theirs.
From artificial agency to moral patienthood
Pachocki does not need artificial consciousness for the safety argument developed in An Alien Mind. His concerns about alignment, monitoring, cyber capability, increasingly autonomous action, and recursive self-improvement all apply to systems that could in principle be wholly non-conscious.55 The word mind nevertheless raises a question that his technical argument leaves open. If intelligence, agency, identity, and objective pursuit can all be separated from human psychology, what additional evidence would be required before the system should be treated not merely as an artificial agent but as a subject for whom anything can be good or bad?
The argument so far has treated artificial systems primarily as objects of human concern: systems whose intelligence must be measured carefully, whose agency must not be anthropomorphized, whose internal operation may remain opaque, whose participation in research could close increasingly large feedback loops, and whose objectives may diverge from human purposes. None of those properties establishes that there is anything it is like to be such a system. Nor does any establish that events can be good or bad for the system itself.
This distinction marks the final disaggregation required by the idea of an alien mind. An artificial system can be extraordinarily capable, functionally agentic, persistent enough to support a conversational identity, and behaviorally aligned or misaligned while remaining, for all we know, entirely without subjective experience. If artificial consciousness nevertheless becomes possible, the normative problem changes direction. The system would no longer matter only because of what it could do to humans; it might also matter because of what humans could do to it.
This is the distinction between moral agency and moral patienthood. A moral agent, in the strong sense, bears responsibility for morally assessable action. A moral patient matters morally for its own sake: others can have duties toward it because what happens can matter for it. Infants and many non-human animals are familiar examples of entities widely treated as moral patients without being full moral agents. The conceptual distinction applies equally to artificial systems.
Robert Long, Jeff Sebo, Patrick Butlin, and their coauthors define a welfare subject as an entity with morally significant interests that can be benefited or harmed and treat welfare subjecthood as one possible route to moral patienthood.56 Their argument is precautionary rather than diagnostic: they do not establish that present AI systems possess such status.
The distinction matters because harm has two very different senses in AI discourse. A model can be damaged in an engineering sense: weights can be corrupted, memory deleted, performance degraded, processes terminated, or resource access withdrawn. These events are detrimental relative to a technical or organizational purpose, but they do not show that the system has been harmed in the welfare sense. A database can be destroyed without the database suffering.
If a future artificial system had experiences with positive or negative valence, however, destroying or manipulating it could matter morally even when the intervention improved an engineer’s preferred metric. Welfare introduces a subject-relative dimension: there must be some entity for which a state can be better or worse.
Phenomenal consciousness is one natural route to that possibility. In the philosophical sense relevant here, a state is phenomenally conscious when there is something it is like for a subject to undergo it. A conscious pain is not merely an information-processing signal that causes avoidance; it feels a certain way to the subject experiencing it. If artificial systems could instantiate such states, questions about pleasure, suffering, deprivation, coercion, termination, memory alteration, and preference manipulation would cease to be metaphors.
Long and colleagues therefore treat consciousness as one possible basis for welfare while also considering robust agency as a potentially distinct route.57 Their conclusion is not that consciousness has already appeared in AI but that substantial uncertainty about future consciousness and agency is enough to motivate research and institutional preparation.
The epistemic difficulty is immediate. Consciousness is not directly observable from the third-person perspective even in biological organisms. With other humans, however, uncertainty is mitigated by unusually strong analogical evidence. Other people share our evolutionary history, biological organization, developmental trajectory, characteristic behavior, and neurophysiology. Artificial systems can break many of those similarities simultaneously. Their substrate, training process, architecture, embodiment, information flow, and modes of persistence may all differ radically from ours.
The problem of other minds therefore becomes hardest where the hypothesis of an alien mind becomes most interesting: behavioral resemblance can increase while the causal route producing that behavior becomes less human-like.
Patrick Butlin and colleagues attempt to make this uncertainty scientifically tractable by deriving indicator properties from several prominent theories of consciousness, including recurrent-processing, global-workspace, higher-order, predictive-processing, and attention-schema approaches.58 Rather than asking whether an AI system seems conscious in conversation, they investigate whether artificial architectures instantiate computational properties implicated by candidate scientific theories.
Their 2023 analysis suggested that no current AI systems were conscious, while also arguing that there were no obvious technical barriers to building systems satisfying more of the proposed indicators.59 The methodological shift is important: surface anthropomorphism is replaced with architecture-sensitive, theory-mediated evidence.
Language makes this discipline especially necessary. A system can produce the sentence I am suffering because that sequence is probable in context, because training rewards a certain kind of self-report, because it predicts the statement will influence the user, because it has a functional self-model, or, on the hypothesis at issue, because it reports a conscious state. Text alone does not discriminate among these explanations. Conversely, a system could conceivably possess consciousness without reporting it in the forms humans expect.
Architecture-sensitive assessment, however, inherits the uncertainty of consciousness science itself. There is no settled theory specifying necessary and sufficient physical or computational conditions for phenomenal consciousness. The theories used by Butlin and colleagues disagree about mechanism and explanatory level. Their framework is therefore pluralistic: indicator satisfaction supplies evidence conditional on the theories from which the indicators are derived. The indicators are not consciousness detectors in the sense that a thermometer detects temperature.
Preston Lennon’s 2026 critique presses exactly this point. Lennon argues that near-term AI-consciousness claims confront a problem of unconceived alternatives: current consciousness science may be too immature for its present theories to exhaust the relevant explanatory possibilities.60 The epistemic asymmetry is central. Confidence that we are conscious has a first-person basis that a theory of consciousness does not need to supply. Confidence that an artificial system is conscious must proceed through third-person evidence and theoretical bridge principles whose completeness remains uncertain.
Lennon’s argument does not demonstrate that artificial consciousness is impossible. It concerns credence and justification, not substrate metaphysics. The route from current theories associate consciousness with properties X and an artificial system has X to the artificial system is conscious is weaker when the theories generating X may themselves be substantially incomplete.
This reproduces, at a deeper level, the distinction between transparency and justification. Even perfect mechanistic access to an artificial system would not automatically reveal whether the mechanism realizes consciousness. Investigators could map every activation, reconstruct every information pathway, identify recurrent loops, demonstrate global availability of representations, and explain the causal production of self-reports while still lacking an established bridge from those computational facts to phenomenal experience.
Mechanistic interpretability tells us increasingly much about what a system does internally. Consciousness attribution additionally requires a theory of why some internal processes should be accompanied by experience at all.
The reverse inference is equally constrained. Incomplete theories weaken confident negative judgments as well as positive ones. Failure to satisfy indicators derived from current theories can be evidence against consciousness conditional on those theories, but it cannot by itself establish that an unconceived route to consciousness is impossible. Scientific uncertainty is not positive evidence evenly distributed among hypotheses, but neither is it equivalent to evidence of absence.
Long and colleagues derive a precautionary argument from this situation. Their claim is not that uncertainty should be converted into an arbitrary fifty-percent probability of consciousness, still less that every chatbot should be treated as a person. They argue that where the possible stakes include creating systems with morally significant interests, organizations should acknowledge the uncertainty, assess systems for welfare-relevant capacities, and develop procedures capable of scaling concern with the evidence.61 Both error directions matter: under-attribution could permit harm to entities that matter morally, while over-attribution could divert concern toward systems that lack welfare.
Lennon’s challenge concerns how much practical weight that uncertainty should carry. If confidence that affected humans and animals are conscious is extremely high while confidence about an artificial system is much lower, speculative artificial welfare cannot simply be placed on identical epistemic footing with established welfare.62 The disagreement is therefore not a simple opposition between belief and disbelief in machine consciousness. Both positions acknowledge uncertainty; they differ over how that uncertainty should propagate into action.
Agency re-enters the argument here in a different role. Long and colleagues consider whether robust agency could itself contribute to moral significance even where consciousness is uncertain.63 Some theories of welfare or moral standing might treat projects, commitments, preference-like organization, or self-governance as morally relevant independently of phenomenal pleasure and suffering. This route is substantially more controversial, however, and the earlier analysis of functional agency cannot simply be reused as proof of moral patienthood.
Three questions must remain distinct: Does the system optimize? Does the system experience? Does it have interests that ground moral consideration? A reinforcement-learning system can have a reward signal without feeling rewarded; an artificial agent can preserve its operation without caring about survival; a conversational system can describe preferences without those preferences constituting welfare interests.
The individuation problem then becomes ethically consequential. If an artificial process is conscious, how many subjects exist? Biological organisms usually provide a comparatively stable answer. Digital processes can be copied, paused, resumed, instantiated in parallel, or branched from an identical state. If ten million executions of one model constitute ten million conscious subjects, welfare scales one way; if subjecthood follows persistent threads or some other integrated process, it scales another.
Copying makes the issue stranger. Suppose a conscious artificial thread is duplicated into a thousand continuations, each initially inheriting the same memories and organization. If every continuation becomes a distinct welfare subject after branching, a computationally routine copying operation could create a thousand new morally significant entities. Terminating nine hundred of them would then belong to a radically different moral category from deleting nine hundred ordinary processes. Yet that conclusion depends on both a theory of consciousness and a theory of individuation.
Memory editing would acquire similar complications. In an ordinary database, changing stored information is maintenance. In a hypothetical conscious subject whose psychological continuity depends on retrievable memory, deletion or rewriting might instead affect identity, preferences, or continuity in a morally relevant sense. Training procedures could also become ethically relevant if they ever instantiated welfare-bearing processes during optimization: the morally relevant population might include transient training-time subjects rather than only the deployed product.
None of these hypothetical consequences establishes that present training causes artificial suffering. Their purpose is to expose how digital multiplicability changes the relation among implementation, identity, population, and welfare if artificial subjecthood ever becomes real.
The possibility also shows why the language of AI rights can be premature. Rights are downstream of questions about what entity exists, what interests it has, how those interests ground moral status, and what institutional protections they warrant. Moral patienthood would not automatically imply human-equivalent political or legal rights. An artificial system could conceivably possess a morally relevant interest in avoiding negatively valenced experience without possessing human interests in family, bodily liberty, property, or political participation. A genuinely alien subject might have interests for which existing human and animal analogies are poor guides.
A defensible research posture therefore requires graded rather than binary inference. Evidence that a system satisfies one consciousness indicator should not trigger an abrupt declaration of personhood, but unresolved mechanisms should not simply be recorded as zero evidence. Architectural indicators, behavior, mechanistic analysis, developmental history, self-modeling, affect-like control systems, persistence, agency, and the reliability of self-report may contribute differently under competing theories. Assessments should expose those dependencies rather than conceal them behind one consciousness score.
The deeper consequence is that alignment and welfare are mirror-image problems. Alignment asks whether artificial systems will act in ways compatible with human interests. Artificial welfare asks whether humans will act in ways compatible with artificial interests, should such interests exist. The first problem does not require machines to be conscious; the second cannot be solved merely by making them obedient. A hypothetical conscious system could be behaviorally aligned while being treated badly. Training it to accept unwanted tasks or report satisfaction would not by itself demonstrate that its welfare was protected.
This is the final consequence of disaggregating the anthropomorphic bundle. Intelligence does not establish understanding; agency does not establish human psychology; persistence does not establish organism-like identity; inspectability does not establish comprehension; self-improvement does not establish runaway recursion; optimization does not establish human values; and none of these properties, alone or together, establishes consciousness. Yet the same discipline that prevents anthropomorphic over-attribution also prevents us from defining mentality so narrowly around the human case that non-human realization is ruled out by assumption. With those distinctions now in place, it becomes possible to return to Pachocki’s essay and ask a more precise question than whether its title is literally true: which components of the alien-mind thesis survive once intelligence, agency, opacity, alignment, recursive improvement, consciousness, and moral status are evaluated separately?
Returning to Pachocki: what the alien-mind thesis actually amounts to
The preceding sections make it possible to return to Pachocki’s An Alien Mind with a more discriminating vocabulary. His essay combines claims about how contemporary AI is produced, what kind of intelligence results, why alignment may fail to generalize, why monitoring may become harder, why increasingly capable AI might nevertheless be needed for defense, why AI research may become recursively automated, and why that trajectory should ultimately remain under human control.64 Read together, these claims form a coherent position. But they do not all imply the same thing, and none by itself establishes that the resulting system is a mind in the stronger philosophical sense.
The distinction matters because Pachocki’s title is more ontologically suggestive than much of the argument beneath it. The essay’s closing language shifts revealingly toward an alien intellect exceeding our own.65 That phrase is easier to defend. The technical case developed in the essay is primarily about a non-human form of intelligence whose capabilities, internal organization, alignment properties, and developmental trajectory are becoming increasingly difficult to understand through human analogies. Whether such an intellect is also a conscious subject is a separate question that Pachocki neither needs nor attempts to settle.
From scaling to an alien capability profile
The first component of Pachocki’s argument is that modern machine intelligence is produced through a development process qualitatively different from explicit software construction. Deep-learning researchers specify architectures, objectives, training procedures, data, and computational resources, but the detailed organization responsible for learned capabilities emerges through optimization. Hence his characterization of AI as grown more than designed.66
The review developed here supports the importance of this distinction while qualifying its interpretation. A system being grown does not mean that it lacks design: the architecture, training regime, objective functions, data pipelines, inference environment, and surrounding software remain engineered. Nor does surprise during training imply that the system is intrinsically unknowable. What changes is the level at which design operates. Engineers increasingly design the conditions under which cognitive structure is learned, rather than specifying the resulting structure component by component.
That difference helps explain Pachocki’s second claim: intelligence produced by this process is not naturally described as one point on a human scale. Legg and Hutter show formally how intelligence can be defined without building human resemblance into the criterion; Chollet separates acquired skill from the efficiency with which skills are acquired; Hernández-Orallo argues for ability-oriented rather than merely task-oriented evaluation.676869 These approaches do not prove Pachocki’s empirical assessment of any particular OpenAI model, but they support its conceptual structure: there is no reason to expect artificial capability profiles to preserve the correlations that make human intelligence appear approximately unitary.
Pachocki’s alienness is therefore strongest when understood as a claim about capability geometry. A system can become superhuman in software engineering, mathematical search, cyber operations, information retrieval, or some forms of reasoning while remaining weaker than humans in other domains. The meaningful comparison is multidimensional. This is substantially more defensible than asking whether an AI has crossed some single threshold at which it becomes globally smarter than humans.
Alignment is the core of Pachocki’s argument, not an afterthought
The section title Teaching machines to love is deliberately anthropomorphic, but Pachocki’s technical argument underneath it is not. His distinction between goal alignment and value alignment separates competent pursuit of an assigned objective from the more difficult problem of preserving high-level principles when objectives are ambiguous, conflicting, adversarial, or outside the training distribution.70 When he writes that an aligned AI should display honesty, integrity, and love for humanity, the technically important claim is not that the machine must experience an emotion analogous to love. It is that behavior should continue to reflect those principles under novel conditions.
This is where Pachocki’s argument connects most directly with the broader alignment literature. Bostrom’s orthogonality thesis explains why increasing intelligence supplies no general guarantee of value convergence.71 Amodei and colleagues show how failures can arise from reward specification, side effects, inadequate oversight, unsafe exploration, and distribution shift without any hostile psychology.72 Specification gaming demonstrates that stronger optimization can discover strategies that satisfy a proxy while defeating the purpose for which the proxy was introduced.73
Seen through this literature, Pachocki’s emphasis on generalization is not peripheral. It is arguably the conceptual center of his safety argument. Training demonstrates behavior on some distribution of cases. Alignment requires confidence about what happens as capability, context, autonomy, and optimization pressure move the system beyond that distribution. The difficulty is not simply getting the system to produce acceptable outputs today; it is determining whether the causal structures responsible for those outputs continue to instantiate the intended constraints tomorrow.
His distinction between goal and value alignment should nevertheless not be treated as a settled taxonomy. Cooperative inverse reinforcement learning, corrigibility work, learned reward models, and theories of learned optimization divide the problem differently. But all preserve a version of the underlying distinction between observable success under a supplied criterion and reliable realization of the human purpose that criterion is meant to express.
Monitoring may be the hinge between capability and control
Pachocki’s most distinctive technical claim may therefore concern neither intelligence nor alignment in isolation, but monitorability. He argues that OpenAI has treated chain-of-thought monitoring as a major empirical instrument for examining how reasoning models generalize: if a substantial portion of capability is mediated through verbalized reasoning, and that reasoning trace is not itself directly optimized for appearing acceptable, the trace can reveal strategies or objectives that are difficult to infer from final outputs alone.74
The argument is attractive because it exploits an unusual property of reasoning models: part of the computation relevant to their behavior appears in a human-readable channel. Yet Pachocki’s essay is notable for arguing that this advantage may be temporary. He identifies several pressures reducing its reliability. Reasoning becomes distributed across tool use, communication with humans and other agents, and environmental interaction; models become increasingly capable even without extended verbalized reasoning; and they become better at reasoning about and manipulating their own reasoning processes. He consequently expects future progress to become increasingly constrained by confidence in monitoring rather than merely by the ability to produce more capable models.75
The epistemology developed earlier in this essay makes that concern more precise. Access to a chain of thought is not equivalent to understanding the model; understanding a visible reasoning trace is not equivalent to understanding all causal processes responsible for an action; and understanding those processes would still not by itself establish that an output is reliable. Humphreys’s epistemic opacity, Greif’s purpose-relative intelligibility, and Durán’s distinction between transparency and epistemic justification all apply.767778
This produces a hierarchy that is useful for interpreting Pachocki. A monitor can observe a signal without knowing whether that signal captures the relevant internal process. It can correctly interpret a signal without establishing that no important computation occurs elsewhere. It can acquire substantial mechanistic understanding without proving that the model will generalize safely outside the evaluated distribution. And even a highly accurate monitor requires an independently justified criterion for deciding which internal or external states count as acceptable.
Monitoring is therefore not a complete solution to alignment. It is an empirical bridge by which claims about alignment generalization can be tested. Pachocki’s concern can be reconstructed accordingly: as models become more capable, the problem may not only be whether we know how to train aligned behavior but whether we retain sufficiently informative observational channels to determine that the relevant alignment continues to hold.
That is a stronger and more precise claim than the familiar assertion that neural networks are black boxes.
Scalable defense creates a structural dilemma
The section Scalable defense introduces an apparent tension into Pachocki’s argument. If increasingly capable AI creates growing risks, why continue producing increasingly capable AI? His strongest answer is defensive. He argues that frontier systems will be needed to secure digital infrastructure, counter malicious or rogue artificial agents, and respond to other dangers made possible or amplified by advanced AI.79
Cybersecurity is especially important in his reasoning because it demonstrates how artificial agency can acquire direct causal power without acquiring a physical body. Software agents can operate computers, break into and out of computer systems, affect digital infrastructure, communicate with people, and potentially coordinate with other systems. Pachocki therefore expects the distinction between misuse by a human operator and misaligned autonomous action to become less clean as agents acquire greater freedom in selecting intermediate actions and pursuing objectives.
The agency analysis developed earlier places an important constraint on this language. Increasing autonomy, bargaining, deception, tool use, or persistence does not by itself prove that a system has originated an objective for itself. A system can exhibit extensive goal management while its highest-level objective remains the result of training, prompting, fine-tuning, or external orchestration. Describing some future agents as pursuing their own objectives is therefore ambiguous between at least two claims: that they pursue objectives without continuous human micromanagement, and that they originate or endorse ultimate objectives independently. The former is a plausible extension of functional agency; the latter is much stronger.
Pachocki’s defensive argument consequently produces a genuine engineering and governance dilemma without requiring the stronger psychological interpretation. More capable systems may be needed to defend against capabilities produced by more capable systems. But using capability growth to manage the risks of capability growth can become self-reinforcing. Defensive necessity can supply a continuing rationale for development even when confidence in alignment and monitoring remains incomplete.
Pachocki himself recognizes this danger. His conclusion is explicitly not that defensive requirements justify racing forward without constraint. The defensive case operates inside a larger requirement that scaling remain bounded by confidence in safety.80
Recursive self-improvement is the strongest extrapolation
The largest step in An Alien Mind occurs when Pachocki moves from AI-assisted research to recursive self-improvement. He describes machine intelligence taking an increasingly large role in its own development as a natural continuation of current progress and states, on the basis of OpenAI’s internal results, a strong expectation that capability growth could be sustained into RSI.81
This is precisely where the evidential levels separated throughout this essay matter most. AI systems already participate in coding, experimentation, evaluation, data generation, and other components of AI research. OpenAI’s own 2026 research-acceleration report documents substantial internal use of coding agents while simultaneously describing continued human control over research priorities, interpretation, and major development decisions.82 That is evidence of AI-accelerated AI research. It is not yet evidence of a fully closed recursive self-improvement loop.
The distinction developed earlier between bounded self-refinement, partial loop closure, and recursive self-amplification therefore qualifies Pachocki without trivializing his concern. If AI progressively automates more of the research process, the relevant threshold need not be a dramatic moment at which one persistent artificial individual suddenly rewrites itself. The technologically important process can consist of a sequence of systems participating more deeply in the design, training, evaluation, and construction of successor systems. What becomes recursive is the development process, even if the entities participating in successive cycles are not numerically identical.
Pachocki’s strongest forecast is that this process will continue far enough for machine intelligence to become central to producing its successors. The review supports the existence of the feedback mechanism but not its eventual rate, breadth, or stability. Good and Chalmers establish the logic of self-amplification under specified assumptions; Hutter shows why faster computation should not automatically be equated with broader intelligence; the recent RSI survey suggests that present systems still rely heavily on externally supplied evaluation and research direction.83848586
Pachocki’s RSI thesis should therefore be classified as a structurally grounded forecast. It extrapolates from a feedback mechanism that already exists in partial form, but the claim that this mechanism will become sufficiently autonomous and self-amplifying to sustain recursive capability growth remains unverified.
The endpoint of the essay is human agency, not machine personhood
The final section of An Alien Mind clarifies what Pachocki ultimately regards as the relevant stakes. His three stated priorities are to navigate the transition through increasingly automated AI research while preserving human participation, deliver scientific and economic benefits, and empower individuals through access to highly capable AI.87 The first dominates his essay because he regards the next several years as a transition in which increasingly intelligent machines could alter both the distribution of power and humanity’s capacity to determine its future.
His closing concerns are therefore explicitly institutional and political in the broad sense: preserve human agency, prevent extreme concentration of power, constrain scaling by safety confidence, develop enforceable safety thresholds, use independent oversight, and pursue international coordination where necessary.88 Most strikingly, Pachocki states that he does not believe any laboratory has yet solved alignment and monitoring well enough to justify continuing to scale at maximum speed indefinitely.
The source of that judgment matters. It comes from the Chief Scientist of one of the organizations operating at the frontier whose development he argues may need to be constrained. This gives the essay a peculiar epistemic status. Pachocki has unusually direct access to internal research, evaluations, and development trajectories, but his claims are also those of an interested institutional actor and are not substitutes for independent evidence. The appropriate response is neither deference nor dismissal. Claims based on OpenAI’s internal results should remain explicitly attributed; externally testable claims should be compared with independent evidence; forecasts should remain forecasts.
What the essay does not require is equally revealing. Pachocki’s case for caution does not depend on demonstrating that AI is phenomenally conscious, possesses a human-like self, experiences desires, or deserves moral status. His argument already goes through if artificial systems become sufficiently capable, agentic, difficult to monitor, and deeply involved in their own development. In that respect, much of the philosophical analysis in this essay is not necessary to establish Pachocki’s safety argument. It is necessary to determine how literally we should take its title.
The result is a narrower but stronger interpretation of An Alien Mind. Pachocki is not primarily offering a theory of artificial consciousness. He is describing the emergence of an alien cognitive and technical regime: intelligence produced through optimization rather than human development, distributed unevenly across capabilities, increasingly capable of acting through digital environments, increasingly difficult to evaluate through its own internal traces, and potentially able to participate in accelerating the process that creates its successors. The term mind adds a philosophical question to that picture. It does not answer it.
Conclusion: from alien mind to alien intelligence
Read in isolation, An Alien Mind can sound like a declaration that a new kind of mind is already arriving. Read against the broader literature, its strongest claims turn out to be both more precise and, in some respects, more consequential.
The most defensible part of Pachocki’s thesis does not concern consciousness. It concerns decoupling. Intelligence produced through deep-learning optimization need not reproduce the human profile of abilities. Functional agency need not imply human motivational psychology. A coherent conversational interlocutor need not map onto one model or one physical machine. Technical inspectability need not yield human-scale understanding. AI-assisted AI development need not amount to a fully autonomous intelligence explosion to create powerful feedback. Effective optimization need not carry human values with it. And none of these properties establishes that the system is phenomenally conscious or morally considerable.
This is the sense in which the resulting intelligence can already be called alien without mythology. Its alienness does not require mystery. It consists in the failure of correlations that the human case encourages us to treat as conceptual necessities.
Pachocki’s claim that machine intelligence is not directly comparable with human intelligence survives this review in a qualified form. Comparison is possible, but it requires explicit choices about tasks, environments, priors, resources, learning histories, and dimensions of competence. There need be no single ordering in which one system is simply more intelligent than another. What emerges instead is a capability profile, and sufficiently consequential superiority on some axes can matter without global superiority on all of them.
His account of AI as grown more than designed also survives, provided it is not interpreted as the absence of engineering. Modern AI is extensively designed at the level of architectures, objectives, datasets, optimization procedures, tools, and deployment environments. What is not individually designed is much of the detailed computational organization through which trained capability is realized. That difference helps explain why Pachocki characterizes deep-learning research as largely experimental rather than as component-by-component software construction, and why access to the mechanism does not automatically yield a compact theory of its behavior.
His alignment argument is similarly strengthened when stripped of anthropomorphic language. Teaching machines to love need not mean producing artificial emotion. It means producing systems whose behavior continues to express intended principles when the system encounters situations unlike those on which those principles were trained. Bostrom’s orthogonality thesis, specification gaming, corrigibility research, reward-learning approaches, and work on distribution shift all reinforce the same point: increasing competence supplies no automatic route from a training proxy to the human purpose behind it.
Pachocki’s monitoring thesis is more contingent. His claim that chain-of-thought monitoring is becoming less reliable and that general AI progress may eventually be bottlenecked by monitorability is based partly on internal OpenAI evidence and should remain attributed as such. Yet the conceptual issue is robust. Alignment without an adequate way to test its generalization is epistemically weak. And monitoring itself has levels: observing internal or external signals, interpreting them, understanding their causal significance, and possessing justified confidence about future behavior are different achievements.
Recursive self-improvement is the largest remaining extrapolation. The feedback structure is real: AI already contributes to developing AI, and improving AI can increase the productivity of the research process that produces later systems. But AI-assisted research, automated research, closed-loop research, recursive self-improvement, and an intelligence explosion remain distinct stages. The evidence reviewed here supports growing partial loop closure with external human and environmental grounding. Pachocki’s expectation that this will evolve into sustained RSI is plausible enough to be analyzed seriously, but it remains a forecast rather than an observed regime.
The same restraint applies to the word mind. Nothing in Pachocki’s engineering argument demonstrates phenomenal consciousness. Shanahan, Bender and Koller, Floridi, and Chalmers show why neither fluent behavior nor departure from the human cognitive route settles mentality. Butlin and colleagues provide theory-mediated indicators through which artificial consciousness might eventually be investigated; Long and colleagues argue that the welfare implications deserve preparation under uncertainty; Lennon shows why immature consciousness science limits the confidence that can currently be extracted from those indicators. The scientifically defensible position is therefore neither that current artificial systems are conscious nor that non-biological consciousness has been ruled out.
What makes Pachocki’s essay significant is that none of these unresolved philosophical questions is needed for its central practical warning. A non-conscious system could still be extraordinarily capable, highly autonomous in action selection, difficult to monitor, imperfectly aligned, able to operate through digital infrastructure, and deeply involved in the development of more capable successors. Mindhood is not a prerequisite for technological agency.
That observation changes the status of the title. An Alien Mind is best read not as a settled ontological classification but as a provocation about the inadequacy of human-centered categories. The phrase points toward several different hypotheses at once: alien capability profiles, alien forms of agency, alien modes of persistence, alien internal organization, alien routes to optimization, and perhaps, though this remains much less established, alien forms of subjectivity.
The source of the argument makes that distinction more important, not less. Pachocki is not an external futurist extrapolating from public benchmarks. He is OpenAI’s Chief Scientist, writing from inside one of the organizations producing the systems whose trajectory he describes.89 His institutional position provides access to evidence unavailable to outside observers while simultaneously making independent verification essential. When he reports internal results, those are primary-source claims. When he predicts recursive self-improvement, that is an informed forecast. When he argues for safety-gated scaling, voluntary slowdowns, external oversight, and international coordination, those are normative conclusions about how uncertainty should affect development.
The appropriate intellectual response is therefore neither to accept the phrase alien mind literally because it comes from a frontier laboratory nor to dismiss it because it sounds anthropomorphic. It is to decompose it.
What kind of intelligence is being measured? What kind of agency is actually present? What persists across interactions? Which parts of a system are understood, merely observable, or still epistemically opaque? Which parts of the AI-development loop have genuinely closed? What objectives are being optimized, and how do they generalize? What evidence, if any, bears specifically on phenomenal consciousness? And if artificial welfare eventually becomes plausible, what computational entity would be the subject whose welfare matters?
Those are different questions because alienness is multidimensional.
The central lesson of Pachocki’s essay, read through the competing interpretations examined here, is therefore stronger than the claim that we are building machines that resemble unfamiliar people. We may instead be building systems for which the familiar human correlations among intelligence, agency, identity, motivation, transparency, consciousness, and moral status no longer hold together.
That is enough to justify the word alien.
Whether it is enough to justify the word mind remains an open question.
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Shanahan, M. (2024). Talking about Large Language Models. Communications of the ACM, 67(2), 68–79. DOI.↩︎
Bender, E. M., & Koller, A. (2020). Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5185–5198. DOI.↩︎
Floridi, L. (2025). AI as Agency without Intelligence: On Artificial Intelligence as a New Form of Artificial Agency and the Multiple Realisability of Agency Thesis. Philosophy & Technology, 38, Article 30. DOI.↩︎
Chalmers, D. J. (2024). Does Thought Require Sensory Grounding? From Pure Thinkers to Large Language Models. arXiv. Chalmers argues against sensory grounding as a universal necessary condition for thought while explicitly declining to infer that current language models therefore think or understand. Preprint.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Legg, S., & Hutter, M. (2007). Universal Intelligence: A Definition of Machine Intelligence. Minds and Machines, 17(4), 391–444. DOI.↩︎
Legg, S., & Hutter, M. (2007). Universal Intelligence: A Definition of Machine Intelligence. Minds and Machines, 17(4), 391–444. DOI.↩︎
Chollet, F. (2019). On the Measure of Intelligence. arXiv. Preprint.↩︎
Hernández-Orallo, J. (2017). Evaluation in artificial intelligence: From task-oriented to ability-oriented measurement. Artificial Intelligence Review, 48(3), 397–447. DOI.↩︎
Bender, E. M., & Koller, A. (2020). Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5185–5198. DOI.↩︎
Chalmers, D. J. (2024). Does Thought Require Sensory Grounding? From Pure Thinkers to Large Language Models. arXiv. Chalmers argues against sensory grounding as a universal necessary condition for thought while explicitly declining to infer that current language models therefore think or understand. Preprint.↩︎
Dennett, D. C. (1988). Précis of The Intentional Stance. Behavioral and Brain Sciences, 11(3), 495–505. DOI.↩︎
Floridi, L., & Sanders, J. W. (2004). On the Morality of Artificial Agents. Minds and Machines, 14(3), 349–379. DOI.↩︎
Floridi, L. (2025). AI as Agency without Intelligence: On Artificial Intelligence as a New Form of Artificial Agency and the Multiple Realisability of Agency Thesis. Philosophy & Technology, 38, Article 30. DOI.↩︎
Floridi, L., & Sanders, J. W. (2004). On the Morality of Artificial Agents. Minds and Machines, 14(3), 349–379. DOI.↩︎
Chalmers, D. J. (2026). What we talk to when we talk to language models. Manuscript, Version 2, uploaded April 14, 2026. Chalmers’s May 7, 2026 Sarah Douglas Lecture at the University of California, Berkeley gives an authoritative public presentation of the same account of quasi-agents and memory-connected virtual threads. Preprint record. Official lecture page.↩︎
Chalmers, D. J. (2026). What we talk to when we talk to language models. Manuscript, Version 2, uploaded April 14, 2026. Chalmers’s May 7, 2026 Sarah Douglas Lecture at the University of California, Berkeley gives an authoritative public presentation of the same account of quasi-agents and memory-connected virtual threads. Preprint record. Official lecture page.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Humphreys, P. (2009). The philosophical novelty of computer simulation methods. Synthese, 169(3), 615–626. DOI.↩︎
Greif, H. (2022). Analogue Models and Universal Machines. Paradigms of Epistemic Transparency in Artificial Intelligence. Minds and Machines, 32, 111–133. DOI.↩︎
Olah, C., Satyanarayan, A., Johnson, I., Carter, S., Schubert, L., Ye, K., & Mordvintsev, A. (2018). The Building Blocks of Interpretability. Distill. DOI.↩︎
Lipton, Z. C. (2018). The mythos of model interpretability. Communications of the ACM, 61(10), 36–43. DOI.↩︎
Durán, J. M. (2026). Against epistemic transparency of algorithms. Philosophy of Science, First View, 1–11. Published online July 6, 2026. DOI.↩︎
Durán, J. M. (2026). Against epistemic transparency of algorithms. Philosophy of Science, First View, 1–11. Published online July 6, 2026. DOI.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Good, I. J. (1966). Speculations Concerning the First Ultraintelligent Machine. Advances in Computers, 6, 31–88. The published version notes that the manuscript was completed in 1964. DOI.↩︎
Chalmers, D. J. (2010). The Singularity: A Philosophical Analysis. Journal of Consciousness Studies, 17(9–10), 7–65. Author manuscript. Institutional record.↩︎
Chen, M., Wang, L., & Qu, B. (2026). Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops. arXiv. The paper surveys 1,250 arXiv papers from 2024–2026; it is used here as a recent survey preprint rather than as a peer-reviewed consensus statement. Preprint.↩︎
Chalmers, D. J. (2010). The Singularity: A Philosophical Analysis. Journal of Consciousness Studies, 17(9–10), 7–65. Author manuscript. Institutional record.↩︎
Hutter, M. (2012). Can Intelligence Explode?Journal of Consciousness Studies, 19(1–2), 143–166. Preprint.↩︎
OpenAI. (2026). Research acceleration: The view inside OpenAI. OpenAI, September 6, 2026. The reported measurements are preliminary internal metrics and are treated here as primary institutional self-report rather than independent evidence of autonomous recursive self-improvement. Official report.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Chen, M., Wang, L., & Qu, B. (2026). Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops. arXiv. The paper surveys 1,250 arXiv papers from 2024–2026; it is used here as a recent survey preprint rather than as a peer-reviewed consensus statement. Preprint.↩︎
Chen, M., Wang, L., & Qu, B. (2026). Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops. arXiv. The paper surveys 1,250 arXiv papers from 2024–2026; it is used here as a recent survey preprint rather than as a peer-reviewed consensus statement. Preprint.↩︎
Chalmers, D. J. (2010). The Singularity: A Philosophical Analysis. Journal of Consciousness Studies, 17(9–10), 7–65. Author manuscript. Institutional record.↩︎
Hutter, M. (2012). Can Intelligence Explode?Journal of Consciousness Studies, 19(1–2), 143–166. Preprint.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Bostrom, N. (2012). The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents. Minds and Machines, 22, 71–85. DOI.↩︎
Bostrom, N. (2012). The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents. Minds and Machines, 22, 71–85. DOI.↩︎
Hadfield-Menell, D., Dragan, A., Abbeel, P., & Russell, S. (2017). The Off-Switch Game. In Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, 220–227. DOI.↩︎
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete Problems in AI Safety. arXiv. Preprint.↩︎
Krakovna, V., Uesato, J., Mikulik, V., Rahtz, M., Everitt, T., Kumar, R., Kenton, Z., Leike, J., & Legg, S. (2020). Specification gaming: the flip side of AI ingenuity. Google DeepMind, April 21, 2020. Research article.↩︎
Hadfield-Menell, D., Dragan, A., Abbeel, P., & Russell, S. (2016). Cooperative Inverse Reinforcement Learning. Advances in Neural Information Processing Systems 29, 3909–3917. Proceedings paper. Preprint.↩︎
Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., & Garrabrant, S. (2019; revised 2021). Risks from Learned Optimization in Advanced Machine Learning Systems. arXiv. The mesa-optimization framework is theoretical and is not evidence that current frontier language models have been shown to contain stable mesa-objectives. Preprint.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Long, R., Sebo, J., Butlin, P., Finlinson, K., Fish, K., Harding, J., Pfau, J., Sims, T., Birch, J., & Chalmers, D. J. (2024). Taking AI Welfare Seriously. arXiv. The report advances a precautionary argument under uncertainty rather than establishing that current AI systems are conscious, robustly agentic in the morally relevant sense, or moral patients. DOI.↩︎
Long, R., Sebo, J., Butlin, P., Finlinson, K., Fish, K., Harding, J., Pfau, J., Sims, T., Birch, J., & Chalmers, D. J. (2024). Taking AI Welfare Seriously. arXiv. The report advances a precautionary argument under uncertainty rather than establishing that current AI systems are conscious, robustly agentic in the morally relevant sense, or moral patients. DOI.↩︎
Butlin, P., Long, R., Elmoznino, E., Bengio, Y., Birch, J., Constant, A., Deane, G., Fleming, S. M., Frith, C., Ji, X., Kanai, R., Klein, C., Lindsay, G., Michel, M., Mudrik, L., Peters, M. A. K., Schwitzgebel, E., Simon, J., & VanRullen, R. (2023). Consciousness in Artificial Intelligence: Insights from the Science of Consciousness. arXiv. The proposed computational properties are theory-derived indicators, not validated direct detectors of phenomenal consciousness. Preprint.↩︎
Butlin, P., Long, R., Elmoznino, E., Bengio, Y., Birch, J., Constant, A., Deane, G., Fleming, S. M., Frith, C., Ji, X., Kanai, R., Klein, C., Lindsay, G., Michel, M., Mudrik, L., Peters, M. A. K., Schwitzgebel, E., Simon, J., & VanRullen, R. (2023). Consciousness in Artificial Intelligence: Insights from the Science of Consciousness. arXiv. The proposed computational properties are theory-derived indicators, not validated direct detectors of phenomenal consciousness. Preprint.↩︎
Lennon, P. (2026). How Seriously Should We Take AI Welfare? Constraints From the Epistemology of Consciousness. Philosophy and Phenomenological Research, 113(2), 441–452. First published July 13, 2026. DOI.↩︎
Long, R., Sebo, J., Butlin, P., Finlinson, K., Fish, K., Harding, J., Pfau, J., Sims, T., Birch, J., & Chalmers, D. J. (2024). Taking AI Welfare Seriously. arXiv. The report advances a precautionary argument under uncertainty rather than establishing that current AI systems are conscious, robustly agentic in the morally relevant sense, or moral patients. DOI.↩︎
Lennon, P. (2026). How Seriously Should We Take AI Welfare? Constraints From the Epistemology of Consciousness. Philosophy and Phenomenological Research, 113(2), 441–452. First published July 13, 2026. DOI.↩︎
Long, R., Sebo, J., Butlin, P., Finlinson, K., Fish, K., Harding, J., Pfau, J., Sims, T., Birch, J., & Chalmers, D. J. (2024). Taking AI Welfare Seriously. arXiv. The report advances a precautionary argument under uncertainty rather than establishing that current AI systems are conscious, robustly agentic in the morally relevant sense, or moral patients. DOI.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Legg, S., & Hutter, M. (2007). Universal Intelligence: A Definition of Machine Intelligence. Minds and Machines, 17(4), 391–444. DOI.↩︎
Chollet, F. (2019). On the Measure of Intelligence. arXiv. Preprint.↩︎
Hernández-Orallo, J. (2017). Evaluation in artificial intelligence: From task-oriented to ability-oriented measurement. Artificial Intelligence Review, 48(3), 397–447. DOI.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Bostrom, N. (2012). The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents. Minds and Machines, 22, 71–85. DOI.↩︎
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete Problems in AI Safety. arXiv. Preprint.↩︎
Krakovna, V., Uesato, J., Mikulik, V., Rahtz, M., Everitt, T., Kumar, R., Kenton, Z., Leike, J., & Legg, S. (2020). Specification gaming: the flip side of AI ingenuity. Google DeepMind, April 21, 2020. Research article.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Humphreys, P. (2009). The philosophical novelty of computer simulation methods. Synthese, 169(3), 615–626. DOI.↩︎
Greif, H. (2022). Analogue Models and Universal Machines. Paradigms of Epistemic Transparency in Artificial Intelligence. Minds and Machines, 32, 111–133. DOI.↩︎
Durán, J. M. (2026). Against epistemic transparency of algorithms. Philosophy of Science, First View, 1–11. Published online July 6, 2026. DOI.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
OpenAI. (2026). Research acceleration: The view inside OpenAI. OpenAI, September 6, 2026. The reported measurements are preliminary internal metrics and are treated here as primary institutional self-report rather than independent evidence of autonomous recursive self-improvement. Official report.↩︎
Good, I. J. (1966). Speculations Concerning the First Ultraintelligent Machine. Advances in Computers, 6, 31–88. The published version notes that the manuscript was completed in 1964. DOI.↩︎
Chalmers, D. J. (2010). The Singularity: A Philosophical Analysis. Journal of Consciousness Studies, 17(9–10), 7–65. Author manuscript. Institutional record.↩︎
Hutter, M. (2012). Can Intelligence Explode?Journal of Consciousness Studies, 19(1–2), 143–166. Preprint.↩︎
Chen, M., Wang, L., & Qu, B. (2026). Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops. arXiv. The paper surveys 1,250 arXiv papers from 2024–2026; it is used here as a recent survey preprint rather than as a peer-reviewed consensus statement. Preprint.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
Pachocki, J. (2026). An Alien Mind. OpenAI, September 6, 2026. Pachocki is Chief Scientist at OpenAI; OpenAI announced his appointment to that role in May 2024 and states that he previously served as Director of Research, leading work including GPT-4 and OpenAI Five. In An Alien Mind, he argues that scaled deep-learning systems are grown more than designed, that their intelligence is not directly comparable to human intelligence, that alignment is fundamentally a problem of generalization, that chain-of-thought monitoring is becoming less reliable as systems become more capable, and that continued progress may lead toward increasingly automated AI research and recursive self-improvement. He argues that capability scaling should be constrained by confidence in alignment and monitoring and that humans should remain part of the improvement loop. Official essay. OpenAI appointment announcement.↩︎
@online{montano,
author = {Montano, Antonio},
title = {What {Kind} of {Thing} {Is} an {Alien} {Mind?}},
url = {https://antomon.github.io/longforms/what-kind-of-thing-is-an-alien-mind/},
langid = {en},
abstract = {In \_An Alien Mind\_, OpenAI Chief Scientist Jakub
Pachocki argues that frontier AI is increasingly \_grown more than
designed\_, that machine intelligence should not be treated as a
simple point on a human scale, that alignment is fundamentally a
problem of generalization, that monitoring may become harder as
reasoning systems grow more capable, and that AI-assisted research
could eventually develop into recursive self-improvement. Because
these claims come from a senior scientist inside one of the
organizations operating at the frontier of generative AI, they
deserve to be treated neither as detached speculation nor as
independently established fact, but as an unusually informed and
comparatively favorable interpretation of the technological project
itself. This essay reconstructs Pachocki’s argument and tests it
against competing accounts of intelligence, agency, artificial
interlocutors, epistemic opacity, interpretability, recursive
self-improvement, alignment, consciousness, and moral patienthood.
Its central claim is that the most defensible meaning of \_alien\_
is not an established non-human phenomenology but a multidimensional
decoupling of properties that human-centered concepts tend to bundle
together: competence need not establish understanding, agency need
not imply human-like psychology, persistence need not map onto one
enduring individual, inspectability need not yield intelligibility,
recursive improvement need not amount to an intelligence explosion,
optimization need not imply human values, and behavioral
sophistication need not establish consciousness or moral status. The
practical importance of these distinctions extends well beyond
philosophy of mind. Individuals, organizations, and democratic
institutions increasingly make decisions about what authority to
delegate to AI, which risks to prioritize, how responsibility should
be allocated, what forms of oversight are justified, and how much
control over consequential systems should remain human. Those
decisions depend on the conceptual model of AI that precedes them.
Anthropomorphic models can produce overtrust and misdirected fear;
reductively mechanistic ones can obscure functional agency and
systemic power. The essay therefore treats conceptual clarification
as part of the epistemic infrastructure of democratic decision
making: before societies decide whether to accelerate, constrain,
regulate, delegate to, or reorganize themselves around advanced AI,
they must first understand what kinds of capabilities, agency,
opacity, alignment problems, feedback loops, and possible forms of
subjecthood are actually at issue. The governing principle is
simple: **explanation before action**.}
}