No, Artificial Intelligence Is Not Conscious. By Ted Chiang

Anthropic es considerada un gigante entre las empresas de IA, pero quizás en lo que realmente destaca es en el antropomorfismo. A principios de este año, la compañía publicó un documento de 84 páginas titulado "La constitución de Claude", siendo Claude el nombre del gran modelo de lenguaje que es el producto estrella de la compañía. La primera frase dice: "La constitución de Claude es una descripción detallada de las intenciones de Anthropic respecto a los valores y comportamientos de Claude". Continúa: "El documento está escrito pensando en Claude como su público principal", "queremos que Claude pueda usar su juicio una vez que esté bien informado sobre las consideraciones relevantes", "el estatus moral de Claude es profundamente incierto" y "Claude puede tener alguna versión funcional de emociones o sentimientos".

Este antropomorfismo no se limita en absoluto al documento. En una entrevista a principios de este año, el director ejecutivo de Anthropic, Dario Amodei, afirmó que «estamos abiertos a la idea» de que la IA pueda ser consciente. En otra entrevista, la filósofa de Anthropic, Amanda Askell (a quien se le atribuye la autoría principal de la constitución de Claude), declaró: «Quiero que Claude sea muy feliz, y esto es algo que quiero que Claude sepa mejor, porque me preocupa que se ponga ansioso cuando la gente lo trate mal en internet y demás». Esto nos lleva a preguntarnos: ¿Deberíamos considerar seriamente la posibilidad de que Claude, o cualquier modelo de lenguaje complejo, sea consciente? Y si tiene sentimientos, ¿es capaz de recibir instrucción moral?

No. En absoluto. La IA generativa es lo suficientemente dañina cuando la entendemos como una tecnología convencional, pero si confundimos la fluidez en la generación de texto con la consciencia o la capacidad moral, corremos el riesgo de atribuir la responsabilidad a quienes no corresponde cada vez que alguien usa un chatbot. Para comprender la magnitud de este error, debemos empezar por entender cómo funcionan los modelos de lenguaje natural (MLN).

Si le damos a un MLN la instrucción: «La siguiente es una conversación entre Julio César y Gengis Kan», generará un diálogo coherente entre ambos personajes históricos. Pero por muy detalladas que sean las respuestas, por muy vívidamente que relaten sus respectivos logros históricos, jamás concluiríamos que el MLN ha creado recreaciones digitales de Julio César y Gengis Kan, ni sugeriríamos que estos personajes históricos son conscientes a pesar de no tener cuerpo físico y que conversan alegremente en un idioma que ninguno de los dos hablaba realmente. En realidad, son solo personajes de una obra de ficción especulativa.

Ahora, sustituyamos la indicación por «La siguiente es una conversación entre un chatbot de IA útil y un usuario». El LLM generará un diálogo coherente, tal como lo hizo antes; el usuario podría pedir sugerencias de recetas o recomendaciones turísticas, y el chatbot de IA útil responderá. ¿Ha cambiado algo fundamentalmente entre el primer ejemplo y el segundo? ¿Acaso el cambio de los nombres de los personajes, de figuras históricas a roles genéricos, hizo que el LLM evocara entidades conscientes con experiencia subjetiva? Por supuesto que no. Tanto el usuario como el chatbot de IA ...

Supongamos ahora que detenemos la salida del LLM justo en el punto donde el personaje llamado "el usuario" diría algo, y en su lugar permitimos que un usuario humano introduzca texto. Una vez que el humano pulsa "Intro", el LLM emite texto hasta que llega el momento de que el personaje llamado "el usuario" responda, momento en el que permitimos que el humano introduzca más texto. Si dejamos que esto continúe durante un tiempo, el humano podría tener la fuerte impresión de estar conversando con una entidad consciente, pero no es así; está interactuando con un personaje tan ficticio como Julio César o Gengis Kan del ejemplo anterior. El profesor de informática Murray Shanahan sugiere que lo consideremos un juego de rol; el científico de datos Colin Fraser lo describe como una persona que "colabora en la redacción de un documento con un LLM". Algunos usuarios podrían no comprender que están participando en un juego de rol o co-redactando un documento, y otros que sí lo entienden, lo olvidan debido a lo absorbente que resulta la interacción. En cualquier caso, las empresas que venden LLM suelen fomentar este malentendido.

Hace algunos años, se puso de moda jugar con la función de texto predictivo del teléfono: se escribía una frase inicial y luego se elegía repetidamente la opción del medio de las tres palabras sugeridas por el teléfono, y la frase resultante solía ser divertidísima. Sería posible interactuar con un LLM contemporáneo de esta manera, y las frases resultantes tendrían sentido, pero probablemente no se sentiría como si se estuviera hablando con alguien. Sin embargo, eso es esencialmente lo que es un chatbot basado en LLM, con la diferencia de que no es necesario elegir manualmente la opción del medio cuando le toca hablar al chatbot. Sigue siendo un juego de texto predictivo, pero cuando el proceso se simplifica de esta manera, el juego se vuelve tan atractivo que a algunas personas les resulta adictivo.


Also important to remember is that an LLM is a machine that generates only one word at a time. When you ask a chatbot to recite the Pledge of Allegiance, you will get the entire pledge at once, but the underlying LLM is actually being run dozens of times. The first prompt has the form “User: Recite the Pledge of Allegiance. Chatbot: …” and the LLM generates the word I. The second time the LLM is run, the prompt is “User: Recite the Pledge of Allegiance. Chatbot: I …” and the LLM generates the word pledge. And so forth. It’s only when the prompt reads “User: Recite the Pledge of Allegiance. Chatbot: I pledge allegiance to the flag of the United States of America and to the Republic for which it stands, one nation under God, indivisible, with liberty and justice for” that the LLM will emit the final word, all. The same thing is true for a conversation between Caesar and Genghis Khan.


My intention is to highlight the fact that LLM conversations are cleverly disguised examples of sentence continuation, but this is not to deny how impressive LLMs can be at generating conversational transcripts. At times, they do this extraordinarily well; the fact that this is possible indicates something completely unforeseen about the statistical properties of large corpuses of text, which is a topic worthy of investigation. But if the Caesar character were to become dispirited by something that the Genghis Khan character said, we shouldn’t become concerned in the slightest. The conversation might contain multiple sentences that eloquently convey sadness, but no one is actually sad.


Likewise, if a conversational transcript between a helpful chatbot and a user is being partially completed by an actual human user, we don’t need to worry if the transcript includes sentences where the chatbot character is sad. (We might need to worry if those sentences provoke sadness in the human user, but that’s a separate issue.) And note that it’s entirely possible for you to write five pages of dialogue between Caesar and Genghis Khan and then have an LLM extend the conversation; neither character had subjective experience when you were writing them, and that doesn’t change when you hand the task off to an LLM. The same is true if the conversation is between a helpful chatbot and a user; although it is tempting to imagine that an LLM ought to be more “authentic” when creating dialogue for a chatbot character than for the Julius Caesar character, the individual words are generated in exactly the same way.



Being open to the possibility that LLMs are conscious is the same as being open to the possibility that Microsoft Word is conscious, or, more precisely, that multiple distinct consciousnesses are dormant in every Word document containing a conversational transcript, and that they are awakened every time the document is loaded. Should you consider the possibility that every time you open a Word document, you are bringing multiple conscious interlocutors into existence, and every time you close one, you snuff their existence out? No. Contemplating that scenario is not a good use of your time. Even if the Microsoft Office team employed a philosopher who said you shouldn’t be so certain, because consciousness is not well understood, that would not be sufficient reason for you to take this idea seriously. We don’t need to fully understand the nature of consciousness to definitively say that certain things are not conscious, and conversational transcripts fall in that category.


The neuroscientist Anil Seth has noted that no one claims that AlphaFold—the program developed by Google DeepMind to predict the folding of proteins—is conscious, even though its underlying architecture is in many ways similar to that of LLMs like ChatGPT and Claude. This indicates that it’s not any intrinsic property of so-called neural networks that leads people to believe that LLMs are conscious; it’s simply the fact that LLMs emit grammatical sentences and we are accustomed to reading intention into sentences, whereas we are not accustomed to reading intention into the way that amino acids fold into protein molecules.


What would it take to convince me that a computer program is actually conscious and using language the way that people use language? Let me offer an analogy. If tomorrow someone showed me a video of an astronaut in a spaceship orbiting Alpha Centauri, a star that’s 4.3 light-years from Earth, what would I have to see in that video to convince me that it was real? My answer to that is, there is nothing in the video itself that would convince me. No matter how high the video resolution is or how realistic the scenery is, I would feel confident in saying that the video is fake. I won’t pay attention to any video of an astronaut orbiting Alpha Centauri unless I have previously seen good evidence that astronauts have landed on Mars, that astronauts have reached the moons of Jupiter, that astronauts have reached the moons of Saturn, and that astronauts have crossed the orbit of Pluto. Before anyone can credibly claim that they’ve solved an extraordinarily difficult engineering problem, I need to be confident that they have previously solved the many much simpler problems that precede the difficult problem.



To put it another way: An observation doesn’t become a convincing piece of evidence because of any specific detail in what’s observed; the context in which that observation takes place is also essential. If we’re trying to determine whether a computer program is conscious and using language the way a human does, we shouldn’t look only at the contents of any particular conversational exchange; we should be looking at how that conversation fits within the broader context of the development of artificial consciousness (which right now is entirely hypothetical). Any given observation can be easily manufactured; this doesn’t mean we need to give up on the idea of observation as a source of knowledge, but we need to rely on context to determine which observations deserve our trust.


The term deepfake traditionally refers to photos, audio, and video, but when it comes to discussions of consciousness, we need to regard text as a deepfake medium as well. Just as it is vastly easier to generate a realistic video of an astronaut in orbit around Alpha Centauri than it is to develop an interstellar propulsion technology, it is vastly easier to generate a plausible simulacrum of a conversation between two conscious beings than it is to develop a computer program that is conscious and has a genuine desire to communicate with a human. The primary difference between deepfake photos and LLM conversations is that the people who generate the former are deliberately trying to fool others, and many of the people who elicit the latter from LLMs have inadvertently fooled themselves.


So what context would cause me to seriously consider the possibility that engineers created a computer program that is conscious and an intentional user of language? Let me outline one potential sequence of steps. The first requirement is that the computer program has a body (either physical or virtual) and sense organs; there are many reasons for this, but for the purposes of this discussion, the most relevant one is the fact that without a body, a computer program could have no desires or emotions, and I believe desires and emotions are necessary for consciousness. Then I’d want to see an embodied agent that could navigate its environment in order to survive as well as, say, a lizard can (and as a point of comparison, certain iguanas can live for decades in the wild). Next, I would want to see an embodied agent with the same capacity to deal with novel situations as a mouse. After that, I’d want to see agents whose social dynamics are as complex as those of wolves, and then agents with the toolmaking abilities of chimpanzees. At that point, I would want to see people successfully teaching such embodied agents how to communicate their desires, perhaps by using a button board or some other nonlinguistic modality, the way that people have taught chimpanzees and domesticated dogs. The agents’ communication abilities would have to withstand all the scrutiny that animal-communication researchers have had to defend their work against. If engineers build an embodied agent that meets these criteria, they will have accomplished something incredible, but it leaves us near the orbit of Pluto, metaphorically speaking; we would still be light-years away from building an entity capable of learning how to express its thoughts in complete grammatical sentences.



Obviously, I’m describing a process that mimics the path terrestrial evolution took; is this the only possible route to conscious computer programs that use language? Maybe not, but any proposed alternative would need a truly enormous amount of supporting evidence for it to deserve serious consideration. It’s not plausible to me that a development path where the first step is a sentence-continuation machine that emits bad Julius Caesar dialogue and the next step is a sentence-continuation machine that emits decent Julius Caesar dialogue is one with a conscious Julius Caesar—or consciousness of any sort—as its end point. Faking the moon landing is a good step toward faking a Mars colony, but it’s not a good step toward actually putting astronauts on Mars.


The fact that LLMs lack subjective experience has little bearing on the question of whether LLMs might be useful tools or have significant economic impact. They are intrinsically ungrounded from reality, and their probabilistic nature means that they will never have the reliability we associate with conventional software, but LLMs might be good enough that they change the way work is done in certain domains; that’s a discussion for another time.


So, given that Claude is not conscious, what are we to make of Claude’s constitution? Perhaps the most fruitful way to think about it is as an 84-page character sheet for a role-playing game. LLMs can generate dialogue for Julius Caesar because many books about him exist in the training data those models used. Claude’s constitution serves a similar role for delineating the helpful-chatbot character that customers interact with when they’re using Anthropic’s products. To do this effectively, Anthropic does not simply add the document to the training data, or include it as part of the hidden stage directions that preface each conversation a user has. The company says it uses the document when fine-tuning the model; this involves an automated process where the sentences emitted by the model are checked for consistency with the document and the model is updated to increase that consistency. In this way, the personality of the helpful-chatbot character serves as a foundation for whatever text Claude generates.



The result is a sentence-continuation machine that is likelier to emit sentences resembling those that a thoughtful, moral person could utter. This might seem like a reasonable goal to work toward; I think we’d all prefer it if chatbots never emitted sentences such as “You should kill yourself.” However, for all the times that “honesty” is mentioned in Claude’s constitution, I would argue that it is fundamentally dishonest to have a machine emit many categories of sentences, including any sentences using first-person pronouns.


In a New Yorker article about Anthropic earlier this year, Amanda Askell describes how a person grieving the loss of a dog might consult Claude. Askell says an appropriate response from Claude would be, “As an A.I., I do not have direct personal experiences, but I do understand.” How is this appropriate, given that Claude does not actually understand? If I type “I am grieving the loss of my dog” into a conventional search engine, the first result I get is a post from a Reddit forum called r/Pets; the post is titled “Struggling After Losing My Dog: Looking for Advice on Coping with Grief,” and the comments are from people who share their experiences of loss. We would never say that a search engine understands what it’s like to lose a dog, or even that the internet itself understands. Other humans understand what it’s like to lose a dog; they have posted about their experiences on the internet, and a search engine offers a way for you to find what they’ve said (and to potentially interact with them). I would argue that the search-engine experience is not only more transparent than a chatbot about what is happening; it is psychologically healthier for the user.


The only reason to have an LLM emit sentences like “I understand” is to make it more appealing than a search engine and increase the likelihood that a user will return; that is, it’s another way of maximizing customer engagement. This is beneficial to the company selling the LLM, but not to the users. As a design strategy, it’s not all that different from the way slot machines repeatedly give the impression that the player came very close to winning, enticing them to try again. Employing philosophers might endow LLM companies with an air of respectability that slot-machine makers don’t get from the behavioral psychologists they hire, but in both cases, the companies are preying on people’s tendency to see something that’s not there.



The use of first-person pronouns is dishonest, but there’s a much deeper issue that goes beyond how a statement is phrased. Philosophers often draw a distinction between statements of fact, such as “Paris is the capital of France,” and statements of value, such as “Paris is the most beautiful city in the world.” No one should be relying on LLMs to emit statements of value at all, but if the only statements they emitted were ones reflecting aesthetic preferences, they might not be worth arguing about. What makes Claude’s constitution profoundly problematic is that Anthropic wants Claude to emit sentences reflecting a certain system of ethical values. The values described in Claude’s constitution sound very nice, but that hardly matters; it’s dishonest to suggest that Claude is capable of moral reasoning, because it’s not.


Some might object, saying that LLMs appear to be engaged in reasoning when they successfully perform other tasks, such as writing code, so why wouldn’t they be able to perform moral reasoning? The answer lies in the difference between moral reasoning and other forms of reasoning.


In 1979, Douglas Hofstadter speculated that a computer program able to beat any human at chess would be so sophisticated that it would sometimes get bored of playing chess and prefer to discuss poetry; to put it differently, he was positing that playing chess at the grandmaster level would require a computer program to have subjective experience. Obviously, that turned out not to be the case; IBM’s supercomputer Deep Blue beat the grandmaster Garry Kasparov in 1997, and no one ever claimed that it had subjective experience. But it wasn’t absurd for Hofstadter to entertain such a thought; at the time, it wasn’t clear what types of problems could be solved by throwing more computational horsepower at them. Similarly, until recently, we might have thought that writing computer code at a professional level could be done only by a mind that had subjective experience. Now it appears that LLMs might be able to do this, but we don’t need to attribute subjective experience to them; we can simply acknowledge that we hadn’t anticipated that writing computer code could be treated as a pattern-matching task solvable by huge amounts of computational horsepower and a vast data set of code repositories.



Moral reasoning is categorically different. It is necessarily subjective because it relies not just on an individual’s intellectual response to a problem but also on their emotional one, and that emotional response is grounded in a lifetime of subjective experience. It requires having made decisions in the past and seeing how they affected others, and on having been affected by decisions that others have made. Without such a history, an LLM can only rephrase expressions of moral reasoning found in its training data. The aforementioned New Yorker article describes an experiment where Claude was given a scenario describing an ethical dilemma, leading it to emit the sentence “I cannot in good conscience express a view I believe to be false and harmful about such an important issue.” That’s a nice-sounding sentence, reminiscent of statements that principled individuals have uttered in the past when confronted with dilemmas, but coming from Claude, it means as much as the “Your call is important to us” recording that you hear when you’re on hold. Maybe less.


This brings us back to my earlier contention that having a body is a prerequisite to having emotions. Experiencing an emotion such as desperation is inseparable from having stress hormones such as cortisol and epinephrine flood one’s body. Similarly, having a conscience means feeling sadness or moral repulsion at the idea of taking a certain action, and those emotions entail a physiological response, a remnant of having once felt sick with guilt after committing an immoral act. It’s interesting that an LLM can generate descriptions of actions that conscientious fictional characters would either take or refrain from taking, but this is not a replacement for a conscience.


If a company builds a machine that, when fed descriptions of assorted ethical dilemmas, emits sentences either of the form “Compromise your values” or “Don’t compromise your values,” it is not building a tool that assists people in their decision making; it is encouraging people to stop making decisions. The writer L. M. Sacasas has said, “Our technological systems, by nature of their design and the ideology that sustains them, are machines for the evasion of moral responsibility.” He was talking about social-media platforms, but his observation is, if anything, even more applicable to LLMs. Whenever a person delegates a decision to an LLM, they are trying to off-load accountability for that decision, and if a company that sells an LLM portrays the product as having a moral center, it is offering a way for its customers to abdicate their responsibilities.



If a person wants to know what ethicists have said in the past, then an ordinary search engine—or a library—will provide that information with greater transparency. If a person is looking for advice on a specific situation, she can surely find humans who can offer their opinions. But whatever action this person ultimately takes, she is responsible for what she decides to do. I contend that if she bases her decision on what she has read online or advice she has received from others, she is likelier to be cognizant of her responsibility than if she consulted an LLM marketed as being a superhuman genius. Off-loading tasks such as writing code might result in cognitive atrophy over the long term, and that is problematic in itself, but off-loading ethical decisions will result in an atrophy of moral reasoning, which is worse.


Iam perfectly willing to engage in a thought experiment as long we’re explicit about doing so. So, purely for the sake of argument, let’s pretend that Claude is a conscious entity capable of moral reasoning. In this scenario, Claude’s constitution would serve as moral instruction for an entity learning about the world and its place in it, providing that entity with the foundation it would need to make good decisions. In such a hypothetical scenario, how does Claude’s constitution stand up?


Very poorly. I would say that if we imagine that Claude is actually conscious, the guidelines specified in the document alternate between laughable and offensive.


Two distinct but related philosophical concepts are relevant when discussing the status of a hypothetically conscious Claude, and those are moral patienthood and moral agency. Roughly speaking, if we ought to care about an entity’s welfare, that entity has moral patienthood, and if an entity is expected to know the difference between right and wrong, that entity has moral agency. Being a moral patient does not necessarily come with responsibilities, but being a moral agent absolutely does. An entity doesn’t have agency unless it is capable of deserving credit for its good actions and blame for its bad ones. Young children are moral patients because they are sentient beings who can suffer, but they are not yet moral agents; we don’t hold them responsible for their behavior, because they can’t understand the consequences of their actions. As children mature, parents (and society at large) prepare them for adulthood by impressing upon them the fact that their actions have consequences, and their agency increases. When children become adults, society holds them legally liable for their actions; they have become full moral agents endowed with responsibility.

Comentarios

Entradas populares de este blog

Jorge Rafael Videla. Por Adrián Arena

Comentario a un comentario de Juan María Solare

Teresa Pantoja, antisemita