A variety of large language models, commonly called “artificial intelligence” (AI) have seen increasing popularity and widespread adoption in recent years. In fact, the popularity of chatbots like Claude, Gemini, and ChatGPT has caused many to think that AI is synonymous with chatbots. However, AI comprises far more than chatbots, encompassing a variety of machine learning algorithms including reinforcement learning [1].
Reinforcement learning is a machine learning paradigm where an agent takes actions in a dynamic environment to maximize a reward (Figure 1). This can also be described as learning a behavior through trial-and-error guided by rewards.
Figure 1: Reinforcement Learning
For example, reinforcement learning may be used to guide an autonomous vehicle to its destination. In this case, the agent is the autonomous vehicle, and the environment is comprised of all the roads, buildings, trees, and pedestrians surrounding the vehicle. The states describe that environment, which may include things like road traffic, weather, and time of day. When an autonomous vehicle takes an action, like pushing on the gas/brake pedal or turning the steering wheel, this changes the surrounding environment since the distance between the autonomous vehicle and the surrounding objects changes. The autonomous vehicle measures that change in position with its sensors, using that information to make its next decision (push on the gas/brake pedal, turn the steering wheel, etc.). In reinforcement learning, the way an autonomous vehicle decides to make its next decision is based on a reward. For example, if the vehicle reaches its destination on time, the vehicle receives a reward of +1. If other drivers start honking at the autonomous vehicle, then it receives a reward of -1. And if the autonomous vehicle gets in a collision, it receives a reward of -100.
The goal of reinforcement learning is to find a set of actions that will maximize this reward. Over thousands of simulated trials, the agent finds a set of actions that achieves this maximum reward, in this case arriving at the destination on time with no collisions or honking from other vehicles. To find this set of actions, an agent simply solves an optimization problem where the objective being maximized (the reward function) is specified by a human.
Reinforcement learning is a powerful and useful tool for training machines, but we must be careful not to treat humans in the same way as machines. Throughout Scripture, God does use the rewards of covenant blessings and curses in response to Israelite behavior to point them back to Him (Deut 28-30; Lev 26). However, the primary means by which He teaches us how to live faithfully is through His Word (2 Tim 3:16-17; Matt 4:4; Ps 119:9, 105). When we learn, we should always interpret God’s act-revelation (the natural world) according to His Word-revelation (Scripture).
Reinforcement learning was originally designed to mimic a behavioristic view and understanding of human beings. B. F. Skinner, a Harvard psychologist, believed that human behavior is shaped by reinforcement rather than free will (Figure 2) [2]. This idea was then applied to machines to produce desired behaviors.
Figure 2: Operant Conditioning
This view stands in contrast to a biblical understanding of human beings since we do have free will (1 Cor 7:37; Judg 17:6; 21:25; Deut 12:8) and are therefore responsible for our actions (Rev 20:12; Matt 12:36; Rom 3:19; 14:12; 2 Cor 5:10; 1 Pet 4:5; Heb 4:13). Behaviorism ultimately encourages the idea that humans are no more than their brains, a view shared with naturalism.
According to naturalism, the human mind emerges out of the processing of inputs by a network of physical neurons (emergentism). Since there is no immaterial reality, both the brain and the mind only refer to the physical network of neurons (mind/brain identity theory) [3]. Daniel Wolpert, neuroscientist at Columbia University, claims that “the brain evolved, not to think or feel, but to control movement” [4]. This statement has been supported by the fact that a sea squirt digests its own brain when it decides not to move anymore.
In contrast, Christianity holds that there is a difference between the brain and the mind. The brain refers to the physical/material aspect while the mind refers to the immaterial aspect. We can distinguish between the material and immaterial aspects of the brain/mind, but they cannot be separated (except in the intermediate state). After all, the internal and external aspects of the creation mandate (Gen 1:28) show that the brain is to be used for more than just external movement. The internal aspect calls us to develop our minds so that we think God’s thoughts after Him, while the external aspect calls us to develop and subdue the earth in a God-glorifying way, which necessarily includes movement.
Christianity and naturalism therefore explain Moravec’s paradox in very different ways. Moravec’s paradox states that “it is comparatively easy to make computers exhibit adult level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility” [5]. Marvin Minsky, computer scientist at MIT, put it similarly by saying that “we’re more aware of simple processes that don’t work well than of complex ones that work flawlessly” [6]. In other words, it is easy for a computer to beat a human when deciding where to move the chess pieces, but when it comes to physically moving the chess pieces, it is easy for a child to beat a robot. The reason for this is that it is particularly difficult for machines to generalize across different environments and situations, like learning how to open a door with a round knob and then trying to generalize that process to doors with lever handles or pull handles.
The typical explanation for Moravec’s paradox is answered according to a naturalistic evolutionary worldview. Hans Moravec, computer scientist at Carnegie Mellon University, argued [7] that
We should expect the difficulty of reverse-engineering any human skill to be roughly proportional to the amount of time that skill has been evolving in animals.
The oldest human skills are largely unconscious and so appear to us to be effortless.
Therefore, we should expect skills that appear effortless to be difficult to reverse-engineer, but skills that require effort may not necessarily be difficult to engineer at all. [8]
Rodney Brooks, roboticist at MIT, similarly states that “intelligence was thought to be best characterized as the things that highly educated male scientists found challenging. Projects included having a computer play chess, carry out integration problems that would be found in a college calculus course, prove mathematical theorems, and solve complicated word algebra problems. The things that children of four or five years could do effortlessly, such as visually distinguishing between a coffee cup and a chair, or walking around on two legs, or finding their way from their bedroom to the living room were not thought of as activities requiring intelligence.” [9]
In contrast to this naturalistic explanation, Christianity teaches that humans are not machines but are made in the image of God (Gen 1:26-27) [10]. We are rational, moral creatures who have a special relationship with God that He does not have with the rest of creation. In this way, humans are created to reflect the nature and character of the Trinitarian God who is unity/generality in one being (God) and diversity/particulars in three persons (Father, Son, Holy Spirit). An image-bearer can therefore seamlessly unite generals and particulars, as seen throughout Scripture. Adam worked in the garden, applying general methods of care to a variety of particular plants (Gen 2:15). Adam named the animals, giving general names to categories of particular animals (Gen 2:19-20). Adam and Eve walked in the garden, generalizing particular steps across various particular terrains (Gen 3:8-10).
Since humans are made in the image of the Trinitarian God who brings together generals and particulars in Himself (unity/generality in one being and diversity/particulars in three persons), humans are able to generalize their movements across a variety of particular environments and situations. In contrast, machines are not made in the image of God, so they do not possess a built-in capacity to unite generals and particulars. This is why it is difficult to design machines that imitate humans by generalizing movements across a variety of particular environments and situations (although we are getting better at designing machines that can generalize their movements). As we evaluate Moravec’s paradox and engage with machine learning algorithms like reinforcement learning, we should strive to interpret these things in the way God would have us see them, taking every thought captive to the obedience of Him (2 Cor 10:5).
[1] Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction, 2nd edition. Cambridge, MA: MIT Press, 2020.
[2] B. F. Skinner, The Behavior of Organisms: An Experimental Analysis. New York: Appleton-Century, 1938.
[3] Julia Haas, “Reinforcement Learning: A Brief Guide for Philosophers of Mind” in Philosophy Compass, 17(9), 2022.
[4] Daniel Wolpert, “The Real Reason for Brains,” TEDGlobal 2011.
[5] Hans Moravec, Mind Children: The Future of Robot and Human Intelligence. Cambridge, MA: Harvard University Press, 1988, p. 15.
[6] Marvin Minsky, The Society of Mind. New York: Simon and Schuster, 1986, p. 29.
[7] Hans Moravec, Mind Children: The Future of Robot and Human Intelligence. Cambridge, MA: Harvard University Press, 1988, pp. 15-16.
[8] https://en.wikipedia.org/wiki/Moravec%27s_paradox
[9] Rodney Brooks, Flesh and Machines: How Robots Will Change Us. New York: Pantheon Books, 2002, p. 36.
[10] Steve VanderLeest and Derek Schuurman, “A Christian Perspective on Artificial Intelligence: How Should Christians Think about Thinking Machines?” in Christian Engineering Conference, 2015, pp. 91-107.