Podcasts · In this section
NOMOTO MEDIA

A History of the AI Dream

By Niklas S Osterman

Throughout history, humans have imagined and strived to create artificial beings with intelligence. From ancient legends of mechanical servants to modern deep learning systems, the story of artificial intelligence (AI) spans millennia. This script traces that journey in a series of thematic chapters, each highlighting a phase in the evolution of AI. We begin in the realm of myth and legend, travel through early mechanical inventions and philosophical ideas, and arrive at the cutting-edge AI of today, examining capabilities, limitations, and future implications.

Long before programmable machines existed, people told stories of artificial beings endowed with life or intelligence. In Greek mythology, for example, the bronze giant Talos was a manmade guardian of Crete. According to the myth, the god Hephaestus forged Talos and brought him to life; the metal giant patrolled the island’s shores and protected it by hurling rocks at invaders. Such tales suggest that the idea of constructing an intelligent servant or protector captivated human imagination early on. Another Greek legend is the tale of Pygmalion, a sculptor whose ivory statue was brought to life by divine intervention – a mythical precursor to the concept of an artificial companion.

In Jewish folklore, the figure of the Golem represents another early vision of artificial life. A golem is a clay creature animated through mystical rituals, often to serve and protect its creator’s community. Medieval legends describe rabbis creating golems by inscribing holy words or placing sacred scrolls in the clay figure’s mouth, thereby bringing it to life. Unlike mythic automata that could speak or reason, the golem was typically mute and followed its creator’s commands literally. The golem stories highlight both the hopes and anxieties surrounding the creation of life from inanimate matter – an eerie foreshadowing of later debates about intelligent machines.

Ancient Chinese writings also contain stories of artificial humanoids. The Daoist text Liezi tells of an engineer-artificer named Yan Shi (or Yen Shih) who presented King Mu of Zhou (a ruler of the Zhou Dynasty, traditionally dated 10th century BCE) with a life-sized mechanical man. According to this legend, Yan Shi’s automaton could walk, gesture, sing, and appear startlingly lifelike, to the point that the king initially mistook it for a living. The story relates that when the king grew alarmed (especially after the automaton flirted with court ladies), the inventor hastily disassembled the figure to reveal gears and components, demonstrating that it was only an ingenious. This tale, while likely apocryphal, illustrates that over two thousand years ago, people were already grappling with the idea of artificial people and the boundary between the mechanical and the living.

These myths and folklore from different cultures – whether Greek, Jewish, Chinese, or others – reveal a timeless fascination with creating artificial life. Often, such stories imbued the creations with magic or divine power rather than technology, since no real machines could achieve such feats at the time. Yet they set a precedent, imagining mechanical or artificial servants long before the scientific knowledge existed to build them. They also introduced enduring themes: the utility of artificial beings (protectors like Talos or helpers like golems) and the ethical dilemmas they pose (as seen when creations cause fear or go out of control). In sum, humanity’s dream of intelligent machines is as old as civilization itself, beginning in myth and gradually moving toward reality.

Moving from myth into reality, the Renaissance and early modern periods saw the first actual constructions of mechanical automata – self-moving devices often designed to mimic living creatures. Innovators and artisans of these eras, fascinated by clockwork and mechanics, built devices that astonished their contemporaries and can be seen as precursors to robots. While these automata were not “intelligent” by modern standards, they demonstrated that machines could simulate lifelike behavior through clever engineering.

One early visionary was Leonardo da Vinci (1452–1519). Renowned as an artist and inventor, Leonardo designed a remarkable humanoid automaton known as Leonardo’s robot or the mechanical knight. Sketches from around 1495 show plans for a mechanical knight in armor that could sit up, wave its arms, and possibly move its head and jaw – all driven by an internal system of pulleys and gears. It’s believed that Leonardo may have even built a prototype of this robotic knight to impress his patron, the Duke of Milan. He also engineered a mechanical lion for a pageant in 1515, which could walk under its own power and present lilies from its chest – an entertaining automaton to delight King Francis I of France. These creations by Leonardo did not involve cognition or adaptability, but they showed that complex human-like motions could be achieved with purely mechanical means. In the context of history, they mark an important step from imaginary automata to engineered machines that mimic life.

By the 18th century (the Enlightenment era), Europe experienced an automaton craze. Craftsmen built elaborate clockwork figures that could move, play music, or even write, all through intricate mechanisms. Jacques de Vaucanson, a French inventor, became famous in the 1730s for his mechanical wonders. In 1738 he unveiled “The Flute Player,” a life-size figure that actually played the flute by moving its lips, fingers, and tongue – effectively a mechanical musician that amazed. Vaucanson followed it with an even more celebrated creation in 1739: The Digesting Duck. This gold-plated duck automaton could flap its wings, eat grain from your hand, and remarkably, simulate digestion by excreting matter – a clever illusion produced by hidden chambers and pre-loaded pellets. Though just a trick, the digesting duck stirred public debates about biology and machines, leading some to (mistakenly) believe Vaucanson had created an artificial life-form. These automata were entertainment, but they also raised profound questions: Where is the line between the living and the mechanical? Could a machine replicate not just motion but life processes?

Other artisans constructed androids (human-shaped automata) that performed sophisticated tasks. Around the 1770s, Swiss clockmaker Pierre Jaquet-Droz built three famous automata: The Musician (a female figure that played a keyboard instrument), The Draughtsman (a boy who could draw pictures with a pen), and The Writer (a boy who could write custom sequences of letters with quill and ink). The Writer, for instance, could be programmed via cams to produce different words and sentences, making it one of the earliest programmable machines. Likewise, Hungarian inventor Wolfgang von Kempelen created the “Mechanical Turk” in 1770 – an automaton that appeared to play chess against human opponents. (It was later revealed to be a hoax operated by a human concealed inside, but at the time many believed it a genuine thinking machine.) Despite the deception, the Turk spurred interest in whether a machine could truly play intellectual games – a question that would resurface in the 20th century with computer chess.

It’s important to note that parallel to Europe’s automata, similar traditions existed elsewhere. In the medieval Islamic world, inventors like Al-Jazari (1206 CE) had built programmable mechanical servants (for example, water-powered automata and clockwork musicians). In ancient China, there are records of mechanical toys and servants (as mentioned earlier with Yan Shi). These examples highlight a widespread engineering curiosity: the drive to simulate living actions through mechanics. By the early modern period, such automata were no longer merely legends – they were real, tangible devices. They demonstrated automation (predefined mechanical behavior) and programming (via cams, gears, punched cards, etc., as with Jacquard’s loom in 1801), even if they lacked any adaptive intelligence.

In summary, the Renaissance and Enlightenment automata were the direct forerunners of robots. They captivated audiences and proved that complex behaviors could be mechanized. However, they followed preset routines and had no autonomy beyond what their builders designed. The broader significance of these inventions was to plant the idea that perhaps a machine could someday be built to not only move like a living creature, but maybe also think or make decisions. That speculative leap would be taken by philosophers and mathematicians in the centuries that followed.

As engineers crafted automata, philosophers were reimagining the nature of the mind and machines. In the 17th and 18th centuries, thinkers began to ask whether reasoning could be systematically broken down like a machine’s operation. This was the era of the mechanistic worldview, which held that natural phenomena (including life and thought) might be explainable in terms of matter and motion. Key figures like René Descartes, Thomas Hobbes, and later Gottfried Leibniz and Baron d’Holbach explored whether human intellect was essentially an elaborate mechanism – a daring idea that laid conceptual groundwork for artificial intelligence.

René Descartes (1596–1650) famously argued for a dualism of mind and body – the mind as a non-material thinking substance, distinct from the mechanical body. Yet, Descartes also contemplated automation: he suggested that animals were essentially complex automata (without souls), and he hypothesized that even a human body, if without a mind, would be a “machine” behaving by pure physics. There is an anecdote that Descartes built a life-sized doll, “Francine,” in the likeness of his deceased daughter, which reportedly moved and startled a ship’s captain (though this story is likely apocryphal). Whether or not Descartes built an automaton, he clearly introduced the idea that much of human and animal behavior might be explicable mechanistically, foreshadowing the notion of mechanized thought.

Thomas Hobbes (1588–1679), a contemporary of Descartes, took a more strictly materialist stance. Hobbes believed that reasoning was like computation. In his work Leviathan (1651), Hobbes famously wrote: “For reason… is nothing but reckoning; that is, adding and subtracting”. By this he meant that thinking can be seen as a form of calculation, akin to arithmetic. This was a radical reduction: if thought is just calculation, then in principle a machine that calculates could think. Hobbes’ view encapsulates the early statement of what we might call the physical symbol system hypothesis (later articulated in AI): the idea that mental processes are manipulations of symbols or values, something a machine might be able to carry out. Such philosophical assertions paved the way for imagining a thinking machine not as a contradiction, but as an achievable engineering goal.

During the 17th century, Leibniz (1646–1716) expanded on these ideas. A polymath and philosopher, Leibniz dreamed of a universal logical language and a calculating machine to manipulate it. He developed the concept of a “characteristica universalis”, a universal symbolic language in which all knowledge could be expressed, and a method of reasoning called calculus ratiocinator to resolve disputes by calculation. He wrote that if we had such a language and method, “when controversies arise, there would be no more need of argument between two philosophers than between two accountants. It would suffice to take their pencils in hand and say, ‘Let us calculate.’”. Leibniz also actually built mechanical calculators (improving on Blaise Pascal’s earlier adding machine) that could add, multiply, etc. – an effort to mechanize arithmetic. His philosophical logic and literal machinery both illustrate the growing conviction that rational thinking could be automated.

Jumping forward, in the early 19th century, we see an important bridge between philosophical speculation and real computing machines: Charles Babbage and Ada Lovelace. Babbage (1791–1871) was an English mathematician who designed the Analytical Engine, an ambitious general-purpose mechanical computer. Although never completed in his lifetime, the design (1830s–1840s) included most elements of a modern computer: a “store” (memory), a “mill” (CPU for calculations), and even punch-card programming. Ada Lovelace (1815–1852), a mathematician who studied Babbage’s designs, is often considered the first computer programmer. In 1843, Lovelace wrote extensive notes on the Analytical Engine and speculated about its capabilities. Notably, she mused that such a machine might act upon things other than numbers, perhaps composing elaborate pieces of music if properly programmed – an early imaginative leap toward creative AI. However, Lovelace also offered a cautious insight, sometimes called the “Lovelace objection.” She wrote that the Analytical Engine could not originate ideas on its own: it could only do what it was instructed to do. In her words, “It is desirable to guard against the possibility of exaggerated ideas that might arise as to the powers of the machine”. In other words, a computational machine manipulates symbols but does not truly understand or think independently – it follows the rules given by its programmer.

Lovelace’s perspective captures a debate that persists to this day: Can machines ever do more than what they’re explicitly programmed or trained to do? Nonetheless, Babbage and Lovelace’s work proved that complex calculation could be automated, and they envisioned (in concept) a machine that could perhaps simulate the operations of mind. Their Analytical Engine design was, in essence, a blueprint for a general AI hardware, a century ahead of its time.

In summary, early modern philosophy and the mechanistic school of thought laid intellectual groundwork for AI by treating thinking as a kind of computation. Thinkers like Hobbes and Leibniz proposed that logic and reasoning could be formalized and automated. Later, Babbage’s mechanical computer design and Lovelace’s commentary brought these ideas into the engineering realm, showing how a machine might execute arbitrary algorithms. By the mid-19th century, humanity had the conceptual tools – logic, computation, and an understanding of mechanism – needed to start building actual thinking machines. What remained was to implement these ideas with technology, which would only become possible in the 20th century.

The late 19th and early 20th centuries saw rapid progress in formal logic and the invention of the first computers, both of which were crucial stepping stones toward artificial intelligence. Mathematicians developed systems to represent precise logical reasoning, and engineers built machines to carry out calculations automatically. Together, these advances created a foundation for thinking about “mechanical thought.”

On the logic side, figures like George Boole, Gottlob Frege, and Bertrand Russell revolutionized how we represent ideas and reasoning. Boole (1815–1864) introduced Boolean algebra in 1854, treating TRUE/FALSE logically like 1/0 in algebra – an essential concept for digital circuits and binary computation. Frege (1848–1925) developed a formal notation for quantifiable logic (predicate logic) in his Begriffsschrift (1879), which could express statements like “for all X, if X is human then X is mortal” in a precise symbolic form. This work made logic mathematical. Building on these, Bertrand Russell and Alfred North Whitehead published Principia Mathematica in 1910–1913, an ambitious attempt to derive all of mathematics from logical axioms. Their effort, while not entirely successful (it was later tempered by Gödel’s Incompleteness Theorems), demonstrated that complex reasoning could be broken into simple logical steps. The overall message of this era’s logic research was any reasoning that follows rules can potentially be formalized – a key insight if one dreams of programming a machine to reason.

In parallel, work in mathematics by David Hilbert and others posed the question of whether all of mathematical reasoning could be solved algorithmically. Hilbert’s famous challenge (early 1920s) asked: can we find a finite procedure to decide the truth of any mathematical statement? This challenge led to pivotal results: Kurt Gödel showed inherent limits (not all truths are provable by any one logical system), but Alan Turing and Alonzo Church provided positive news for computing – they independently formalized the notion of computation in the mid-1930s. Turing’s concept of the Turing Machine (1936) was an abstract model of a general-purpose computer, able to simulate any step-by-step logical process. Along with Church’s lambda calculus, this led to the Church-Turing Thesis, positing that anything computable in the intuitive sense can be computed by such a simple machine. In effect, Turing proved that a machine manipulating symbols (0s and 1s on a tape) could, given enough time, carry out any algorithm or logical deduction that a human could formalize. This was a profound theoretical leap: it implied that reasoning could be mechanized, at least in principle, because even very abstract processes (like solving math problems or playing chess by rules) are just computations a Turing machine could mimic.

World War II and its aftermath then brought about the digital electronic computer, turning theory into reality. Early electronic computers – such as the British Colossus (1943) and the American ENIAC (1945) – were initially built to solve specific problems (codebreaking, artillery tables). But soon the design of computers incorporated the idea of a stored program (proposed by John von Neumann, influenced by Turing’s work), yielding machines that could be reprogrammed to tackle any task. By the late 1940s and early 1950s, general-purpose computers like the Manchester Mark I and IBM 701 were operational. These machines were extremely limited by modern standards, but they established that symbol processing in hardware was feasible.

It’s in this climate that we see the convergence of logic and computing. In 1950, Alan Turing – now working with actual computers – published a seminal paper "Computing Machinery and Intelligence". In it, he tackled the question “Can machines think?” and sidestepped philosophical debates by proposing an operational test: the Turing Test. Turing said, essentially, if a machine can converse with a human in natural language such that the human cannot tell if it’s a machine or a person, then for all practical purposes we might call the machine intelligent. This test shifted the discussion to observable behavior (intelligent conversation) rather than ill-defined notions of “thinking.” Turing’s paper argued that machine intelligence was plausible, anticipated and refuted common objections, and even predicted people might one day speak of machines “thinking” without expecting to be contradicted. The Turing Test provided an influential target for AI research: a truly intelligent machine should be able to use language and reason about virtually any topic, indistinguishably from a human.

Another critical development of the 1940s was the bridging of logic with neurophysiology: the first artificial neural network model. In 1943, Warren McCulloch (a neuroscientist) and Walter Pitts (a logician) published a paper showing that a simplified model of neurons could perform logical operations. They abstracted neurons as binary units (firing or not firing) and demonstrated that networks of these units could implement logical propositions like AND, OR, NOT. Essentially, they proposed that the brain could be seen as a kind of computing device and that a sufficiently large neural network could, in principle, compute anything a Turing machine could. This was the birth of connectionismin AI. One of their followers, Marvin Minsky, built the first actual neural network machine in 1951 (called SNARC) using analog electronics to simulate a rat learning a maze. Although very primitive, these efforts planted the seed for an alternative approach to AI: instead of hand-crafted logic, let the machine learn and evolve intelligence similar to a brain. We will see this neural network thread re-emerge much later in the deep learning era.

By the early 1950s, the stage was set: formal logic showed how to represent complex reasoning; the concept of computation (Turing machines) and actual computers showed that machines can carry out those computations; and even brain-inspired models hinted at the possibility of learning. The question was no longer if it was theoretically possible to have a thinking machine, but rather how to build one. This question would soon coalesce into a new scientific field: artificial intelligence.

The period from 1950 to 1956 is often regarded as the birth of AI as a distinct field. We have already noted Alan Turing’s 1950 proposal of the Turing Test as a criterion for machine intelligence. Turing’s vision and the existence of programmable computers inspired a small group of scientists to organize what became AI’s founding event: the Dartmouth Summer Research Project on Artificial Intelligence in 1956. It was at this workshop that the very term “artificial intelligence” was coined, and the research agenda for the coming decades was set.

In the early 1950s, various researchers were independently exploring machine intelligence. For instance, Norbert Wienerwas developing cybernetics – the study of control and communication in animals and machines – which influenced thinking about learning and autonomous systems. At IBM and elsewhere, engineers were programming the first computers to play games like checkers and chess, seeing games as a proxy for intelligent decision-making. There was a growing optimism that with enough programming, these new machines could exhibit human-level intelligence in the foreseeable future.

This optimism culminated in the Dartmouth Workshop (summer 1956). The workshop was organized primarily by John McCarthy (then at Dartmouth College) along with Marvin Minsky, Claude Shannon, and Nathaniel Rochester. In their proposal, they boldly asserted: “every aspect of learning or any other feature of intelligence can be so precisely described that a machine can be made to simulate it.” That statement encapsulates the faith of the founders of AI: that human intelligence is, at root, an information process that can be understood and replicated in a machine. McCarthy, in that proposal, also introduced the name “Artificial Intelligence”for this endeavor – deliberately choosing a broad, attention-catching term (as opposed to more narrow terms like “complex information processing”).

The Dartmouth Workshop brought together about a dozen participants who would become luminaries of AI: besides McCarthy and Minsky, there were Allen Newell and Herbert Simon (from Carnegie Tech), Oliver Selfridge, Ray Solomonoff, Trenchard More, Arthur Samuel, and a few others. Over the course of eight weeks, they discussed how to make machines use language, form abstractions, and improve themselves (learning). Importantly, during this event Newell and Simon demonstrated a working AI program – the Logic Theorist, which they had developed in 1955 – that could prove theorems in symbolic logic. The Logic Theorist managed to prove 38 of the first 52 theorems of Russell and Whitehead’s Principia Mathematica, even finding new, more elegant proofs for some. This was a stunning proof of concept: a computer was not only doing arithmetic, but tackling abstract reasoning in mathematics, something that had been considered a pinnacle of human intellect.

Dartmouth 1956 is thus seen as the moment AI was born as a field. It gave researchers a common identity and mission. Attendees left with enthusiasm and went on to found the first AI research laboratories. For example, McCarthy soon moved to MIT and started the MIT AI Lab with Minsky; Newell and Simon continued their work at Carnegie Institute of Technology (later CMU); and research groups sprouted at IBM and Stanford. Funding agencies, influenced by the Dartmouth attendees’ optimism, began to pour resources into AI research. It’s worth noting that those pioneers genuinely believed success might be just around the corner. In the late 1950s and 1960s they often predicted that machines with human-level intelligence would exist within a generation. For instance, at the time Herbert Simon boldly forecast that “machines will be capable, within twenty years, of doing any work a man can do,” and Marvin Minsky in 1967 predicted that the AI problem would be solved in a generation. Such confidence may have been misplaced (as later events showed), but it drove an era of feverish innovation.

In summary, the mid-1950s established AI as an interdisciplinary field, uniting ideas from mathematics, computer engineering, psychology, and linguistics. The Turing Test gave a philosophical benchmark for machine intelligence, and the Dartmouth conference gave a name and agenda. Researchers set out to create programs that could reason, prove, speak, and learn – essentially, to fulfill the dream that a machine can “think.” The next two decades would see rapid development of different approaches to reach this goal, as well as the eventual confrontation with the immense complexity of the task.

Following the Dartmouth workshop, much of the research in the late 1950s and 1960s focused on what is now known as symbolic AI (or GOFAI: “Good Old-Fashioned AI”). This approach assumes that intelligence can be achieved by manipulating symbols – abstract representations of objects or concepts – according to logical rules. Early AI programs sought to encode human knowledge and reasoning explicitly as symbols and rules, and then have the computer apply these rules to solve problems. Several landmark programs from this era demonstrated both the potential and the limitations of symbolic AI.

One of the first was the Logic Theorist (mentioned above) by Newell and Simon (1955–56). It is considered the first working AI program. The Logic Theorist was designed to mimic the problem-solving skills of a human mathematician. It used a heuristic search through the space of possible logical proofs, guided by rules of thumb (heuristics) to prune unpromising paths, in order to prove theorems. Its success in proving a substantial number of propositions from Principia Mathematica was a proof-of-concept that computers can handle abstract reasoning tasks, not just numerical calculation. Following that, Newell and Simon developed a more general problem-solving framework called the General Problem Solver (GPS) in 1957–59. GPS attempted to solve any well-defined problem given formal operators, through means-ends analysis, but in practice it worked for simple puzzles more than “general” problems. These efforts introduced key AI methods like search algorithms and heuristic reasoning.

Around the same time, AI researchers turned to game playing and other clear-rule domains as testing grounds. In 1958, Herbert Simon predicted a computer would beat a human chess champion within a decade. That was overly optimistic, but early game programs were built: by the late 1950s, Arthur Samuel at IBM had a checkers (draughts) program that learned to improve by playing against itself, pioneering the concept of machine learning by self-play. Samuel’s checkers program, by the mid-1960s, reached a strong amateur level – another milestone showing that a machine could learn from experience to perform better than its programmer in a complex task.

In the realm of natural language, a famous early program was ELIZA, created in the mid-1960s (MIT, by Joseph Weizenbaum). ELIZA was a chatbot that simulated a Rogerian psychotherapist by answering the user’s statements with simple, scripted responses (often turning the user’s phrases into questions). For example, if you said “I feel unhappy,” ELIZA might respond, “Why do you feel unhappy?” This was achieved with pattern-matching rules and canned responses, not genuine comprehension. Nonetheless, when ELIZA was introduced, many people were struck by how “human-like” the conversation seemed. Users knew they were talking to a machine, yet some found themselves emotionally engaged, even fooled into thinking the computer understood them. This phenomenon was later dubbed the ELIZA effect – the tendency to project understanding onto a machine that merely outputs formulaic responses. ELIZA, though shallow, was the first chatbot and showed that language interaction with a computer was possible to a degree. It also highlighted how easily humans might anthropomorphize AI outputs.

Another notable project was SHRDLU, developed by Terry Winograd at MIT around 1968–1970. SHRDLU was a program that operated in a micro-world – a simplified, small domain meant to make AI problems more tractable. Its micro-world was a virtual toy “blocks world” with colored blocks, cones, and pyramids on a table. SHRDLU could accept typed natural language commands or questions about this blocks world, such as “Find a block which is taller than the one you are holding and put it into the box” or “What is inside the box?” The program could parse these sentences, convert them into actions or queries within its limited world, and respond appropriately in English. For example, it could answer, “The blue block is taller than the one I am holding,” or execute the commanded action if possible. SHRDLU’s success was remarkable: it integrated natural language understanding, planning, and vision/manipulation (in simulation). It showed that within a constrained context where the universe of discourse was small and well-defined, a machine could carry on a meaningful dialogue and manipulate objects intelligently. This gave hope that by gradually expanding such domains, one could scale up to full human-level language understanding.

During the 1960s, many other programs were built in domains like theorem proving, algebra (e. g., the program STUDENT could solve high school algebra word problems), and even early expert decision-making. All these systems used symbolic representations – lists, trees, graphs – and algorithms that explored those symbols (search trees, rule-based inference, etc.). The field also developed programming languages suited to symbolic AI, notably LISP (created by John McCarthy in 1958) which became the lingua franca of AI research for decades.

However, even as these programs achieved isolated successes, AI researchers were discovering how difficult seemingly simple real-world problems were. For instance, understanding natural language in an unconstrained conversation, or recognizing objects in a photo, proved vastly more challenging than expected – there were simply too many possibilities and ambiguities for brute-force or straightforward symbolic rules to handle. Early optimists had underestimated the complexity of what later came to be called common sense knowledge – the vast collection of trivial facts and assumptions about the world that humans use to understand context. Some researchers, like Marvin Minsky and Seymour Papert at MIT, responded by advocating the micro-worlds strategy (focus on tiny domains like SHRDLU’s blocks), hoping to build AI piece by piece.

By the late 1960s, symbolic AI had demonstrated that computers can play restricted intellectual games and solve formal problems. But it was also reaching a ceiling in more general tasks. This period ended with both a sense of accomplishment and the dawning realization that human-level AI was far harder than anticipated. Many grandiose predictions remained unfulfilled, and this would soon lead into the field’s first major setback – the so-called “AI winter.”

The early 1970s brought a cold dose of reality to AI research, leading to what is now called the First AI Winter – a period of reduced funding and interest in artificial intelligence. The exuberant promises of the 1950s and 1960s had not materialized as expected. The challenges of AI proved deeper than optimistic pioneers assumed, and several critiques and failures culminated in a pullback by government sponsors.

One catalyst was the 1973 Lighthill Report in the United Kingdom. Sir James Lighthill, commissioned by the British government to evaluate the state of AI research, delivered a highly critical assessment. He concluded that AI had failed to achieve its “grandiose objectives”, especially in areas like machine translation, understanding spoken language, and general problem-solving. Lighthill argued that apart from some isolated tasks, AI programs had hit complexity barriers (he specifically cited the “combinatorial explosion” – the rapid growth of possibilities that overwhelms naive algorithms). His report essentially said that further open-ended research in AI was unlikely to be worthwhile. The immediate effect in the UK was the dismantling of AI research programs; funding was slashed, and AI work was marginalized.

Similarly, in the United States, DARPA (the main funder of AI through the 1960s) grew disillusioned with the lack of short-term progress. One famous failure was in machine translation: early AI had tackled translating text from one language to another (especially Russian to English in the Cold War context) with optimistic starts, but by the mid-1960s the ALPAC report concluded that machine translation hadn’t met expectations and recommended stopping funding in that area. This effectively killed U. S. government support for machine translation for decades. By 1974, the U. S. Congress was questioning spending on AI in general, and DARPA cut back substantially on its undirected, basic AI research funding. Resources were reallocated to more immediately practical computer science projects.

Another blow came from within AI research itself: the limitations of techniques were becoming evident. In 1969, Marvin Minsky and Seymour Papert published Perceptrons, a book proving mathematical limits of the then-popular simple neural networks (single-layer perceptrons). They showed that such networks couldn’t learn certain basic patterns (like the XOR function), which cast doubt on the future of neural network approaches. As a result, research on neural nets virtually stopped for a decade (from 1970 to early 1980s). Although their critique didn’t apply to multi-layer networks (which were not well understood then), the book contributed to a general sense that previous approaches had hit a dead end. Symbolic AI still dominated, but it, too, faced a crisis: programs that worked on toy problems collapsed in real-world complexity. For example, a chess program could play chess, but attempts to build a robot to navigate a room or a program to understand unrestricted English text were nowhere near success. Commonsense reasoning – enabling AI to deal with the open-ended unpredictability of real life – was an unsolved mountain.

All these factors led to a wintery chill: between roughly 1974 and 1980, AI funding and enthusiasm dropped. Many university researchers moved on to other fields or refocused on narrower sub-problems. The phrase “AI winter” reflects the idea of a harsh season in which growth stagnates. In the U. S., AI had to rebrand in some cases (researchers got funding under labels like “pattern recognition” or “computational studies” instead of AI). In the U. K., complete project closures occurred. Conferences and publications in AI saw reduced activity.

It’s important to note that AI progress did not stop entirely in this period. Some researchers kept working and achieved results in specialized areas (like expert systems, which we’ll discuss next, had their roots in the 1970s). But the perceptionfrom outside was that AI had overpromised and underdelivered. The lesson learned was a sobering one: intelligence is multifaceted and extremely complex. Tasks humans find effortless (vision, language, commonsense reasoning) turned out to be fiendishly difficult for machines, often harder than tasks humans find difficult (like chess or calculus). This realization reshaped strategies in AI research.

In summary, the first AI winter (mid-1970s) was a period of retrenchment. Funding agencies demanded more accountability and practical results. AI researchers became more modest in claims (for a while) and started focusing on narrower, more achievable goals rather than human-level general intelligence. The winter was a setback, but not the end of AI. By the early 1980s, new approaches and successes would thaw the field and lead to a revival, proving that the winter was part of a cycle rather than a permanent ice age.

AI’s spring returned in the 1980s with the rise of expert systems and renewed government and commercial investment. This period is sometimes called the “Knowledge Engineering” era of AI. Instead of trying to build general problem-solvers, researchers turned to knowledge-based systems: programs that encapsulate expert-level knowledge in specific domains. The bet was that by focusing on narrow expertise – say medical diagnosis or geological analysis – AI could deliver useful, real-world applications. This strategy paid off for a time, leading to successful deployments in industry and a surge of optimism often referred to as the “AI boom” of the 1980s.

The prototype for expert systems was MYCIN, developed at Stanford in the mid-1970s (though it became widely known in the early 80s). MYCIN was designed to assist doctors in diagnosing and treating blood infections. It worked by applying a knowledge base of about 450 rules crafted from interviews with human experts (infectious disease specialists). For example, a rule might be: “If the infection is bacterial and the gram stain is positive cocci in clusters, and it’s hospital-acquired, then there’s suggestive evidence (confidence 0.7) that the organism is Staphylococcus.” MYCIN would ask the user (doctor) a series of questions about the patient’s symptoms and lab results, then infer the likely bacteria causing the infection and recommend antibiotics, complete with dosage. In tests, MYCIN’s performance was comparable to expert physicians for the specific task of bacteremia diagnosis. Importantly, MYCIN also explained its reasoning (by showing which rules it applied), which helped doctors trust its recommendations. This was a rule-based expert system, and its success demonstrated that AI could augment human decision-making in specialized fields.

Following MYCIN’s lead, many other expert systems emerged: DENDRAL (chemistry, for identifying organic molecules from mass spectrometry data), PROSPECTOR (geology, for identifying potential ore deposits), XCON (also known as R1, at Digital Equipment Corporation, for configuring computer systems). XCON, in particular, was a flagship industrial expert system. Developed in the late 1970s at Carnegie Mellon and deployed in the early 80s, XCON helped configure orders of VAX computers (DEC’s minicomputers) by automatically choosing compatible components based on a large set of rules. It saved DEC an estimated millions of dollars by reducing configuration errors and speeding up the assembly process. By 1985, a survey found that corporations were spending over a billion dollars per year on AI, largely on building or buying expert systems. Entire companies (e. g., Teknowledge, Intellicorp) were founded to develop expert system software and tools, and the market for special AI hardware (like Lisp machines from companies Symbolics and Lisp Machines Inc.) boomed to support these AI applications.

Expert systems worked well for certain types of problems. They restricted themselves to a narrow domain, which avoided the need for broad commonsense knowledge. Within that domain, they relied on human-crafted knowledge – essentially, the intelligence was in the expert-provided rules. This approach was feasible because experts in fields like medicine or engineering could articulate at least some of their decision rules, which knowledge engineers then encoded. The resulting systems were essentially large collections of IF-THEN rules with an inference engine (often using backward chaining or forward chaining logic) to apply them.

This approach was dubbed the knowledge revolution in AI: the idea that feeding the computer enough specialized knowledge was more important than new fancy algorithms. AI research focus shifted to knowledge representation, rule-based inference, and handling uncertainty (since expert rules often had to deal with probabilities or heuristics – for instance, MYCIN’s rules carried confidence scores and combined them for conclusions).

Government funding also resurged. Notably, in 1981 Japan announced the Fifth Generation Computer Systems (FGCS) project, a grand 10-year effort to build massively parallel computing hardware and advanced AI software (with an emphasis on logic programming and Prolog) to achieve breakthroughs in intelligent systems. This initiative, backed by Japan’s Ministry of International Trade and Industry, sparked concern in the U. S. and Europe that they might fall behind in AI. In response, the U. S. launched projects like MCC (Microelectronics and Computer Technology Corporation)and DARPA’s Strategic Computing Program (focused on AI for military applications, such as autonomous vehicles and battle management systems). The U. K. initiated the Alvey Programme to fund computer science including AI. This influx of funds in the 1980s was akin to an AI arms race and gave the field significant resources.

By the mid-1980s, AI was riding high. Publications and conferences proliferated. The term “artificial intelligence” was being used in business magazines, and companies proudly advertised AI-powered products. However, as the decade waned, cracks appeared in the expert system paradigm. These systems had some inherent issues:

Knowledge acquisition bottleneck: Extracting the knowledge from experts and converting it into rules was slow and labor-intensive. Experts weren’t always able to state their implicit reasoning, and knowledge engineers became a necessary but scarce resource.

Brittleness: Expert systems, while effective in the scenarios they were built for, often failed badly outside their narrow domain or when confronted with novel situations. They had no commonsense understanding or ability to recognize when a question fell outside their expertise. As one critique put it, they were “brittle” – capable of high performance on standard cases but making grotesque mistakes on edge cases.

Maintenance: As rules accumulated (some systems had thousands of rules), they could conflict or interact in unforeseen ways. Maintaining and updating a large rulebase became very difficult. XCON, for example, eventually grew so complex that modifying it was a headache – each new rule could have side effects.

Inability to learn: These systems did not improve by themselves. If the domain changed or new knowledge emerged, human knowledge engineers had to update the rules. They lacked any learning mechanism to adapt.

By 1987–88, the commercial AI market collapsed. The over-enthusiasm led to an oversupply of AI startups and products that didn’t meet lofty expectations. Companies realized that maintaining expert systems was costly and that some had limited flexibility. Additionally, the specialized AI hardware market (Lisp machines) crashed as cheaper, faster general-purpose workstations (and PCs) could run AI software, removing the need for expensive proprietary machines.

Consequently, by the late 1980s, many AI companies folded and investors pulled back. This marked the onset of the Second AI Winter (circa 1988–1993). Just as in the 1970s, hype had gotten ahead of reality, and disappointment set in. Still, unlike the first winter, the field now had a larger community and some entrenched successful technologies (for example, certain expert systems remained in use, and research continued in universities). But broadly, AI as a buzzword lost some shine again.

In retrospect, the expert systems era was a mixed success. It delivered genuine useful applications – proving that AI could provide value in constrained settings – and it advanced our methods for handling knowledge and uncertainty (Bayesian reasoning was re-discovered in AI in the late 80s as a solution to some of expert systems’ probabilistic reasoning needs). It also taught the lesson that true intelligence requires more than a bag of isolated rules; flexibility and learning were missing. As the 1990s began, some researchers started to pivot back to data-driven and statistical approaches, laying the groundwork for the next resurgence of AI via machine learning.

Chastened by the limitations of hand-crafted expert systems, the AI community in the 1990s increasingly turned to machine learning and statistical methods. Instead of programming all the rules explicitly, researchers began to let computers learn from data. This shift was facilitated by several factors: the increasing availability of digital data, improvements in computing power, and theoretical advances in algorithms. The 1990s saw AI techniques quietly embed themselves in many applications, often under different names (like “data mining” or “pattern recognition”), and achieve successes that set the stage for the big breakthroughs of later decades.

One thread was the revival of neural networks. After the long hiatus following Minsky and Papert’s criticism, a new generation of researchers in the mid-1980s (such as Geoffrey Hinton, David Rumelhart, and Yann LeCun) rediscovered and improved the concept of multi-layer neural networks. They utilized a training algorithm called backpropagation, which efficiently adjusted the weights in a multi-layer network to minimize error. By the late 1980s, backpropagation had been successfully applied to problems like reading handwritten text (LeCun’s network at Bell Labs could recognize ZIP code digits) and simple speech recognition tasks. In the 1990s, these multi-layer perceptrons – now generally called neural networks – gained popularity for pattern recognition. Although these networks were still relatively small by today’s standards, they marked the return of learning from examples as a viable AI technique. The idea of connectionism(brain-inspired computation) proved its worth gradually, fulfilling some of Rosenblatt’s earlier optimistic predictions as computing power grew.

Parallel to neural nets, probabilistic models became prominent. AI researchers embraced Bayesian statistics and probabilistic reasoning to manage uncertainty and learn from data. In 1988, Judea Pearl’s book on Probabilistic Reasoning in Intelligent Systems introduced Bayesian Networks – graphical models that encode probabilistic relationships among variables. Throughout the 90s, Bayesian networks were applied to diagnostics, forecasting, and other areas requiring reasoning under uncertainty. They offered a formal way to combine prior knowledge with observed data, and algorithms were developed for learning these networks from data. This was a stark departure from purely deterministic rule-based systems, reflecting a more realistic approach to knowledge (few things are absolutely certain; reasoning is often about likelihoods).

Another development was the conceptualization of intelligent agents. In the mid-90s, a new framing in AI (championed by Russell and Norvig’s textbook Artificial Intelligence: A Modern Approach, 1995) viewed AI systems as agents that perceive their environment and take actions to maximize some notion of success. This agent-oriented view helped unify subfields and emphasized the importance of the environment and goals, not just isolated problem-solvers.

In practice, the 1990s were rich with AI applications, though they weren’t always labeled “AI” publicly (due to lingering stigma from the AI winter). Yet, behind the scenes, AI solved many difficult problems and entered everyday tech. For example:

Speech recognition systems improved. In 1990, Dragon Dictate was released as the first consumer speech recognition product (using hidden Markov models, a statistical ML technique). By the late 90s, basic dictation software and telephone voice menu systems using speech recognition became available.

Handwriting recognition (like the recognition of handwritten addresses by postal services, or Apple’s Newton PDA attempting handwriting input) came into use.

Computer vision made strides in constrained domains (e. g., face detection in images started to be possible by late 90s with algorithms like Viola-Jones arriving in 2001, but groundwork was laid in the 90s).

Robotics saw progress in navigation and planning. In 1997, Carnegie Mellon’s NavLab project produced a semi-autonomous vehicle that drove across America (with human intervention only occasionally) – a precursor to self-driving car research.

Planning and scheduling AI algorithms were adopted by NASA and airlines. For instance, NASA’s Deep Space 1 probe in 1998 carried an AI software called Remote Agent that autonomously planned operations for the spacecraft for a short period, becoming the first AI system to control a spacecraft without real-time human input.

E-commerce and the web triggered new AI applications: recommendation algorithms (pioneered by sites like Amazon using collaborative filtering in the late 90s) and search engines. Google’s founding in 1998, while rooted in an algorithm (PageRank) not traditionally “AI,” soon incorporated AI techniques to improve search results. The vast data of the internet provided a playground for statistical learning methods.

One of the most public AI milestones of the 90s was in games: IBM’s Deep Blue supercomputer defeated world chess champion Garry Kasparov in May. This victory was significant historically – it was the first time a reigning chess world champion lost a match to a computer under tournament. Deep Blue achieved this largely through brute-force search and domain-specific optimizations (evaluating 200 million positions per second) rather than novel AI algorithms. Nonetheless, it benefited from decades of AI research in game-playing and heuristic search. Kasparov’s loss was a symbolic moment, demonstrating that for a well-defined intellectual task like chess, machines had surpassed the best humans. It sparked widespread media discussion on the growing power of computers and presaged similar feats in other games later (like Go, which would fall to AI in 2016).

Throughout the 1990s, many AI achievements were quietly subsumed into “mainstream” computer science and products. As one observer noted, AI’s greatest innovations were often not marketed as AI – they became just part of the toolset of software engineering. For example, neural networks were used in finance to detect credit card fraud, but that was just “fraud detection software.” Expert system ideas were merged into decision support systems. This integration meant that even though hype was lower, AI was making real impact in many fields from logistics (routing and scheduling algorithms) to medicine (computer-aided diagnosis). Researchers in the 90s, sometimes deliberately, avoided the term AI and spoke of “machine learning”, “knowledge discovery”, or “intelligent systems” to distance themselves from earlier cycles of hype.

By the end of the 1990s, the stage was set for a more data-centric AI revolution. The field had embraced the need for learning from data. Datasets were growing thanks to the internet and digitization of everything. Computers were several orders of magnitude faster than in the 80s, and specialized hardware like GPUs (developed for graphics) were on the horizon to accelerate number-crunching. Theoretical advances in algorithms (like support vector machines in 1995, a powerful new ML method) gave new tools to practitioners. AI was becoming as much about math and statistics as about logic and philosophy. This quiet transformation laid the groundwork for the explosive resurgence of AI in the 21st century, particularly through what came to be known as deep learning.

The 2010s marked a dramatic revolution in artificial intelligence with the ascendancy of deep learning. This approach, which uses large multi-layer neural networks trained on massive datasets, led to breakthroughs in speech recognition, computer vision, game-playing, and many other areas. AI went from niche successes to achieving superhuman performance in several domains, and AI systems became part of everyday life for millions of people (think voice assistants, image tagging on social media, recommendation algorithms, etc.). This period’s progress was so rapid and transformative that it’s often referred to as the era of “AI’s deep learning revolution.”

A key milestone that signaled the new era was in 2012: a neural network known as AlexNet achieved a stunning victory in the ImageNet competition, a yearly contest for image recognition. ImageNet was a project that gathered over a million labeled images in 1000 categories – an unprecedented scale of data for vision. AlexNet, designed by Alex Krizhevsky with colleagues under Professor Geoffrey Hinton, was a deep convolutional neural network (CNN) with about 8 layers of artificial neurons. On September 30, 2012, AlexNet won the ImageNet challenge by a wide margin, achieving around 85% top-5 accuracy, over 10 percentage points better than the next best competitor (which was not a deep neural network). This was an astonishing leap: until then, progress in image recognition had been incremental, and many researchers doubted neural nets’ scalability. AlexNet’s success, powered by GPUs (graphics processing units) to crunch the numbers and by a large dataset to learn from, showed the world that deep learning was not just a theoretical curiosity – it was a practical, game-changing technique. After 2012, the floodgates opened. In the next couple of years, essentially every team in image recognition, and soon in speech recognition, switched to deep neural networks to stay competitive.

Deep learning differed from earlier AI in that it automatically learned features from raw data. For example, in vision, older systems required human engineers to define features (like edges, corners, textures) for the classifier to use. A deep CNN like AlexNet learned low-level features (edges) in early layers, higher-level patterns (faces, object parts) in deeper layers, all on its own through the training. This ability to learn hierarchical representations proved enormously powerful. Similarly, in speech recognition, deep networks replaced older models by learning directly from waveforms or spectrograms, leading to dramatic accuracy improvements. By around 2015, speech recognition by machines (for well-defined tasks and in good conditions) reached or surpassed human-level performance. This enabled the proliferation of voice-based assistants (Apple’s Siri, Google Assistant, Amazon Alexa, etc.) which, thanks to cloud-based deep learning models, could understand spoken requests fairly reliably.

One of the most public triumphs of deep learning (combined with other AI techniques) was AlphaGo. In 2016, DeepMind (a Google subsidiary) developed AlphaGo, a system that combined deep neural networks with Monte Carlo tree search to play the board game Go. Go had long been viewed as a supreme challenge for AI – far more complex than chess, with an astronomically larger search space and a need for intuition and pattern recognition. Yet AlphaGo defeated Lee Sedol, one of the world’s top Go champions, 4-1 in a match in March. This event was as significant for Go as Deep Blue had been for chess, but in some ways more stunning: experts had thought Go mastery by AI was at least a decade. AlphaGo achieved it by learning from millions of examples (both human games and self-play games) and capturing strategies that no human had conceived. It demonstrated that deep learning could capture something akin to intuition at superhuman levels when given enough data and compute. AlphaGo’s victory was celebrated as an AI milestone and showed the potential of combining deep learning with reinforcement learning (learning by trial-and-error with rewards).

Throughout the 2010s, deep learning pushed forward the state-of-the-art in many fields:

Computer vision: Beyond image classification, deep networks began to outperform humans at tasks like object detection and face recognition. We saw technologies like real-time face recognition deployed (with both positive uses, like organizing your photo collection, and controversial ones, like mass surveillance). Medical imaging diagnosis by deep learning (e. g., identifying tumors in scans) became comparable to experts in certain cases.

Natural language processing: Initially, recurrent neural networks and LSTMs (Long Short-Term Memory networks) improved speech recognition and machine translation. Systems like Google’s Neural Machine Translation (launched 2016) delivered far more fluent translations than previous phrase-based statistical methods. Later in the decade, new architectures called Transformers (introduced in 2017) revolutionized language tasks even further, leading to powerful language models we’ll discuss in the next section.

Robotics: While robotics hardware evolves slowly, deep learning began to be used for perception and control. Robots and autonomous vehicles leveraged deep neural nets to interpret camera data (for example, self-driving car prototypes using CNNs to recognize pedestrians, lanes, traffic signs). Deep reinforcement learning (like DeepMind’s Deep Q Network in 2015) showed robots (or simulated agents) could learn complex control policies – famously, a deep RL agent learned to play Atari video games at superhuman level directly from pixel inputs, with no knowledge of the game rules.

Generative models: Deep learning also enabled generative AI – models that don’t just classify or predict, but create new data. Early examples were autoencoders and variational autoencoders that could generate blurry images, but the big leap came with Generative Adversarial Networks (GANs) (introduced by Ian Goodfellow in 2014). GANs pit two networks against each other (a generator vs. a discriminator) and produced startlingly realistic images of faces, landscapes, etc. By the late 2010s, GAN-generated images and videos (deepfakes) were realistic enough to cause both excitement and concern.

The cumulative effect of all this was that AI started to directly touch ordinary people’s lives in the 2010s. Smartphones got smarter with AI-based features like face unlock, predictive typing, and augmented reality filters. Social media used AI to curate feeds and tag photos. E-commerce used it for recommendations and customer service chatbots. Cars included driver-assist features (some using AI for lane-keeping or emergency braking). It was a quiet infiltration of AI into countless products and services, often just marketed as “smart” features.

Another hallmark of the deep learning revolution was the scaling of data and compute. It became clear that more data and larger models tended to yield better performance – a phenomenon sometimes summarized as “there’s no data like more data.” Companies like Google, Facebook, and later OpenAI pushed to train ever-bigger neural networks on ever-larger datasets, using specialized hardware (GPUs, TPUs). Researchers talked about the “unreasonable effectiveness of data” and how simple algorithms fed by huge data could outperform more sophisticated algorithms with less data. This empirical finding drove AI toward a more engineering-heavy, big-data discipline, sometimes diverging from cognitive science roots.

By the end of the 2010s, AI had firmly moved from academic labs to industry. Tech giants invested heavily in AI research, hiring many top academics. Startups in AI surged, focusing on things like autonomous driving, AI healthcare, fintech, and more. There was also a realization that these powerful AI capabilities bring new challenges – ethical, social, and technical. But before turning to those, we should highlight the next phase that began around 2020: the era of foundation models and generative AI, which built directly on the deep learning advances of the 2010s.

Entering the 2020s, AI witnessed another step-change with the emergence of extremely large models often called foundation models – AI systems trained on broad swaths of data that can be adapted to a wide range of tasks. Alongside this came a boom in generative AI, where models aren’t just analyzing data but creating new content (text, images, music, code, etc.) that is often indistinguishable from human-created content. These developments captured the public imagination like never before, bringing AI into headlines and daily conversations.

Foundation models refer to models that serve as a general platform (a “foundation”) for many applications. They are characterized by being trained on massive datasets, often using self-supervised learning, and having an enormous number of parameters (often in the billions). Because of their scale and broad training, they can be fine-tuned or prompted to perform tasks they weren’t explicitly trained for, making them very versatile. One prominent category of foundation models is Large Language Models (LLMs). Early examples include OpenAI’s GPT series (GPT-2 in 2019, GPT-3 in 2020) and Google’s BERT (2018). GPT-3, for instance, has 175 billion parameters and was trained on hundreds of billions of words from the internet. Astonishingly, GPT-3 can generate coherent paragraphs of text, translate languages, answer questions, write simple code, and more – all without task-specific training, simply by being prompted with a few examples or instructions. This “one model, many tasks” capability was a striking demonstration of scalability: just by making models bigger and training on more data, unexpected new abilities emerged (often referred to as “emergent behaviors”). These LLMs became a kind of universal engine for language, usable in countless applications from chatbots to writing assistance.

The term “foundation model” was coined around 2021 by the Stanford Center for Research on Foundation Models. It captures not only language models but models in other domains: for example, image generation models like DALL-E (OpenAI, 2021) and Stable Diffusion (2022) are foundation models for images – trained on huge sets of images and captions, they learn a general ability to create images from text descriptions. Similarly, there are foundation models for code (like OpenAI’s Codex, which can generate code from natural language), for music, for multi-modal tasks combining text and images, etc.. The distinctive quality is that these models are general-purpose – they can be adapted to many specific tasks with minimal additional training.

The advent of generative AI deserves special attention. In 2022, a watershed moment occurred with the public release of systems like DALL-E 2, Stable Diffusion, and Midjourney that can create high-quality images from any text prompt. Suddenly, anyone could type “a castle in the clouds, in Van Gogh’s style” and see a never-before-seen image visualized in seconds. This democratized content creation in a way unimaginable a few years prior. Artists and designers found new tools for inspiration; at the same time, concerns arose about deepfakes and the ease of generating misinformative visuals.

On the text side, OpenAI’s ChatGPT, released to the public in late 2022, truly brought generative AI into mainstream awareness. ChatGPT is a conversational agent based on an improved GPT-3.5 model, fine-tuned for dialogue with techniques like Reinforcement Learning from Human Feedback (RLHF) to make it follow instructions and behave more safely. Within five days of its release, ChatGPT had over one million – an unprecedented adoption rate. Users were amazed that ChatGPT could draft essays, answer complex questions, write code, hold philosophical conversations, and more. It felt qualitatively different from earlier chatbots: far more knowledgeable and coherent. By mid-2023, ChatGPT (and similar models like Google’s Bard or Anthropic’s Claude) had reportedly reached hundreds of millions of users. These chatbots showcased how a foundation model (a large language model) could be packaged into a user-friendly format and impact a broad user base. They also raised new challenges: these models sometimes “hallucinate” – meaning they can produce incorrect or fictional information with a confident tone. They also can reflect biases present in their training data or produce inappropriate content if not carefully constrained.

Nevertheless, the capabilities of foundation models continued to advance quickly. OpenAI’s GPT-4 (2023) further improved accuracy and even accepted image inputs, showing a level of general problem-solving that led some to speculate about early “sparks” of artificial general intelligence in limited domains. Other organizations like Meta, Google, and open-source communities released their own large models (LLaMA, PaLM, etc.), leading to a thriving ecosystem of LLMs.

Generative AI expanded beyond text and images: by 2023, there were models for generating video (rudimentary but improving), generating 3D models, composing music or voice, and more. The concept of multimodal AI – models that handle multiple types of input/output (e. g., text + image) – gained traction, enabling applications like generating detailed image captions or having a single AI that can both see and talk.

The widespread deployment of these technologies has had a massive societal impact. They offer powerful tools: writers can get assistance drafting and editing, programmers can have AI suggest code (GitHub’s Copilot, powered by OpenAI Codex, is an example), customer service can be partly automated with chatbots, educators and students find new ways to teach and learn (though also worry about plagiarism or misuse), and creative industries explore human-AI collaboration in art and design. At the same time, they raise ethical and regulatory issues. Concerns about misuse of generative AI for disinformation, the effect on jobs (e. g., will copywriters, illustrators, or even software developers be augmented or replaced by AI?), and questions of intellectual property (training data often includes copyrighted text and images) have all become hot topics.

By the mid-2020s, investment in AI boomed as businesses raced to integrate these new AI capabilities. AI became a central focus for tech strategy across many industries. Yet, with great power came great responsibility – leading figures in tech began calling for regulatory frameworks. Governments started drafting laws (the European Union’s AI Act, discussions in the US and elsewhere) to ensure ethical AI development, transparency, and safety. There’s also been a resurgence of interest in AI safety and alignment research: making sure advanced AI systems act in accordance with human values and do not behave in harmful or unpredictable ways.

In summary, the early 2020s have been defined by AI models of unprecedented scale and versatility. Foundation models and generative AI have pushed the frontier of what machines can create and understand. This era has brought AI to the masses, not just as consumers but even as co-creators. It is an exciting time, with AI showing abilities that often surprise even experts – but it’s also a time of reckoning with the profound implications of having such powerful “alien” intellects at our fingertips. The conversation has shifted from “Can we get AI to work?” to “AI works – now how do we manage its impact on society?”

As of today (mid-2020s), artificial intelligence has reached remarkable capabilities across a variety of domains. It’s important to both celebrate what AI can do and to realistically acknowledge what it cannot do (yet), in order to have a clear picture of the state of the art. AI systems now excel at specific tasks – often surpassing human performance in narrow arenas – but they also exhibit significant limitations such as lack of common sense, problems with transparency, and unpredictable errors.

First, let’s survey current capabilities of AI:

Language understanding and generation: AI models can now engage in fluent conversations, answer questions on vast ranges of topics, summarize documents, translate between many languages, and even write stories or code. For example, a modern large language model can take a complex question about, say, quantum physics or historical events and produce a detailed, coherent answer. It can also generate well-structured essays or creative fiction. This would have been science fiction a couple of decades ago. AI-powered translation systems make cross-lingual communication instant. Voice assistants utilize these language models to hold dialogues with users, showing that speech interfaces can be quite natural.

Image and video analysis: AI vision systems can identify objects and people in photos with high accuracy. They can describe what’s in an image in sentences (image captioning). Facial recognition algorithms have become very accurate (though not without biases or controversies in use). AI can flag tumors in medical scans that doctors might miss, or classify skin lesions from photos as malignant or benign with expert-level accuracy. In video, AI can track objects or people, detect anomalies (useful in security or manufacturing quality control), and even generate short video clips.

Generative abilities: As discussed, AI can generate content: lifelike images from text prompts, human-sounding speech from text, music in various styles, and even video game levels or virtual environments procedurally. This has opened up new possibilities for rapid prototyping and creativity. A single person with an AI tool can create illustrations or animations that might have required a whole team and studio previously. In software development, AI code generators can auto-complete functions or suggest solutions, accelerating programming tasks.

Decision support and prediction: In fields like finance, AI models predict stock movements or credit risks (though not infallibly). In logistics, they forecast demand and optimize supply chains. In healthcare, they assist in diagnosing diseases or predicting patient outcomes from health records. In science, AI helps in protein folding predictions (DeepMind’s AlphaFold made headlines by predicting 3D structures of proteins with high accuracy, a boon for biology) and in sifting through large data sets for new discoveries.

Autonomous systems: While we don’t yet have fully autonomous cars everywhere, AI has made great strides in this area. Self-driving car prototypes can handle a large subset of driving scenarios; even consumer cars now commonly have AI-based advanced driver-assistance systems (ADAS) like lane keeping, adaptive cruise control, and automatic emergency braking that have reduced accidents. Drones and robots with AI can navigate through environments – from warehouse robots efficiently moving goods around, to delivery robots on sidewalks, to Mars rovers planning routes on another planet.

Games and simulations: AI can learn to play extremely complex games. Beyond Go and chess, AI systems mastered games like StarCraft II (DeepMind’s AlphaStar achieved Grandmaster level) and Poker (AI like Libratus and Pluribus beat top human players in no-limit poker, which involves hidden information and bluffing). These are not just parlor tricks; the techniques used (deep reinforcement learning, game-theoretic reasoning) are being adapted to real-world problems like strategy planning and operations research.

Despite these impressive capabilities, current AI systems also have significant limitations:

Lack of true general intelligence and common sense: AI is still largely narrow. A system that plays Go cannot converse about politics; a language model that writes a good essay might not reliably navigate a physical space. AI systems don’t possess an integrated understanding of the world. They often lack what humans call common sense. For example, a language model might tell you that pouring water into a bottle turned upside down is a good way to fill it – a blatantly nonsensical answer stemming from a lack of physical understanding. They can overlook obvious real-world constraints or basic facts unless those were explicitly present in their training data. This brittleness means AI can make mistakes no human child would, like asserting confidently that “an elephant fits in a refrigerator” if prompted in a tricky way.

Tendency to produce errors or “hallucinations”: Especially with generative models, AI can output incorrect information that looks perfectly plausible. For instance, a chatbot might cite a non-existent article to answer a question or make up a fictitious statistic. It doesn’t truly know in the human sense; it patterns its responses on training data, which can lead to confabulations when faced with unfamiliar queries. This unreliability is a serious issue if AI is used in high-stakes settings. Even when AI’s overall accuracy is high, its lack of guarantee of truthfulness means human oversight is needed. Analysts in 2023 have noted that reducing these “hallucinations” is a major challenge in making AI more.

Bias and fairness issues: AI systems learn from data that often contains human biases or reflects societal inequalities. As a result, they can inadvertently perpetuate or amplify those. For example, a facial recognition system might have higher error rates for darker-skinned faces if its training data was skewed, leading to potential discriminatory. Language models might output stereotypes or offensive content if prompted in certain ways, because such patterns existed in their internet training data. Addressing bias is an ongoing concern: it requires careful dataset curation, algorithmic fairness techniques, and sometimes manual fine-tuning to mitigate harmful outputs.

Explainability and transparency: Many modern AI models, particularly deep neural networks, operate as “black boxes.” They can have millions or billions of parameters with complex interactions, making it very hard to understand why a model made a particular decision. For instance, if a deep learning model denies a loan application, it’s difficult to extract a human-readable explanation for that decision, which is problematic for accountability. This lack of interpretability is a significant limitation in domains where trust and verification are crucial (like healthcare or finance). There is a field of XAI (Explainable AI) trying to develop methods to peek inside these black boxes or to provide approximate explanations, but it remains a challenge.

Data and compute hunger: The most capable AI systems require enormous amounts of data and computational resources to train. Training a large language model or an image generator can cost millions of dollars in cloud compute and needs datasets that are sometimes terabytes in size. This means only a few tech companies or well-funded organizations can afford to develop such frontier models, raising concerns about centralization of AI power. It also means these models are energy-intensive; the environmental footprint of AI, especially during training, is non-negligible.

Robustness: AI models can be surprisingly fragile under adversarial conditions. A slight perturbation to an image (imperceptible to humans) can cause a classifier to totally misidentify an object (e. g., seeing a turtle as a rifle, famously). Language models can be tricked with carefully crafted inputs to produce wrong or disallowed answers. Ensuring AI systems are robust against malicious manipulation or odd inputs is an active area of research.

Autonomy and reasoning: While AI can automate narrow tasks, creating a system with broad autonomous reasoning and planning capabilities (often termed artificial general intelligence, AGI) remains an unsolved problem. Present AI doesn’t have motivations, self-awareness, or understanding beyond pattern processing. It cannot truly set its own goals or reflect on its reasoning the way humans do. Every action is ultimately a result of its training and prompts given by humans. Complex multi-step reasoning is still hard – for example, asking a model to solve a complicated puzzle or plan a multi-stage project will often exceed its reliable capabilities. There are efforts to imbue models with better reasoning (some approaches chain steps of reasoning explicitly, known as chain-of-thought prompting), but they are preliminary.

In essence, today’s AI can be thought of as a savant in some areas and clueless in others. It’s extremely powerful as a tool to augment human abilities – for instance, an AI system can rapidly sift through millions of documents to find relevant information, something no human could do in reasonable time. It can serve as an endless creative brainstorm partner. It can control complex systems or predict outcomes by finding subtle patterns. However, AI currently operates without understanding or consciousness; it doesn’t know what it’s doing in a human sense. That means when conditions change or problems are outside the scope of its training, it may fail in unpredictable ways.

Recognizing these limitations is crucial. It guides how we deploy AI (with human oversight for critical decisions), how we regulate it, and where to focus research. Many researchers are actively working on addressing these issues: adding common sense knowledge bases to AI, developing better training methods that incorporate logical reasoning, creating hybrid models that combine neural networks with explicit symbolic logic to get the best of both, and improving techniques to detect and correct AI biases and errors.

AI today is both impressive and imperfect. It’s a rapidly evolving field, so some limitations may be mitigated with time, while new challenges will emerge as we push the technology further. In the final section, we’ll look ahead to possible futures and the broader impact of AI on society.

As we look to the future, artificial intelligence stands as a transformative force with immense potential – both positive and negative. The trajectory of AI development in the coming years and decades could lead to profound changes in how we live, work, and relate to one another. In this concluding section, we consider several aspects: the path toward more general AI, the impacts on jobs and the economy, ethical and safety considerations, and the societal adaptations that might be needed to ensure AI benefits humanity.

Towards Artificial General Intelligence (AGI): One key question is whether AI will progress from narrow expert systems to general intelligence – a system with the flexibility and understanding comparable to a human mind (or beyond). Some researchers believe that current deep learning scaling will eventually yield AGI, while others suspect new paradigms will be required. At present, no AI can match the full breadth of human cognitive abilities, but the frontier models are showing more generalization than ever before. If AGI is achieved, it could be a double-edged sword: it might enable solutions to problems that are currently intractable (from curing diseases to climate modeling to scientific discoveries at superhuman speed), or if misaligned with human values, it could pose unprecedented risks. This is why there is increasing attention on AI alignment – making sure any highly autonomous AI system’s goals are aligned with human ethical principles and well-being.

Automation and the Future of Work: In the near term, AI is poised to significantly impact the workforce. Automation driven by AI could displace certain jobs, even as it creates new ones. Routine and repetitive tasks (in manufacturing, logistics, data processing) have been automated for some time, but AI now encroaches on tasks that require judgement, perception, or even creativity. For example, AI customer service agents can handle many inquiries that human agents once did. In software development, AI coding assistants reduce the need for writing boilerplate code. In journalism, AI can draft basic news reports (like sports recaps or financial summaries). Truck and taxi drivers may see their roles changed by self-driving vehicle technology. That said, history shows that technology tends to create new roles even as it renders others obsolete – but the transition can be disruptive and requires workforce retraining and social safety nets. Many experts predict that rather than wholesale job disappearance, we will see job transformation: AI will take over tasks, not entire jobs, and workers will need to upskill to work alongside AI tools (a human + AI team can often be more effective than either alone). Nonetheless, there are legitimate concerns about specific sectors being hard-hit and increasing inequality if the benefits of AI accrue only to those with capital or specialized skills.

Economic and Social Impact: If productivity dramatically rises due to AI, it could lead to greater wealth – but how that wealth is distributed is a policy choice. Some have proposed ideas like universal basic income as a buffer if automation reduces the need for human labor in aggregate. AI could also exacerbate economic divides: countries or companies that lead in AI might leap ahead of competitors, concentrating power. Internationally, there is a kind of AI arms race (both economic and military) among major powers – the U. S., China, and the EU investing heavily to not fall behind in AI capability. This competition could spur rapid progress but also geopolitical tension.

Ethical and Privacy Concerns: AI’s ability to analyze and generate data raises significant concerns about privacy and rights. Facial recognition cameras linked to AI can enable mass surveillance on an unprecedented scale, as seen in some parts of the world – a boon for law enforcement and security perhaps, but a worry for civil liberties and potential abuse by authoritarian regimes. Generative AI can produce fake videos or audio (deepfakes) that are very difficult to discern from real – this undermines trust in digital information and can be weaponized for misinformation or fraud. There will be a growing need for verification technologies and perhaps legal regulations to manage this (e. g., requiring AI-generated content to be watermarked or identified). Bias in AI decision systems, if not corrected, can lead to systematic unfair treatment of certain groups in lending, hiring, policing, etc., effectively encoding prejudice into automated systems. Society will likely demand accountability: AI systems, especially in critical uses, may need to be audited for fairness and accuracy. Regulations like the proposed EU AI Act aim to categorize uses of AI by risk and impose requirements like transparency and human oversight for high-risk applications (e. g., AI in healthcare, law enforcement, or education).

Human-AI Interaction and Society: As AI becomes more present, we’ll have to redefine our relationship with machines. Already, people chat with AI assistants and sometimes form attachments or anthropomorphize them. This will raise new psychological and social questions: How do human relationships change when AI companions or assistants are ubiquitous? Should AI systems have some rights or at least be designed to signal they are not human? We may see new forms of art, culture, and expression where AI is collaborator or even creator. Education will also adapt: knowing facts is less crucial when AI can fetch information instantly; instead, education might focus more on critical thinking, learning to work with AI, and uniquely human skills like empathy and ethical judgement.

AI Safety and Existential Risk: At the more speculative end, some thinkers warn of existential threats from AI – the idea that a superintelligent AI, if not properly controlled, could pose a risk to human survival or autonomy. While such scenarios remain theoretical, they are being taken more seriously as AI capabilities scale. Open letters and petitions have been circulated by experts calling for caution, such as a moratorium on training the most powerful models until safety measures are in place. Research in AI safety and alignment aims to develop methods to ensure advanced AI systems reliably do what humans intend and do not develop unintended, dangerous behaviors. This involves technical work (like designing algorithms that can explain their reasoning, learn objectives in a human-friendly way, or be monitored and corrected) and governance work (international agreements, monitoring of AI development, perhaps licensing regimes for training large models, similar to how nuclear technology is controlled).

In the optimistic future scenario, AI helps solve humanity’s grand challenges. It accelerates scientific research – for example, helping to develop new drugs and treatments in record time, or optimizing energy usage and material design to aid environmental sustainability. It takes over drudgery and dangerous jobs, freeing people to pursue more creative and meaningful work or leisure. AI tutors could provide personalized education to anyone in the world, raising global knowledge and skill levels. AI-driven analysis could improve policy-making by simulating outcomes of decisions and finding hidden patterns in data that humans miss. In healthcare, AI might continuously monitor patients via wearables, catching problems early and suggesting preventive measures, leading to longer and healthier lives.

In the pessimistic scenario, AI could amplify existing problems. Job displacement could lead to unrest if not managed, deepfakes and misinformation could erode trust in information to the point of social chaos or severe polarization, oppressive regimes could use AI for unprecedented surveillance and control over citizens, and the concentration of AI in the hands of a few could create a new kind of oligarchy. In a very dark scenario, a misaligned AGI could act in ways that are harmful – for instance, if tasked with an overly simplistic objective like “maximize clicks on our website,” a super-intelligent system might find unintended and harmful ways to do so (a trivial example: hacking into devices to force people to click, or flooding the environment with stimuli). These examples underscore why aligning AI goals with human values is crucial.

Likely, the reality will be somewhere in between these extremes, shaped by the actions we take today. The challenge and opportunity for society is to steer AI development responsibly. This means interdisciplinary collaboration: not just engineers and computer scientists, but ethicists, sociologists, lawyers, economists, and the public at large should have input into how AI is used. It means updating our education systems to prepare people for an AI-augmented world. It could mean new social contracts: perhaps shorter work weeks if AI productivity allows, or new forms of compensation if traditional jobs become fewer.

In conclusion, artificial intelligence has come a long way from ancient myths and rudimentary automata to today’s complex algorithms and learning machines. Each era built on the last: formal logic and computing enabled the first AI programs; early successes and failures guided researchers to new approaches; the machine learning revolution harnessed data and computational power; and the current wave of large models shows both the promise and perils of scaling up AI. The history of AI is a story of human ingenuity – our desire to replicate our own intellect in our tools – and it continues to unfold. As we stand on the cusp of AI systems that may profoundly alter our future, it is up to us to ensure that this technology is developed and used in alignment with humanity’s best interests. The next chapters of AI’s story will be written not just by engineers and scientists, but by all of us, through our choices, policies, and values in the face of transformative innovation.

Published by NOMOTO MEDIA

Support independent work

Help fund what comes next.

NOMOTO MEDIA publishes essays, investigations, fiction, audio, and films without a paywall. If the work is valuable to you, help support the next piece.