01
Intellectual precursors
The modern field of artificial intelligence grew from several lines of work that were originally separate: mathematical logic, the theory of computation, statistics, information theory, neuroscience and engineering. The central question was whether processes associated with reasoning, learning or perception could be described precisely enough to be reproduced by machines.
Alan Turing's 1936 work on computability provided a formal model of mechanical computation. In 1943 Warren McCulloch and Walter Pitts proposed a mathematical model of neurons as simple computational units. Their work did not constitute modern machine learning, but it established an important conceptual connection between neural systems and computation.
In 1950 Turing reframed the question 'Can machines think?' through the imitation game. The importance of the paper lies partly in the methodological shift: instead of attempting to define intelligence in purely philosophical terms, it proposed an operational test and then examined objections concerning consciousness, learning and machine limitations.
02
The birth of a named field
The 1955 Dartmouth proposal is one of the canonical founding documents of AI. McCarthy, Minsky, Rochester and Shannon proposed a summer research project based on the conjecture that aspects of learning and intelligence could be described sufficiently precisely for machines to simulate them.
The proposal named problems that remain central today: language, abstraction, concept formation, problem solving and self-improvement. This is striking because the computational resources available in 1956 were tiny by modern standards. The field therefore began with an ambitious theory of intelligence before it had the hardware, data or algorithms needed to realise that theory.
03
Symbolic AI
Early AI programs represented knowledge explicitly as symbols and transformed those symbols through search, logic and rules. The Logic Theorist and General Problem Solver demonstrated that machines could solve structured problems by manipulating formal representations.
Symbolic methods offered a powerful engineering property: the internal representation of the system was inspectable. Developers could examine the rules, modify them and trace the reasoning process. This made symbolic systems attractive for theorem proving, planning and expert systems.
The limitation was equally important. Real environments contain uncertainty, ambiguity, enormous search spaces and information that is difficult to encode manually. The cost of explicitly specifying the world became a recurring bottleneck.
04
AI winters
AI's early demonstrations often performed well on carefully bounded problems but failed to scale to the complexity of real environments. When research promises repeatedly exceeded practical results, funding and institutional support declined. These periods became known as AI winters.
The most important lesson from the winters is not that AI research stopped being useful. Instead, they demonstrated that capabilities observed in controlled demonstrations do not automatically generalise to open-ended environments. The gap between a prototype and a robust system remains a central engineering problem.
05
Expert systems
During the 1970s and 1980s researchers pursued domain-specific expert systems. Instead of attempting to solve intelligence in general, these systems encoded specialised knowledge using rules and inference procedures. DENDRAL, MYCIN and XCON became important examples.
Expert systems demonstrated that explicit domain knowledge could produce useful commercial and scientific systems. They also revealed the knowledge-acquisition problem: building and maintaining a sufficiently complete rule base was expensive, brittle and difficult to scale.
06
The statistical turn
From the late twentieth century onward, probability, statistics and optimisation increasingly became central to AI. Rather than encoding every decision explicitly, systems could learn parameters from observations and quantify uncertainty.
This transition connected AI more tightly to statistical learning. Bayesian models, support-vector machines, hidden Markov models and later neural networks became increasingly important because they offered mechanisms for dealing with noisy observations and large datasets.
07
Deep learning and foundation models
The deep-learning era emerged from the convergence of larger datasets, specialised hardware and improved optimisation methods. Deep neural networks were able to learn hierarchical representations directly from data rather than depending entirely on manually designed features.
The 2012 ImageNet result of AlexNet demonstrated the practical impact of this combination for visual recognition. Later systems such as AlphaGo showed that learned representations could be combined with search and reinforcement learning to solve difficult structured problems.
The 2017 Transformer architecture was another major transition. By relying on attention rather than recurrence or convolution for sequence transduction, Transformers enabled highly parallel training and became the architectural foundation for many large pretrained models.