The History of Calculus Who found it, who fought over it, and why your notation looks the way it does

☰ Contents Search
Preset
Details

Chapter 12

Calculus now

9 minute read

The chain rule is what trains artificial intelligence

T​‌‍‌‌‍‍‌‍‌‍‍‌‌‍‌‍‌‍‍‌‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‍‌‍‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‌‌‌‌‍‍‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‍‌‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‌‍‌‍‌‌‌‍‍‌‍‌‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‌‌‍‌‌‌‌‍‍‌‍‌‌‍‌‍‍‍‌‌‍‍‌‍‍‍‌‍‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‍‍‌‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‌‌‍‍‌‌‌‍‌‌‌‌‌‌‍‌‌‌‌‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‍‍‌‌‌‍‍‌‍‍‍‌‍‌‍‌‍‍‌‍‍‌‌‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‍‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‌‍‌‌‍‌‍‍‌‌‍‍‍‌‍‍‌‍‌‌‍‌‍‍‌‍‍‍‌‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‍‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌​his is the strongest "why does this matter to me" hook in the document for a reader in 2026. Two peer-reviewed surveys confirm it (Source 436, Baydin, Pearlmutter, Radul and Siskind; Source 437, Schmidhuber).

Every chatbot, image generator, and recommendation system on Earth is trained by running the chain rule backwards. The rule taught in about week three of a derivatives course:

dydx=dydududxfor y=sin(x2+3x):dydx=cos(x2+3x)(2x+3)at x=0.7:dydx=3.7474403185central difference check=3.7474403187
Read this equation in words

d y over d x equals d y over d u, times d u over d x. For y equals sine of, x squared plus three x: d y over d x equals cosine of, x squared plus three x, times, two x plus three. At x equals zero point seven, that is negative three point seven four seven four four zero three one eight five; a central difference check gives negative three point seven four seven four four zero three one eight seven.

(Appendix E.18.) That rule is not an exam technique. It is the algorithm that trains artificial intelligence. It is called backpropagation, it is reverse-mode automatic differentiation, and it is the chain rule.

Guess before you read on

A modern network has one output, the loss, and inputs numbering in the billions. Training it means getting the derivative of that one output with respect to every one of those billions. Say how many passes through the network the backward direction costs.

I have a guess

One. Full stop. About two to three times the cost of a single forward evaluation, and it does not matter whether the network holds a thousand parameters or a trillion (Source 436, Baydin, Pearlmutter, Radul and Siskind).

If you guessed one pass per input, that is the honest arithmetic for the forward direction, and it is why nobody trains anything that way. Every model you have used was trained in the direction that costs one pass.

W​‌‍‌‌‍‍‌‍‌‍‍‌‌‍‌‍‌‍‍‌‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‍‌‍‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‌‌‌‌‍‍‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‍‌‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‌‍‌‍‌‌‌‍‍‌‍‌‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‌‌‍‌‌‌‌‍‍‌‍‌‌‍‌‍‍‍‌‌‍‍‌‍‍‍‌‍‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‍‍‌‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‌‌‍‍‌‌‌‍‌‌‌‌‌‌‍‌‌‌‌‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‍‍‌‌‌‍‍‌‍‍‍‌‍‌‍‌‍‍‌‍‍‌‌‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‍‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‌‍‌‌‍‌‍‍‌‌‍‍‍‌‍‍‌‍‌‌‍‌‍‍‌‍‍‍‌‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‍‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌​hy it has to run backwards is a counting argument a student can check. A modern network has one output (the loss) and inputs numbering in the billions (the parameters). Forward mode costs one pass per input. Reverse mode costs one pass, full stop, at roughly two to three times the cost of a single forward evaluation, regardless of parameter count (Source 436, Baydin, Pearlmutter, Radul and Siskind).

The history is a lesson in how credit works:

And in 1969 Minsky and Papert published a book about what single-layer networks cannot do. The tools to fix the problem had already been published. The field mostly stopped for a decade anyway (Source 437, Schmidhuber). Jürgen Schmidhuber calls this "surprising in hindsight".

Before you read on: the mathematics that fixed the problem was already published when the book came out. Why did the field stop anyway?

Why a correct book can still stop a field

Perceptrons was not wrong. Single-layer networks cannot compute XOR, and Minsky and Papert proved it. The book was right and the conclusion people drew from it was not.

The result was about one layer. Readers heard it as being about neural networks. The tools for more than one layer existed, in work that had already been published and that almost nobody in the field had read (Source 437, Schmidhuber).

So the field spent a decade acting on a true theorem plus a false generalization, and the generalization traveled because the theorem was rigorous and the book was well written.

A correct proof of a narrow claim is the most persuasive way to spread a wide one. That is worth remembering every time you see a study quoted one sentence at a time.

T​‌‍‌‌‍‍‌‍‌‍‍‌‌‍‌‍‌‍‍‌‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‍‌‍‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‌‌‌‌‍‍‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‍‌‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‌‍‌‍‌‌‌‍‍‌‍‌‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‌‌‍‌‌‌‌‍‍‌‍‌‌‍‌‍‍‍‌‌‍‍‌‍‍‍‌‍‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‍‍‌‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‌‌‍‍‌‌‌‍‌‌‌‌‌‌‍‌‌‌‌‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‍‍‌‌‌‍‍‌‍‍‍‌‍‌‍‌‍‍‌‍‍‌‌‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‍‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‌‍‌‌‍‌‍‍‌‌‍‍‍‌‍‍‌‍‌‌‍‌‍‍‌‍‍‍‌‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‍‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌​he line from Leibniz to ChatGPT is five steps and unbroken: Leibniz writes the chain rule in 1676; Euler builds the calculus of variations in 1744; Kelley and Bryson use it to steer rockets around 1960; Linnainmaa writes it down as an algorithm about rounding error in 1970; and in 2026 it trains every large language model (Source 437, Schmidhuber). Five steps, three hundred and fifty years, one rule.

Two more modern threads worth a paragraph each

Laurent Schwartz and distributions. Physicists had been using Dirac's delta: an object that is zero everywhere except one point, infinite there, with area 1. No such function exists. They used it anyway because it worked. Schwartz's answer was not to ban it but to enlarge the definition of what can be differentiated until the delta became a respectable citizen. That is the exact opposite of the nineteenth-century strategy of tightening definitions until the monsters were excluded (Source 433, OConnor and Robertson).

He worked it out lying in the dark, in bed, in occupied France, under a false name, and ran next door in the middle of the night to tell his neighbor. The neighbor's reply is the best one-line summary of the theory anyone has produced: "Now we'll never again have functions without derivatives" (Source 433, OConnor and Robertson). He later lost his professorship for saying the French Army tortured people, and was proved right.

Bourbaki. The most influential mathematics textbook of the twentieth century was written by a man who did not exist. "Nicolas Bourbaki" was a group of young French mathematicians annoyed at the calculus textbook they had to teach from. Their solution was to rewrite mathematics from the ground up. Their style, definition then theorem then proof, with no diagrams, no motivation, and no history, is why so many university textbooks feel the way they do (Source 434, OConnor and Robertson). It is also why a book like the one you are building has to argue for its existence.

And the statistics connection: the Radon-Nikodym derivative is why "probability density" means anything. Every time a statistics course writes p(x) dx, the reason that expression is legitimate is a theorem finished in 1930. And Stone-Weierstrass is why it is not crazy to believe a complicated function can be approximated by simpler ones, which is the entire premise of machine learning (Source 435, OConnor and Robertson).

How calculus reached teenagers

This thread is partially sourced, and the gaps are named.

Verified. The College Board states plainly that "The Advanced Placement Program was established in 1955" (S448, Access to Excellence, 2001, p. 1). In 1976 the entire AP Program, every subject nationwide, was about 100,000 examinations taken by roughly 76,000 students, and Calculus AB and BC were already both in place by 1977 (Source 493, Coleman). In 2024, Calculus AB drew 278,657 exams and Calculus BC drew 148,191, a total of 426,848, out of 5,744,259 exams taken by 3,079,134 students (Source 446, College Board; Source 447, College Board).

How calculus reached teenagers: 3 rows.
Year AP Calculus exams Total AP exams, all subjects
1976 not sourced 100,000
2004 not sourced 1,887,770
2024 426,848 5,744,259

A nice statistical aside for a class: the BC exam is harder and has the better score distribution, because its students are self-selected. That is a concrete illustration of selection effects (Source 447, College Board).

N​‌‍‌‌‍‍‌‍‌‍‍‌‌‍‌‍‌‍‍‌‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‍‌‍‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‌‌‌‌‍‍‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‍‌‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‌‍‌‍‌‌‌‍‍‌‍‌‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‌‌‍‌‌‌‌‍‍‌‍‌‌‍‌‍‍‍‌‌‍‍‌‍‍‍‌‍‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‍‍‌‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‌‌‍‍‌‌‌‍‌‌‌‌‌‌‍‌‌‌‌‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‍‍‌‌‌‍‍‌‍‍‍‌‍‌‍‌‍‍‌‍‍‌‌‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‍‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‌‍‌‌‍‌‍‍‌‌‍‍‍‌‍‍‌‍‌‌‍‌‍‍‌‍‍‍‌‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‍‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌​ot sourced, and flagged as such: the Ford Foundation and Kenyon Plan origins, the 1952 report, the 1954 pilot, the 1956 first exams, the figure of 1,229 candidates, the year AP Calculus specifically began, and when and why the AB/BC split appeared. Rothschild (1999) and Schneider (2009) are both non-open-access, AP Central has no history page, and ERIC holds the histories without full text. These are blanks, not estimates, and a number you cannot cite is worse than an empty space.

The best primary source in this thread is public domain and from 1901. John Perry, addressing the British Association at Glasgow on 14 September 1901, complained that:

Some boys of ten years of age study the methods of the differential calculus... (Source 492, Perry, p. 19)

His needling comparison, that ten-year-olds handle differentiation while nineteen-year-olds who have "worked at mathematics" all their lives cannot grasp what a derivative is, is the most quotable line in the whole schools-and-calculus thread (Source 492, Perry). Silvanus P. Thompson, who spoke at that same meeting, published Calculus Made Easy nine years later. Its famous prologue, "What one fool can do, another can", is the same argument in a book for the general reader (Source 492, Perry).

And the first calculus textbook in English (1704) belongs here, uncomfortably. Charles Hayes's A Treatise of Fluxions was written by a director of the Royal African Company, a slave-trading corporation (Source 439, Hayes). That is a true fact and it belongs in a chapter about who gets to write the books. Its first sentence defines an infinitely small quantity, and its first proof shows there are infinitely small quantities inside the infinitely small quantities. Thirty years later Berkeley would call these the ghosts of departed quantities. Two hundred and sixty-two years later Robinson would prove Hayes had been entitled to them all along (Source 439, Hayes).

Hayes's own warning, "this Doctrine may seem hard to most Readers at first", is the oldest sentence in English calculus pedagogy (Source 439, Hayes).

One question before you go

Every statistics course writes p(x) dx, and nobody stops to say why that expression is allowed to mean anything. Which theorem makes it legitimate?

Show the answer

T​‌‍‌‌‍‍‌‍‌‍‍‌‌‍‌‍‌‍‍‌‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‍‌‍‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‌‌‌‌‍‍‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‍‌‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‌‍‌‍‌‌‌‍‍‌‍‌‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‌‌‍‌‌‌‌‍‍‌‍‌‌‍‌‍‍‍‌‌‍‍‌‍‍‍‌‍‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‍‍‌‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‌‌‍‍‌‌‌‍‌‌‌‌‌‌‍‌‌‌‌‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‍‍‌‌‌‍‍‌‍‍‍‌‍‌‍‌‍‍‌‍‍‌‌‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‍‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‌‍‌‌‍‌‍‍‌‌‍‍‍‌‍‍‌‍‌‌‍‌‍‍‌‍‍‍‌‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‍‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌​he Radon-Nikodym derivative, finished in 1930 (Source 435, OConnor and Robertson). It is what turns a probability measure into a density you can integrate, which is the move behind every sentence that says the area under the curve is the probability.

Integrals 3.7 is where you meet that sentence. The reason it is not a fudge is a piece of measure theory that landed two and a half centuries after Leibniz wrote the integral sign.