The History of Calculus Who found it, who fought over it, and why your notation looks the way it does

☰ Contents Search
Preset
Details

Appendix E

Worked mathematics, re-derived and checked in code

I​‌‍‌‌‍‍‌‍‌‍‍‌‌‍‌‍‌‍‍‌‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‍‌‍‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‌‌‌‌‍‍‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‍‌‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‌‍‌‍‌‌‌‍‍‌‍‌‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‌‌‍‌‌‌‌‍‍‌‍‌‌‍‌‍‍‍‌‌‍‍‌‍‍‍‌‍‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‍‍‌‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‌‌‍‍‌‌‌‍‌‌‌‌‌‌‍‌‌‌‌‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‍‍‌‌‌‍‍‌‍‍‍‌‍‌‍‌‍‍‌‍‍‌‌‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‍‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‌‍‌‌‍‌‍‍‌‌‍‍‍‌‍‍‌‍‌‌‍‌‍‍‌‍‍‍‌‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‍‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌​ re-derived every result below and checked it numerically or symbolically in Python (sympy, fractions.Fraction, decimal.Decimal at 60-digit precision). The script and its raw output are in verify/math_verification.txt. Eighteen results, eighteen passes.

One check failed on the first run, and the log keeps the failure rather than hiding it. The test was written wrong: it compared a 400-term partial sum of a geometric series to its exact limit and expected exact equality. The mathematics was right and the test was wrong. The corrected version verifies Archimedes' closed form symbolically and then takes the limit. Both versions are in the log, which is the point of keeping one.

Numeric columns are right-aligned to a consistent number of decimal places. Any column that can go negative shows an explicit sign.


E.1 Archimedes: quadrature of the parabola (c. 250 BCE)

Archimedes proves the finite identity and then argues by contradiction. He never writes the third line.

k=0n(14)k=4313(14)nk=06(14)k=54614096=1.3332519531limn[4313(14)n]=43
Read this equation in words

The sum from k equals zero to n of one quarter to the k equals four thirds minus one third times one quarter to the n. At n equals six, the sum is five thousand four hundred sixty-one over four thousand ninety-six, which is one point three three three two five one nine five three one. The limit as n grows without bound is four thirds.

Check: sympy confirms the closed form of the finite sum equals Archimedes' expression identically, and that both the infinite sum and the limit equal 4/3.


E.2 Zu Chongzhi's bounds on π (5th century CE)

3.1415926<π<3.1415927355113=3.1415929203π=3.1415926535|355113π|=2.668×107
Read this equation in words

Pi lies between three point one four one five nine two six and three point one four one five nine two seven. Three fifty-five over one thirteen is three point one four one five nine two nine two zero three, and so on. Pi is three point one four one five nine two six five three five, and so on. The difference between them is two point six six eight times ten to the negative seventh.

Check: both bounds hold; 355/113 agrees with π to 6 decimal places. It is also the best rational approximation to π with denominator below about 16,000.


E.3 Mādhava's π, decoded from verse (c. 1400)

T​‌‍‌‌‍‍‌‍‌‍‍‌‌‍‌‍‌‍‍‌‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‍‌‍‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‌‌‌‌‍‍‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‍‌‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‌‍‌‍‌‌‌‍‍‌‍‌‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‌‌‍‌‌‌‌‍‍‌‍‌‌‍‌‍‍‍‌‌‍‍‌‍‍‍‌‍‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‍‍‌‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‌‌‍‍‌‌‌‍‌‌‌‌‌‌‍‌‌‌‌‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‍‍‌‌‌‍‍‌‍‍‍‌‍‌‍‌‍‍‌‍‍‌‌‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‍‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‌‍‌‌‍‌‍‍‌‌‍‍‍‌‍‍‌‍‌‌‍‌‍‍‌‍‍‍‌‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‍‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌​he bhūta-saṅkhyā word-numerals decode to 2827433388233.

π28274333882339×1011=3.14159265359222π=3.14159265358979difference=2.429×1012
Read this equation in words

Pi is approximately the thirteen-digit number two eight two seven four three three three eight eight two three three, over nine times ten to the eleventh, which is three point one four one five nine two six five three five nine two two two, and so on. Pi is three point one four one five nine two six five three five eight nine seven nine, and so on. The difference is two point four two nine times ten to the negative twelfth.

Check: agrees with π to 11 decimal places.


E.4 The Mādhava-Leibniz series, and why it is unusable raw

π4=113+1517+
Read this equation in words

Pi over four equals one, minus one third, plus one fifth, minus one seventh, and so on.

E.4 The Mādhava-Leibniz series, and why it is unusable raw: 4 rows.
Terms Partial sum Absolute
error
Correct
decimals
103.04183961899.98 × 10⁻²0
1003.13159290361.00 × 10⁻²1
10003.14059265381.00 × 10⁻³2
100003.14149265361.00 × 10⁻⁴3

Ten thousand terms buys three decimal places. This is exactly why Mādhava built correction terms, and why the Yuktibhāṣā states an exactness criterion for them rather than guessing.


E.5 Oresme: the harmonic series diverges (c. 1350)

Group the terms into blocks that each sum to at least one half.

k=11k=1+12+(13+14)1/2+(15++18)1/2+H2m1+m2
Read this equation in words

The sum from k equals one to infinity of one over k equals one, plus one half, plus a group of one third plus one quarter, which is at least one half, plus a group of one fifth through one eighth, which is also at least one half, and so on. So the harmonic sum through two to the m is at least one plus m over two.

E.5 Oresme: the harmonic series diverges (c. 1350): 4 rows.
n H(n) Oresme's
lower bound
Bound
holds
163.3807293.000000yes
2566.1243455.000000yes
40968.8951047.000000yes
1638410.2813078.000000yes

Check: the bound holds at every power of two up to 2^14. The series diverges, and slowly: passing 10 takes more than 12,000 terms.


E.6 Fermat's adequality: maximize x(ax) (c. 1636)

f(x)=x(ax)f(x+e)f(x)=ae2xee2f(x+e)f(x)e=a2xesuppress e:0=a2xx=a2
Read this equation in words

f of x equals x times, a minus x. Then f of x plus e, minus f of x, equals a e minus two x e minus e squared. Dividing by e gives a minus two x minus e. Suppress e: zero equals a minus two x, so x equals a over two.

Check: sympy expands the quotient to a - e - 2*x and solves, at e = 0, to [a/2].

The pedagogical point: the last step is a suppression, not a substitution of zero. You divided by e, which requires e ≠ 0, and then you delete it, which requires e = 0. That is precisely the crack Berkeley drove a wedge into in 1734.


E.7 Cavalieri and Wallis: the area under x^n on [0, 1]

01xndx=limN1Nk=1N(kN)n=1n+1
Read this equation in words

The integral from zero to one of x to the n, d x, equals the limit as capital N grows without bound of one over N times the sum from k equals one to N of, k over N, to the n. That limit is one over, n plus one.

E.7 Cavalieri and Wallis: the area under x^n on [0, 1]: 4 rows.
n Riemann sum,
N = 200,000
1/(n+1) Absolute
error
10.5000025000.5000000002.50 × 10⁻⁶
20.3333358330.3333333332.50 × 10⁻⁶
30.2500025000.2500000002.50 × 10⁻⁶
40.2000025000.2000000002.50 × 10⁻⁶

E.8 Newton's binomial series (1665)

(1+x)1/2=1+12x18x2+116x35128x4+7256x51.2 (12 terms)=1.0954451150347271.2 (exact)=1.095445115010332error=2.44×1011
Read this equation in words

One plus x, all to the one-half power, equals one, plus one half x, minus one eighth x squared, plus one sixteenth x cubed, minus five over one twenty-eight x to the fourth, plus seven over two fifty-six x to the fifth, and so on. Twelve terms give the square root of one point two as one point zero nine five four four five one one five zero three four seven two seven; the exact value begins one point zero nine five four four five one one five zero one zero three three two. The error is two point four four times ten to the negative eleventh.

Check: coefficients confirmed symbolically against sympy's series expansion.


E.9 Leibniz's product-rule error, 11 November 1675

He guessed d(xy) = dx·dy, tested it on one example, got the right answer by accident, and refuted himself the same afternoon.

(x+dx)(y+dy)xy=xdy+ydx+dxdy
Read this equation in words

x plus d x, times y plus d y, minus x y, equals x d y, plus y d x, plus d x times d y.

W​‌‍‌‌‍‍‌‍‌‍‍‌‌‍‌‍‌‍‍‌‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‍‌‍‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‌‌‌‌‍‍‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‍‌‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‌‍‌‍‌‌‌‍‍‌‍‌‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‌‌‍‌‌‌‌‍‍‌‍‌‌‍‌‍‍‍‌‌‍‍‌‍‍‍‌‍‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‍‍‌‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‌‌‍‍‌‌‌‍‌‌‌‌‌‌‍‌‌‌‌‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‍‍‌‌‌‍‍‌‍‍‍‌‍‌‍‌‍‍‌‍‍‌‌‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‍‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‌‍‌‌‍‌‍‍‌‌‍‍‍‌‍‍‌‍‌‌‍‌‍‍‌‍‍‍‌‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‍‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌​ith x = 3, y = 5, dx = dy = 1/1000:

E.9 Leibniz's product-rule error, 11 November 1675: 4 rows.
Quantity Exact Decimal Verdict
Actual change in xy8001/10000000.008001000the truth
dx times dy1/10000000.000001000Leibniz's first guess: smaller than the actual change by a factor of 8001
x dy + y dx1/1250.008000000correct to first order
Actual minus (x dy + y dx)1/10000000.000001000exactly dx dy, the discarded second-order term

Check: the identity actual(xdy+ydx)=dxdy holds exactly in rational arithmetic.

Why this is the best teaching artifact in the subject: the wrong answer and the right answer differ by exactly the term you are entitled to throw away, and Leibniz left both on the page.


E.10 Wallis's product for π (1656)

π2=n=14n24n21=212343456567
Read this equation in words

Pi over two equals the infinite product of four n squared over, four n squared minus one: that is two over one, times two over three, times four over three, times four over five, times six over five, times six over seven, and so on.

E.10 Wallis's product for π (1656): 1 rows.
Factors 2 x product Absolute
error
200003.14155338493.93 × 10⁻⁵

E.11 Euler and the Basel problem (1734/35)

Euler treats sin(x)/x as an infinite-degree polynomial and factors it by its roots at x = ±nπ.

sinxx=1x23!+x45!=n=1(1x2n2π2)coefficient of x2:16=1π2n=11n2n=11n2=π26=1.6449340668
Read this equation in words

Sine x over x equals one, minus x squared over three factorial, plus x to the fourth over five factorial, and so on. It also equals the infinite product over n of, one minus x squared over n squared pi squared. Matching the coefficient of x squared: negative one sixth equals negative one over pi squared, times the sum of one over n squared. So the sum from n equals one to infinity of one over n squared equals pi squared over six, which is one point six four four nine three four zero six six eight, and so on.

Check: sympy extracts the x2 coefficient of sin(x)/x as exactly 1/6, and (1/6)π2 simplifies to π2/6, matching zeta(2). The partial sum to 200,000 terms gives 1.6449290669, error 5.00 × 10⁻⁶.

The step that is not rigorous is assuming an infinite-degree polynomial factors like a finite one. Making that legitimate is a large part of what Chapter 10 is about.


E.12 Lagrange's program fails: a smooth function that is not its Taylor series

f(x)={e1/x2x00x=0f(n)(0)=0for all n0n=0f(n)(0)n!xn=0f(x)0for x0
Read this equation in words

f of x equals e to the negative one over x squared when x is not zero, and zero when x equals zero. Every derivative of f at zero is zero, for all n at least zero. So the Taylor series, the sum of the n-th derivative at zero over n factorial, times x to the n, is identically zero. But f of x is not zero for any x other than zero.

This is why Théorie des fonctions analytiques (1797) could not do what Lagrange wanted.


E.13 Ibn al-Haytham's summation identity (c. 1000)

(n+1)i=1nik=i=1nik+1+p=1ni=1pik
Read this equation in words

n plus one, times the sum from i equals one to n of i to the k, equals the sum from i equals one to n of i to the k plus one, plus, for each p from one to n, the sum from i equals one to p of i to the k, all added together.

He states it only for n = 4 and k = 1, 2, 3, proving each by induction on n (Source 61, Katz, p. 166).

Check: verified exactly in integer arithmetic for k = 1, 2, 3, 4 and n = 4, 5, 9. His closed form for the fourth powers also matches the standard one:

i=1ni4=n(n+1)(2n+1)(3n2+3n1)30=n55+n42+n33n30
Read this equation in words

The sum from i equals one to n of i to the fourth equals n, times n plus one, times two n plus one, times three n squared plus three n minus one, all over thirty. Expanded, that is n to the fifth over five, plus n to the fourth over two, plus n cubed over three, minus n over thirty.


E.14 The paraboloid is eight fifteenths of its cylinder (c. 1000)

Rotate the parabola x = ky² about the line x = kb². Slice into discs; the disc at height y has radius kb² - ky².

Vparaboloid=0bπ(kb2ky2)2dy=πk2b5(123+15)=815πk2b5Vcylinder=π(kb2)2b=πk2b5VparaboloidVcylinder=815
Read this equation in words

The volume of the paraboloid is the integral from zero to b of pi times, k b squared minus k y squared, squared, d y. That works out to pi k squared b to the fifth, times, one minus two thirds plus one fifth, which is eight fifteenths pi k squared b to the fifth. The cylinder's volume is pi times, k b squared, squared, times b, which is pi k squared b to the fifth. So the paraboloid fills exactly eight fifteenths of its cylinder.

C​‌‍‌‌‍‍‌‍‌‍‍‌‌‍‌‍‌‍‍‌‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‍‌‍‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‌‌‌‌‍‍‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‍‌‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‌‍‌‍‌‌‌‍‍‌‍‌‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‌‌‍‌‌‌‌‍‍‌‍‌‌‍‌‍‍‍‌‌‍‍‌‍‍‍‌‍‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‍‍‌‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‌‌‍‍‌‌‌‍‌‌‌‌‌‌‍‌‌‌‌‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‍‍‌‌‌‍‍‌‍‍‍‌‍‌‍‌‍‍‌‍‍‌‌‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‍‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‌‍‌‌‍‌‍‍‌‌‍‍‍‌‍‍‌‍‌‌‍‌‍‍‌‍‍‍‌‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‍‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌​heck: sympy returns exactly 8/15 (Source 61, Katz, p. 168).

What makes it a calculus result and not just an answer: he brackets the solid between eight fifteenths of the cylinder and eight fifteenths of the cylinder less its top slice, then makes the top slice as small as he likes. That is the same squeeze Fermat and Roberval use in 1636.


E.15 The cycloid's area, which Galileo weighed and threw away (1599)

For one arch of x = a(t - sin t), y = a(1 - cos t):

A=02πydxdtdt=02πa2(1cost)2dt=3πa2Aπa2=3
Read this equation in words

The area A equals the integral from zero to two pi of y times d x over d t, d t, which becomes the integral from zero to two pi of a squared times, one minus cosine t, squared, d t. That evaluates to three pi a squared. So the area over pi a squared is exactly three.

Check: exactly 3. Galileo's balance gave him "about three times as heavy" in 1599 and he discarded the result, believing the ratio incommensurable (Source 64, Whitman, p. 310). The experiment was right and the intuition about it was wrong.


E.16 Wren's rectification of the cycloid (1658)

L=02π(dxdt)2+(dydt)2dt=02πa22costdt=02π2asint2dt(sint20 on [0,2π])=8aL2a=4
Read this equation in words

The arc length L is the integral from zero to two pi of the square root of, d x over d t squared plus d y over d t squared, d t. That becomes the integral of a times the square root of, two minus two cosine t, which simplifies to the integral of two a times sine of half t, since sine of half t stays nonnegative from zero to two pi. The integral evaluates to eight a. So L over two a is exactly four.

Check: exactly 8a, confirmed symbolically and numerically (8.00000000000000 for a = 1).

Why it mattered in 1658: a curved length exactly equal to a straight one, four times the diameter of the generating circle, with no π in it. Curved lengths were widely believed not to be expressible by straight ones at all (Source 64, Whitman, p. 314).


E.17 Sarasa 1649: hyperbolic areas behave as logarithms

S​‌‍‌‌‍‍‌‍‌‍‍‌‌‍‌‍‌‍‍‌‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‍‌‍‌‍‍‍‌‍‍‌‌‌‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‌‌‌‌‍‍‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‍‌‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‌‌‌‌‌‍‍‌‌‍‌‌‌‍‍‌‍‍‌‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‌‍‌‍‌‌‌‍‍‌‍‌‌‌‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‌‌‍‌‌‌‌‍‍‌‍‌‌‍‌‍‍‍‌‌‍‍‌‍‍‍‌‍‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‍‍‌‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‌‌‍‍‌‌‌‍‌‌‌‌‌‌‍‌‌‌‌‍‍‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‍‍‌‌‌‍‍‌‍‍‍‌‍‌‍‌‍‍‌‍‍‌‌‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‍‌‌‍‌‌‌‌‌‌‌‍‌‍‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‍‍‍‍‌‍‍‍‌‌‍‌‌‍‍‌‍‌‌‍‌‍‍‌‌‍‍‍‌‍‍‌‍‌‌‍‌‍‍‌‍‍‍‌‌‍‍‌‌‌‌‍‌‍‍‌‍‍‌‌‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‍‌‍‌‍‌‍‍‍‌‌‍‌‌‍‍‍‌‌‍‍‌‍‍‌‌‍‌‍‌‌‍‌‌‌‌‌‌‍‍‌‌‌‍‍‌‍‍‌‍‍‍‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌‌‍‍‌‌‍‌‍‌‍‍‌‍‍‍‌‌‍‍‍‌‍‌‌​arasa's claim is that in place of the logarithms of magnitudes in continued proportion, you may take the areas under a hyperbola. Here is why that works.

abdxx=lnbaabscissas in continued proportion a,ar,ar2,ar3,aardxx=arar2dxx=ar2ar3dxx=lnraarndxx=nlnr
Read this equation in words

The integral from a to b of d x over x equals the natural log of b over a. Take abscissas in continued proportion: a, then a r, then a r squared, then a r cubed, and so on. The integral over each step, from a to a r, from a r to a r squared, from a r squared to a r cubed, is the same every time: the natural log of r. Across n steps, the integral from a to a r to the n equals n times the natural log of r.

Equal ratios give equal areas. Geometric progression in x becomes arithmetic progression in area, which is precisely the behaviour of the numbers 6, 7, 8, 9, 10 that Sarasa starts from.

The functional equation, which is what makes it a logarithm rather than merely log-like:

1uvdxx=1udxx+1vdxx
Read this equation in words

The integral from one to u v of d x over x equals the integral from one to u of d x over x, plus the integral from one to v of d x over x.

Products become sums. That is the property the whole seventeenth century wanted logarithms for, and here it falls out of an area.

His numerical case, checked with common ratio 2:

E.17 Sarasa 1649: hyperbolic areas behave as logarithms: 6 rows.
Magnitude Area from 1 Difference from
the previous row
10.000000-
20.6931470.693147
41.3862940.693147
82.0794420.693147
162.7725890.693147
323.4657360.693147

Check: every difference is identical, at ln 2 = 0.693147. The magnitudes double; the areas step up by a constant. That is a logarithm, found as an area, in 1649, and stated in a pamphlet defending somebody else's book (Source 66, Sarasa).


E.18 The chain rule, run backwards, is backpropagation

y=sin(x2+3x)dydx=dydududx,u=x2+3x=cos(x2+3x)(2x+3)at x=0.7:dydx=3.7474403185central difference, h=106=3.7474403187difference=1.28×1010
Read this equation in words

y equals sine of, x squared plus three x. d y over d x equals d y over d u, times d u over d x, with u equal to x squared plus three x. That gives cosine of, x squared plus three x, times, two x plus three. At x equals zero point seven, d y over d x is negative three point seven four seven four four zero three one eight five. A central difference with h equal to ten to the negative sixth gives negative three point seven four seven four four zero three one eight seven; the two differ by one point two eight times ten to the negative tenth.

The cost argument, which is the whole reason reverse mode exists: for a function with m inputs and n outputs, forward mode costs roughly m passes and reverse mode costs roughly n. Training a neural network has n = 1 (the loss) and m in the billions. Reverse mode costs one pass regardless of parameter count, at roughly two to three times a single forward evaluation (Source 436, Baydin, Pearlmutter, Radul and Siskind).