English (unofficial) translations of posts at kexue.fm
Source

How Textbooks Should Be Written: What Kind of Mathematics Education Do We Need?

Translated by DeepSeek V4 Pro. Translations can be inaccurate, please refer to the original post for important stuff.

Reposted from: matrix67.com

Note: This article contains many personal views and carries a strong subjective color. Some of these ideas may not necessarily be correct, and there are things I may not even be qualified to say. I simply want to share some of my thoughts with everyone. Please remember to maintain your own opinions. Also, please include this disclaimer when reposting.

I am not a mathematician. I am not even a mathematics major. I am a pure amateur math enthusiast—just more persistent and more obsessed than the average hobbyist. Having been fast-tracked through middle and high school and not majoring in math in university, I have been able to study the mathematical knowledge I am interested in without the goal of passing exams. This has allowed me to maintain such a strong interest in the subject. Since establishing this blog in 2005, whenever I see a stunning conclusion or a beautiful proof, I take the time to record it no matter how busy I am, for fear of forgetting it. However, I am well aware that these clever tricks, while impressive, are hardly the true beauty of mathematics. I fear I have not yet grasped even a fraction of the truly profound and vast ideas of mathematics.

I have told people many times that my life’s ideal is to one day learn all the various branches of mathematics, stand at a high vantage point, overlook the entire field, and truly experience the beauty of mathematics. However, achieving this is very difficult. The greatest difficulty is the lack of a path to learn mathematics. Read textbooks? That is what I want to talk about today—textbooks are extremely unreliable.

I have deep experience with this. In the past two years, I have been doing middle school math tutoring and have developed some of my own views. Mathematics education can be roughly divided into three stages: seeing mountains as mountains and waters as waters; seeing mountains not as mountains and waters not as waters; and finally, seeing mountains as mountains and waters as waters.

==================================

In the earliest stage, math education was simply: teaching you a few theorems, telling you how they are proven, and then asking you to prove some new theorems.

Later, the requirements changed: learning math alone was not enough; one had to use math. Math education rose to a new level: everyone should apply math to life and explain phenomena in life. For a time, whether in textbooks or high school entrance exam questions, everything was filled with math application problems closely related to real life, as if math were truly everywhere you looked. Department stores selling goods, bookstores selling books, farmers tilling land, and workers laying bricks once again flooded textbooks, supplementary materials, and exam papers. In fact, while math can explain life, we don’t actually do it that way. There are too many variables in life; even the most powerful mathematical model cannot account for everything. For the average person, the only place they truly use math is for balancing accounts.

One day, math education will rise to the third level: returning to simplicity, where the truly impressive thing about math is math itself. You will find that those great mathematical ideas and brand-new mathematical theories were not initially motivated by an urgent need to explain some weird phenomenon around us, but by their own inherent beauty. The emergence of linear algebra is largely credited to the mysterious Cramer’s Paradox; the birth of group theory was a product of Galois studying the structure of solutions to polynomials; Euler founded graph theory because of the “pointless” Königsberg Bridge problem which had no practical value; and the appearance of non-Euclidean geometry was entirely due to the charm of the problem itself. What about calculus? It indeed has very wide practical value, and various definitions in physics depend on it; but unfortunately, it is not a “disruptive” mathematical idea.

When seventh-grade textbooks introduce negative numbers, they repeatedly talk about their practical significance, such as altitude, scores, temperature, income and expenditure, etc., turning negative numbers into a tangible existence. In fact, this was not the primary motivation for people to use negative numbers. The value of negative numbers lies in the fact that they can turn subtracting a number into adding a negative number. Many complex additions and subtractions that might even require case-by-case discussions can be unified into a single expression. For example, the “surplus and deficit” problems in elementary school: if 3 apples are given to each person, 8 are left over; if 5 apples are given to each person, 2 are left over. How many people and how many apples are there? The solution is: the difference in leftover apples between the two methods is 6, which is caused by each person getting 2 more apples; therefore, there are 3 people, and one can calculate there are 17 apples. However, what if the problem is changed to “if each person gets 3, there are 8 left over; if each person gets 5, there are 2 short”? The formula above changes; you can’t subtract 2 from 8, you have to add 2. Therefore, elementary school surplus/deficit problems are discussed in three cases: “surplus-deficit,” “surplus-surplus,” and “deficit-deficit.” In fact, if “2 short” is understood as “plus -2,” the problem is exactly the same, and the previous formula applies just as well. The new idea of negative numbers immediately unified the three cases, making their essence identical.

This is the example I always give when teaching negative numbers to seventh graders. This is the true meaning of negative numbers. This is what textbooks should repeatedly emphasize with examples.

Once I saw someone on a forum asking, “What’s the point of group theory?” Someone replied, “Group theory is very interesting, it’s just that textbooks make it boring. For example, how can you teach group theory without mentioning the Rubik’s Cube?” I disagree with this reply. The attraction of mathematics lies not in its applications in life, but in its own beauty. Why not talk about Lagrange’s Theorem? Why not talk about Sylow’s Theorems? For me, what attracts me most to studying a mathematical topic is a series of non-trivial conclusions and their brilliant proofs.

The science fiction story “The Mourner” lists many mathematical theories that went without practical application for a long time, but it didn’t mention an even more extreme example. The crown jewel of mathematics—number theory—had no practical application for 2,000 years; it was the purest mathematics. It wasn’t until the advent of computers, especially modern cryptography, that number theory stepped out of mathematics and into people’s lives for the first time. What supported the research of number theory? It could only be mathematics itself.

When I set geometry problems for middle school children, I try to give general problems: prove that the average length of two sides of a triangle is greater than the median on the third side; prove that the sum of the reciprocals of the three altitudes of a triangle equals the reciprocal of the in-radius, and so on. Even for pure algebra and analytic geometry problems, I can always compose problems with simple descriptions that are highly challenging. How many integer solutions are there where the sum and product of two numbers are equal? What is the equation of the line obtained by reflecting the line y=x across y=2x? While feeling the beauty of the conclusions, they also get excited because they have independently solved a real mathematical problem.

==================================

However, this is not yet the main problem with education. Once, while chatting about the Riemann Hypothesis with a math major, she said she had never heard of it. I was shocked—how could a math major not know the Riemann Hypothesis? I immediately realized this was also thanks to math education. Opening a math textbook, one always finds a complete theoretical system, starting with definitions and followed by proofs, all very logical. But where did these things come from? What detours did mathematicians take in the process of deriving these things? Textbooks mention none of this. Textbooks only ever teach what is right, but never what is wrong. Math exams only ask you to prove a conclusion; they never ask you to refute one.

The 2010 Jiangsu Gaokao math problems were controversial for being “too difficult.” One of the final large problems was as follows: Given that the three side lengths of \triangle ABC are all rational numbers, (1) prove that \cos(A) is rational; (2) prove that for any positive integer n, \cos(nA) is rational. In fact, this is a very beautiful and good problem: simple description, universal problem, interesting conclusion, and clever proof. This is how entrance exam questions should be. However, I feel that if one more small question were added, the problem would be truly perfect: prove or disprove that \sin(A) must be a rational number. Of course, the problem itself is not hard; an equilateral triangle is the simplest counterexample. The key is that refuting a conclusion and finding a counterexample is also a fundamental ability in mathematical research, which is rarely emphasized in secondary math education.

Consequently, when I teach middle school math, every homework problem I assign invariably starts with “Prove or Disprove.” Occasionally, some problems truly require the students to refute them. For example, prove or disprove that two triangles with the same perimeter and the same area are congruent. Different people find different counterexamples—some simple, some complex, some profound, some random. Then, using an entire class period to explain and critique the counterexamples everyone constructed brings far more benefit to the children than simply explaining the problem directly.

==================================

However, I still haven’t reached the most significant problem in math education. Some time ago, I went to a Turing author-translator exchange meeting, during which I had a brief chat with Teacher Liu Jiang. Teacher Liu mentioned a website called Better Explained. He said that the reason people fail to understand the wonder of mathematics is that it isn’t taught well; math could originally be explained more intuitively and popularly.

I very much agree with Teacher Liu Jiang. Let’s take an example. If a student asks, “What is a prime number?” a teacher might say, “A prime number is a number that has no divisors other than 1 and itself.” No, that’s not the answer the student wants. What the student really wants to know is: what is a prime number, really? In fact, prime numbers are the “indivisible” numbers, the basic elements that make up all natural numbers. 12 is composed of two 2s and one 3, just as H_2O is composed of two H atoms and one O atom. Only, unlike the chemical world, the elements of the arithmetic world are infinite. All objects, theorems, and methods within the arithmetic world are composed of these basic elements—that is why prime numbers are so important.

When learning complex numbers in high school, I believe many people wonder: what is an imaginary number? Why should we accept imaginary numbers? How does an imaginary number represent rotation? In fact, people established the theory of complex numbers not because they sometimes needed to handle cases where there is a negative number under a square root, but for the following irresistible reason: if we accept imaginary numbers, then an n-th degree polynomial will have exactly n roots, and the number system suddenly becomes as perfect as a crystal ball. But complex numbers cannot be visually reflected on a number line, not only because real numbers are already complete on the number line, but for another reason: there is no geometric operation that, when performed twice, results in taking the negative. For example, “multiplying by 3” represents expanding the distance of a point on the number line from the origin to three times its original; “3 squared,” which is “multiplying by 3 and then by 3,” is performing the above operation twice, i.e., expanding to 9 times. Similarly, “multiplying by -1” represents flipping the point to the other side of the number line; “-1 squared” will flip the point back again. But how do we represent the operation of “multiplying by i” on the number line? In other words, what operation performed twice can turn 1 into -1? A revolutionary creative answer is: rotate the point 90 degrees around the origin. Rotating 90 degrees twice naturally lands you on the other side of the number line. Exactly—this extends the number line to the entire plane, perfectly solving the problem of where to represent complex numbers. Thus, the multiplication of complex numbers can be explained as scaling plus rotation, and complex numbers themselves naturally have the representation z = r(\cos\theta + \sin\theta i). Following this logic, everything becomes natural. Complex numbers not only have a geometric interpretation but can sometimes handle geometric problems more conveniently.

I have always been interested in linear algebra, so I took a linear algebra course in university, but the harvest was almost zero. The reason was simple: I was expecting a moment of total enlightenment, but after studying for a semester, I still didn’t know what a matrix actually was, why matrix multiplication was defined that way, what it meant for a matrix to be invertible, or what a determinant actually represented.

It wasn’t until I saw a certain webpage today that I saw someone explain the essence of linear algebra in one sentence (which is also the direct reason I finally decided to write this article). I finally found what I had been searching for that whole semester. Just like changing x to 2x, we often need to change (x, y) into something like (2x + y, x - 3y); this is called a linear transformation. Thus, we thought of defining matrix multiplication to represent all linear transformations. Geometrically, moving every point (x, y) on the plane to the position (2x + y, x - 3y) is equivalent to performing a “linear stretch” on the plane.

Matrix Multiplication - Geometry

Matrix multiplication is actually the effect of superimposing multiple linear transformations; it clearly satisfies the associative law but not the commutative law. A matrix with 1s along the main diagonal corresponds to a linear transformation that means “no change,” which is why it is called the identity matrix. Matrix A multiplied by matrix B resulting in the identity matrix means that after performing linear transformation A and then linear transformation B, you return to the original state—no wonder we say matrix B is the inverse of matrix A. The definitions of determinants in textbooks are all sorts of strange, involving recursion or inversions, even with mnemonics to help everyone remember. In fact, the true definition of a determinant is just one sentence: the area of a unit square after a linear transformation. Therefore, the determinant of the identity matrix is naturally 1; a determinant with a row of all 0s is clearly 0 (because one dimension is ignored, and the linear transformation squashes the entire plane); and |A \cdot B| is clearly equal to |A| \cdot |B|. If the determinant is 0, the corresponding matrix is naturally not invertible, because such a linear transformation has already squashed the plane into a line, and nothing can change it back. Of course, higher-order matrices correspond to higher-dimensional spaces. In an instant, everything is explained clearly.

Incredibly, such exciting things were not mentioned at all in the textbooks we used! Why don’t those textbooks that start with the definition of a determinant first treat the area under linear transformation as the definition of the determinant, then derive the calculation method, and then add a supplementary note: “Actually, logically speaking, we should first use this calculation formula to define the determinant, and then say that the determinant can be used to represent area”? Sacrificing readability for the sake of rigor is simply not worth it. Writing this, I really want to pick up a linear algebra textbook immediately, look at all the definitions and theorems with fresh eyes, and then rewrite a true linear algebra textbook.

Calculus textbooks are equally absurd. Mainstream calculus textbooks always teach derivatives first, then indefinite integrals, and then definite integrals, completely reversing the order. Many people, after finishing calculus, can use it skillfully but still don’t understand what is going on. The reason, again, is the problem of math teaching.

My ideal calculus textbook would teach definite integrals first, then derivatives, and then indefinite integrals. Teach definite integrals first, but definitely do not use the current definite integral symbol, to avoid students mistakenly thinking that definite integrals evolved from indefinite integrals. Teach the integral ideas that have existed since ancient times, teach the method of partitioning, summing, and taking limits, and create a custom set of symbols for definite integrals. Then start fresh and teach differentiation, infinitesimals, and rates of change. Finally, mention that as x increases bit by bit, the rate of change of the area under the curve is just the height of those vertical lines—isn’t that just the function value of the curve itself? Therefore, conversely, to find the area under the curve corresponding to a function, one only needs to find a new function such that its derivative is exactly the original function. Pop—calculus is born.

Formal derivations alone are useless. This is the way to truly make someone understand calculus. Strict definitions and strict proofs should be placed after intuitive understanding. Unfortunately, I have yet to see a textbook written this way.

==================================

Having said so much, it all boils down to one sentence: Our process of learning mathematics should be the same as the process of human understanding of mathematics. We should learn mathematics in the order of its historical development. We should start from ancient counting, learn arithmetic and geometry, learn coordinate systems and calculus, understand the motivation for the creation of each branch of mathematics, and the winding developmental history of that branch. We should experience every bottleneck in the development of mathematics, experience the greatness of every new theory, experience the feeling of mathematicians being caught off guard by every mathematical crisis, and experience the difficulty of moving from intuitive thinking to formal descriptions.

Unfortunately, I have not found any path to learn mathematics in this way.

But that’s okay. Since there is no shortcut, let me read through those formal definitions and proofs myself and then experience the reasoning within them. Looking at it this way, our education isn’t wrong either: first use exams to force everyone to learn what should be learned, even if they don’t know what they are learning; then, one day in the future when a certain height is reached, look back at what was learned in the past and suddenly have an epiphany, understanding what was actually being studied. This is undoubtedly a more enjoyable thing. I hope that one day I can, like today, realize what higher algebra is actually about, realize what category theory is actually for, realize why the Riemann Hypothesis is so incredible, realize what a Hilbert space is, and then write them all down.

I suppose that might take me most of my life.

Please include the original address when reposting: https://kexue.fm/archives/1324

For more detailed reposting matters, please refer to: Scientific Space FAQ