I never want to dwell too much on the definitions found in textbooks, because I know that a true understanding of a definition is something that must be gradually realized through long-term practical application. Therefore, the only thing we need to do is to continue our research.
However, a few days ago, some friends asked me about my understanding of differentials, such as "must dx be very small?" and so on. So, I decided to write down my understanding here.
Something closely related to differentials, and something we are very familiar with, is of course the "increment," such as \Delta y, \Delta x, etc. An increment can obviously be arbitrarily large (as long as the independent variable remains within the domain). Now, consider a function y=f(x). How does the differential of the function arise? It arises because studying the increment of a function directly is quite cumbersome. Therefore, we introduce the differential dy. When \Delta x is small, it represents the principal part of the increment: \Delta y = dy + o(\Delta x) = A \Delta x + o(\Delta x), where A is a constant.
We usually say that o(\Delta x) is a higher-order infinitesimal term compared to \Delta x, but this result is not intuitive. o(\Delta x) can usually be written as O(\Delta x^2), which means there exists some positive constant k such that |\Delta y - dy| can be controlled within k \Delta x^2!
One must realize that dy is the principal part of \Delta x. Since \Delta x can be arbitrarily large, dy can certainly be very large. For the independent variable x, it can be viewed as the function y=f(x)=x; therefore, its differential is clearly dx = \Delta x, with A=1. We already know that for other functions, A=f'(x), which means dy = f'(x)dx = f'(x)\Delta x. Obviously, as long as the derivative is not zero, dy can be arbitrarily large because \Delta x can take any value.
The main reason many people believe dx and dy must be very small is due to the following rules: \begin{aligned} f'(x) &= \lim_{\Delta x \to 0} \frac{\Delta y}{\Delta x} \\ f'(x) &= \frac{dy}{dx} \end{aligned} Thus, dx and dy are seen as equivalent to \Delta x \to 0 and \Delta y \to 0. I understood it this way at first and felt there wouldn’t be any issues, but it was only after studying university-level calculus and understanding concepts in quantum mechanics that I confirmed I was wrong. The understanding mentioned above is the correct one. If one asks why f'(x) = \frac{dy}{dx}, it is better to answer straightforwardly: because the definition of dy is dy = f'(x)dx! This avoids putting the cart before the horse.
Some students might raise the following objection: for a composite function y=f(g(x)) where u=g(x), we know from the chain rule: \begin{aligned} dy &= f'(u)u' dx = f'(u)u' \Delta x \\ dy &= f'(u)du = f'(u) \Delta u \end{aligned}
Doesn’t that imply \Delta u = u' \Delta x? Isn’t this an obvious contradiction?
Indeed, this is an easily discovered contradiction. However, this is merely because our notation is too simplistic. One must understand that the differential of a function is related to the independent variable. When x is the independent variable, we have: \Delta y = dy + o(\Delta x) But when u is the independent variable, we have: \Delta y = dy + o(\Delta u)
Since u and x are not equal, how can we assume that the dy in both equations is the same? At this point, I think many readers who have just gained clarity might become confused again. If the dy for different variables is inconsistent, why don’t textbooks use different symbols to distinguish them? This is actually a matter of "implicit convention," much like how we use a prime to denote a derivative f'(x) without specifying which variable we are differentiating with respect to; we have already defaulted to the content inside the parentheses as the variable. As for dy, although it varies depending on the variable, whenever it is used, it is always paired with dx, du, etc. At that point, it is equivalent to specifying the symbol of the independent variable. Furthermore, although the forms of dy differ, the difference between them is only a second-order infinitesimal term; therefore, using the same notation does not cause confusion.
Please include the original address when reposting: https://kexue.fm/archives/1815
For more details on reposting, please refer to: Scientific Space FAQ