Section Insights
Introduction to Mathematics Essentials
What are the six essential pillars of mathematics?
Terrence Tao introduces six fundamental concepts in mathematics: numbers, algebra, geometry, probability, analysis, and dynamics. He emphasizes that these concepts, while basic, have evolved into sophisticated mathematics over time.
- The six essential pillars of mathematics are foundational yet complex.
- Mathematics serves as a precise language to describe intuitive concepts.
- Understanding these pillars is crucial for grasping advanced mathematical ideas.
Understanding Probability in Real Life
How does probability relate to real-world uncertainty?
Tao discusses the difference between sanitized mathematical problems and real-world unpredictability. He explains that while exact predictions are difficult due to uncertainty, understanding the frequency of outcomes is essential, a realization first made by gamblers.
- Real-world problems often involve uncertainty and unpredictability.
- Probability helps in understanding and managing outcomes rather than seeking exact answers.
- The origins of probability theory are linked to gambling strategies.
Advancements in Predictive Mathematics
How has mathematics improved weather prediction?
Tao highlights the achievements of atmospheric scientists in reducing error rates in weather forecasting through the application of dynamical systems. While predictions are not perfect, they are significantly more reliable than mere guessing.
- Mathematics has played a crucial role in improving weather prediction accuracy.
- Dynamical systems theory helps model complex natural phenomena.
- Human systems, like stock markets, remain challenging to predict accurately.
Sphere Packing and Communication Technology
What is the significance of sphere packing in modern technology?
Tao explains how the mathematical problem of sphere packing relates to digital communication. Efficiently packing signals in high-dimensional space prevents interference, making it crucial for technologies like cell phones.
- Sphere packing has practical applications in digital communication.
- Mathematics can optimize signal separation to avoid interference.
- Understanding high-dimensional geometry is essential for modern technology.
The Role of AI in Scientific Discovery
How is AI changing the landscape of scientific research?
Tao discusses the increasing role of AI in automating various aspects of scientific research, from experiments to data analysis. However, he cautions that while AI can accelerate processes, it may also overlook valuable insights gained through slower, manual methods.
- AI enhances efficiency in scientific research but may miss deeper insights.
- Automated tools can perform tasks faster than human scientists.
- There is value in traditional methods of scientific inquiry.
Challenges of AI in Science
What are the potential pitfalls of using AI in scientific research?
Tao warns that AI's speed may lead to overfitting and the creation of overly complex models that do not accurately reflect reality. He emphasizes the need for careful integration of AI into the scientific process to ensure meaningful advancements.
- AI can produce rapid results but may not always lead to scientific progress.
- Overfitting is a risk when relying too heavily on AI-generated models.
- The integration of AI into science requires thoughtful consideration to avoid pitfalls.
Transcript
0:00 My name is Terrence Tao. I'm a professor of mathematics at the University of California, Los Angeles. And I have a forthcoming book, six math essentials. Today on Big Think, I'll be talking about six essential pillars of mathematics, how math interacts with science historically, and how it has anticipated many of the great developments in the sciences, and how new developments in AI will impact math and science going forward. Thank you for watching Big Think. If you'd like to support our work, we encourage you to join our members community. As a member, you'll receive our quarterly print magazine, a beautifully designed collection of the ideas and interviews that matter most.
0:34 Bigthink is for curious people who want to take deep dives into big ideas. To support the media you want to see in the world, go to bigthink.com/membership. And now back to the interview. Chapter one, the six essential elements of mathematics. I decided to organize my book around six really fundamental concepts that have origins from thousands of years ago or centuries ago which are very familiar in the early stages to to most people but mathematicians have developed over time to become extremely sophisticated.
1:09 Numbers is the first concept and then algebra, geometry, probability, analysis and dynamics. These are all very basic concepts but they've evolved into very sophisticated mathematics. But if you strip away all the technical complexities they really are just extremely intuitive concepts and the ma the mathematics that we we we developed it's just a precise language to describe it really really carefully and in a way that that allows you to think really clearly about these concepts. Numbers are one of the oldest mathematical inventions and still the most useful really. We have records of carvings on bones that predate the alphabet or or other writing. And it was invented multiple times by multiple civilizations. We didn't have numbers.
1:56 we would have to always speak poetically when trying to describe any situation and it would always be a little bit imprecise and the next person who's communicating what what you're saying would get it slightly different. It allows for precision. Numbers are placeholders for concepts like quantity and size and magnitude into very portable things that you can communicate to other people who may not have directly interacted with the objects you're describing. And once once you try and describe a very complicated scenario with many many u moving parts, you need numbers and everything else that's born on top of that. Humans are not really wired to think in numbers. If you don't have the ability to think quantitatively to measure both the benefits and the costs of an action and which one is bigger, you can make some life choices that you'll regret later that you spend a lot of of resources for very little gain. The first step in being able to think more quantitatively u for these things is is to understand numbers and then more advanced mathematical topics like probability and and algebra on top of that. Basic things like agriculture or or trade could not have happened without the ability of numbers to measure large quantities of you know grain and actually the ability to tax is you you can't have a civilization without taxation unfortunately and that requires mathematics and numbers. Of course not everything is is is quantitative.
3:20 You know if if you want to go on a date you shouldn't be measuring the cost and benefits of of of your of your prospective partner. Okay. Some things should be still be very subjective and and personal but there are increasingly in this modern world there are lots of decisions for example in finance or in medicine where some quantitative thinking is very helpful. The thing about numbers is that they take on a life of their own because once you have the concept of number, you can study numbers in abstractly divorced from their actual application. And you find patterns and you find often that it's very natural to extend the number system that you have to create new numbers which you wouldn't have thought would be applicable to your your original context, but they fit very very well into the number system. If you're counting sheep, you know, sometimes you want to add sheep and you want to subtract sheep. And so very soon you develop the notions of addition and subtraction. But you realize if you are only counting counting numbers 0 1 or actually not even zero 1 2 3 4. You find that you can always add to these numbers together, but you can't always subtract.
4:21 If you subtract four from three it doesn't make any sense. You can't take away four sheep from three sheep. But patterns in the numbers themselves are so regular that if you just blindly apply the rules of of of arithmetic, it feels like you should be able to take away three from four and have a new number. And it it took a while, but eventually people realize that you can invent these negative numbers and and you can add negative 1,2g3 to your number system and zero. And zero took a long time to to actually realize it was a good addition to to the number system. and you still get all the nice laws of of of arithmetic. For example, if you take a number A and you subtract B and then you add back B again, you you get back to A.
5:04 That's one of the laws of arithmetic and it still works even when you have negative numbers. Similarly, we learned to divide numbers by another and and we created fractions and fractions fit very well into this number system and then there was a shock. We found that there were numbers somehow between all these rational numbers, all these fractions, there were numbers like the square root of two which could not be expressed as any ratio. and this was a big shock actually that they I mean these numbers are literally called irrational numbers which is Latin for you know insane not unreasonable numbers but they do exist and they're very useful. You can never write down all their digits or whatever in a in a on a on a finite piece sheet of paper. it is very very useful to have all these extra numbers lying around and then eventually we we tried to take square roots of regular numbers which we couldn't do in a regular number system and we invented complex numbers and that turned out to be extremely useful for electromagnetics and and quantum mechanics. It's remarkable that these number systems which were often invented just so that we could be better at solving equations and and solving practical problems end up actually being the most natural language to describe very very complicated phenomena in the real world like quantum quantum mechanics for instance.
6:21 Algebra is the second layer of abstraction over numbers. So with numbers, we took concrete things like a bunch of sheep or a quantity of water or whatever and we replaced these these quantities with numbers that you can then apply operations to addition, subtraction, division and so forth. Algebra is goes one step further and tries to not look at specific numbers like seven or 17. First of all, replace numbers by even more generic placeholders. We give them names like x and y, but to also study the operations themselves, plus and times and and all these other operations and ask what properties do the operations have, not just not just the numbers. And so people discovered that the these operations that are so useful in arithmetic themselves have many fascinating properties. so addition has a property called commutativity. If you add A to B, that's the same as adding B to A. And that turns out to be extremely useful property for helping you solve problems involving addition.
7:20 Similarly for multiplication and all the other basic arithmetic operations obey these very simple laws. Later on we discovered that these laws also hold for other operations. For example, if I want to take say an object and I rotate it, I can rotate it by 30° and then rotate by another 60°. But if if I rotate it in the in the other order, rotate by 60° first and 30° next, I end up with the same position that I started with, these two rotations are commutive.
7:47 So even though that operation has nothing to do with it's not really an addition or multiplication in a traditional sense of numbers, it it has the same algebraic structure. On the other hand, some things are do not obey the commutive law. Like if I put on my socks and then I put on my shoes, I get a different outcome than if I put on my shoes first and then I put on my socks. Those two operations do not commute. Certain operations obey nice laws like like the commutive law and certain operations don't. Sometimes you you see a match. The laws that that that you you you you are present are very similar to laws that we already understand for say numbers. and then because of that we can take intuition and and ideas and and and proofs and and a theory of numbers and we can transfer them to a a a different setting. For example, matrices are a much more complicated concept than a numbers. Not just one number, but it's a whole square array of numbers. But it turns out that matrices obey very similar laws of algebra to numbers. And if you are very good at manipulating numbers, you can start to manipulate matrices the same way. And many of our modern technologies for example large language models are based on being able to manipulate matrices very very efficiently. Once you have abstracted to get to numbers and into algebra then you are working with equations that involve variables like x and y and x may have had some physical meaning and some specific value. But often it can clarify your thinking to not focus on the specific values of these numbers or or what they represent and just manipulate these these equations by pure algebra just by moving symbols around. When you first learn this it it it feels very disconnected from your actual experience. But it is a very powerful technique. One early use of algebra was the story of Johannes Kepler who was this astronomer and scientist who was one day walking down the streets of his hometown and he saw the wine market and and wine sellers was was selling wine by the barrel and so some large barrels and some small barrels but they were able to figure out how much wine was in each barrel and and so that they could pay the wine the wine sellers for for for these barrels. if you have a barrel and you want to compute how much how much wine there is there. You know, you could pour it out and into cups and so forth, but it was very tedious. But what fascinated Kepler was that the person in charge of the market had a very very efficient way to to measure the volume of the barrel.
10:16 He just had a a stick with various markings on it and there was a bung hole in the middle of the barrel and he just poked the stick down the barrel into the corner and just saw how far the stick went and he could measure and just based on what that marking was, he could say, "Oh, this is, you know, 30 gallons or whatever of wine and then they could price it." This astounded Kepler how you could just take this one measurement of just this one sort of diagonal and work out the shape of the volume because some bowels could be very tall and skinny or short and wide and somehow this this one measurement was somehow able to compute the volume. He did the math. This was a puzzle to him. So he he went home and he wrote out some equations. Assume my barrel is maybe the radius of the barrel is r and the height is h. You know nowadays with modern algebra like this is a question you can assign to a high school student. It's almost precisely one of these word problems that that that we we love to give our students and you can compute the volume and you compute the this length and he found that the length did not completely determine the the the volume but a wine seller would want us to to sell as much wine as possible and you know you would want to sort of maximize how much volume you could get for a given for a given length. Reasoning that you know all these merchants were trying to maximize their profit. even though he couldn't quite solve the equations right away, if he added this profit incentive, he did some rudimentary what we would now call calculus. And that turned out to almost exactly match the shape of the barrels that actually were sold in in the marketplace. And his formulas did actually match what what people used very high precision. So he actually explained these these rules which I think the the marketplace had come up with over time just by empirical measurement. But he had found a very satisfactory explanation and I think this was part of the inspiration for the the theory of calculus developed by Newton and Lipets a few centuries later.
12:04 Geometry is literally is Greek for measurement of the earth. Since antiquity it was important to to know you know how many miles it was to travel from one place to another and how to navigate you know in the ocean or or or out in in in the wilderness by the stars. We needed to understand how to use things that we could observe like angles and distances to for very practical problems like transportation. Just like numbers have got various patterns like a plus b= b plus a. Once you start measuring distances between different points and angles, there are lots and lots of relations between those all those measurements as well. So geometry obeys laws just like numbers obey laws. So for example, one law is similarity. Once you know that two shapes are similar, you know they have the same angles and things, then all their sides are proportionate. Once you know that the side of the big triangle is say two times as big as the side of the small triangle, then you know that all the other sides of the big triangle are also five times as big. The proportions are equal. The reason why this is so useful is that it allows you to predict the measurement of of distances or scales that you couldn't directly reach, you couldn't directly measure. You can look at a distant mountain and you might be able to actually estimate how how far away it is because you know something about how tall it is and the angle of elevation is you can use a similar triangle and you make a scale model. So for example, now you can navigate and you can see how to get to a distant location even if you can't see it directly from your sailboat or whatever. And even in ancient Greek times, they were able to measure distances to the moon and the sun, which they could not possibly measure directly. But just through the laws of geometry, which they had similar triangles and things like that, which they had already worked out at that time, they could actually get reasonably good measurements, you know, without any satellites or advanced technology. Once you know geometry, you can really extend your senses well beyond what you can just touch and measure directly.
13:55 The fourth math essential is probability which is the standard way that mathematicians try to encapsulate one of the basic features of the real world which is uncertainty. In your primary or high school classes when you teach mathematics we often present very sanitized predictable word problems. You know you know Annie has has 30 apples. She gives half of them to to James. You know how much does does James have and whatever where everything is precise and you know all the information. But in the real world, there's there is uncertainty and there's unpredictability. So, you know, when you flip a coin, maybe it is heads, maybe it is tails. Now, if you know exactly how much force you apply to a coin and and you you you knew all these measurements, you could maybe do enough of a simulation to actually predict exactly which way the coin will land. But this is extremely difficult. and often you don't have that data. What was eventually realized is that rather than try to compute exact answers all the time for every single outcome, sometimes you just have to accept that there's a range of outcomes to a to a given measurement. But what's important is which ones are more frequent and which ones are less frequent.
15:07 And the first people to realize this was important were gamblers because you know gamblers would gamble on certain events like trying to to to bet that a certain die rolls would sum up to to a certain number or whatever and if they could calculate the odds correctly they could make money and if they didn't calculate the odds correctly they would lose money on in the long run. The mathematics of probability was was was created by several letters that some gamblers wrote to their mathematician friends asking for help trying to to to optimize their their their gambling strategies probably has has has blossomed well beyond it it its gambling roots. you know anytime we have a system too complicated to model all the way from first principles there's going to be some stoasticity and we we want to have a proistic model. So whether it's a stock market or whether a drug is going to be be successful, we turn to probability now quite often.
16:01 Now sometimes we don't know what the odds are. So probability works best when it involves an event that happens over and over again. You know, there are thousands and thousands of trials and we can start getting a good measure of the odds. If it's an event that only happens once in a century, it may not quite be the right the right mathematics and we're still trying to work out mathematics for extremely rare events. that is a better replacement for probability and there are some miracles that make it effective. so you may think that every time you do a different experiment, you know, so instead of doing a medical trial, you're you're trying to understand the just the outcome of of a of a die roll or what genetic traits are going to emerge from evolution or whatever. and and you would think that in every different circumstance you get all these different distributions, you know, so some distributions will be heavy tailed and some will be narrow and whatever. But there are these these funny laws in probability called universality laws that that even very general types of of of of random systems various common shapes emerge. the most famous of which is the Gaussian or what's called the beller shape that many distributions like if you take the distribution of men or heights of women, they form almost a perfect bell curve. We actually have good explanations now. We we have understood pro probably well enough that that we can explain quite a few of these universal laws from mathematics although some are still mysterious.
17:25 Analysis is how mathematics deals with two things. One is inaccuracy in our measurements that sometimes it's not so much randomness but just imposition. Analysis is sort of the mathematics of error bars where we've realized that sometimes we we need to to understand not only what the numerical value of various quantities are but how much uncertainty we have what what what are the plus and minus error bars around them but these are concepts that you you you you can't even talk about I mean if it just having only qualitative language of big and small is is is not it only works up to a point. We may measure say the length of an object and it it's roughly 2 meters but may 2 meters plus or minus 10 centimeters. There's an approximation and there's an error. And ideally you want the errors to be zero but in in the real world we can't always make the errors entirely zero. But sometimes if we we can make the errors smaller and smaller and if we keep making more and more precise measurements we can make the errors shrink to zero. But it may take an infinite amount of of of of precision or infinite amount of time to get the error all the way down to zero.
18:35 Analysis is also about how we we take limits and how we deal with infinities. In algebra, the laws of algebra work very well when you're just taking a a finite number of operations. If you're just taking five things and you're adding them together, there's there's no problem. You can you can rearrange them in any order. But once you start trying to to work with an infinite number of operations, there are some funny paradoxes that that that show up that that sometimes you can rearrange an infinite number of objects and they they end up you end up with a different sum than than before you rearranged which doesn't happen with finite sums. For example, if you're always betting on say roulette or red and black, you win 50% of the time and you lose 50% of the time. There is a theorem that there is no strategy that will allow you to constantly win that will guarantee you a win in the long run as long as you only have a finite amount of money. No betting strategy can beat the house basically.
19:33 And so there is a strategy of always doubling down when you lose and just betting bigger and bigger numbers. And the moment as long as you win at least once you will get your dollar. So there was some way to constantly beat the house. But the problem is that it assumes that you have an infinite amount of money. What the strategy is doing is that it is compressing all the risk of losing money into this very very small event where you're always losing. At some point you know you're you're betting millions of dollars and at some point you become bankrupt. So analysis helps you understand exactly what these terror risks are and how to reason with infinities in a way which is in which you avoid all these paradoxes. Yeah, there there was a lot of very inaccurate mathematics that was that predated analysis where people were just saying, "Oh, if I do this infinitely often, I can get rid of of all of all these problems." And it took a while to realize that there was a that infinity is a very dangerous beast if you if you're not trained to deal with it properly. But we we need to deal with infinities all the time in the real world. Well, maybe not infinities, but we need to deal with very large numbers.
20:41 But infinity is a very good approximation for understanding how we deal with large numbers. But it has to be used with care. So one famous demonstration of of the unintuitive nature of infinity is what's called the infinite monkeys theorem. The common way to to phrase this is that if you have an infinite number of monkeys in a room or maybe just one monkey typing infinitely for forever on a typewriter and just hitting keys at random then usually there the monkey will just type nonsense in gibberish but every so often the monkey will type a a a word know the or it and sometimes it will it will type a sentence but the infinite monkey theorem is that if you wait long enough or you have enough monkeys you're almost certainly guaranteed eventually that the monkey will type whatever you like like the complete works of Shakespeare or Hamlet or Wikipedia or or anything. And you can you can prove this mathematically that that as as long as the probability of the monkey doing it at least once is positive. It doesn't matter how how small as long as it's positive if you wait long enough eventually the probability that this this particular pattern gets hit will eventually go to one. It's like if you play Russian roulette and you you you you only have one bullet in the revolver and you keep you keep firing. it doesn't matter how many chambers you have, eventually it will eventually fire and you you you'll get your hit for any given word or sentence or or paragraph or whatever. Your monkey will eventually create this text, but the time taken grows exponentially with with the size of the text. So if if it's just a four-letter word, it might just take an hour or so of typing before the monkey gets it. But if it's already say a sevenletter word that you know the word Shakespeare for instance then that may already take years a sentence you know it might take millennia and actually to get even a fraction of Hamlet think even just a page would would take way more than the age of the universe before you'd actually see it.
22:39 So infinity is actually just a placeholder for a number anything which is which could potentially be be far larger than any fixed number that you could come up with. You know, even though in the real world you don't have infinitely many monkeys, you don't have infinite budget. Often we we reason with infinity first in as an idealized situation to see what is possible and then from there we we can we can turn to the more quantity questions of of exactly what can we do with finite resources. But first you you understand what to do with infinite resources. When I was a kid I used to play a lot of computer games. For many computer games there were certain cheats you could you could put in your games. you could give yourself infinite health or infinite ammunition. And it was sometimes helpful to play that game first with those cheats where you didn't have to worry about managing your your health potions or or your your ammo or whatever and just see how to solve how to solve the game and then you could play it on hard on a harder mode and then see how to to do things more more efficiently.
23:35 So many problems in math are actually solved this way. One thing that distinguishes math from other disciplines is that we have the freedom to fail because failure is very cheap in in mathematics. You know, if you're running a business and you make a bad business decision and your company goes bankrupt, that's a terrible mistake. You know, if you're a surgeon and you cut the wrong thing, that's a terrible mistake. But if you're trying to solve a math problem and you make an incorrect assumption or something and it doesn't work, it's not really that much of a bad mistake. You just try again. It's it's actually a good move when you're trying to solve a math problem is to just first make an idealized assumption. assume that you have, you know, an infinite amount of of of energy or or zero friction or or some other unrealistic assumption, solve the problem then and then try to get from the infinite world back to the finite world. And this is where analysis comes in to sort of carefully see what features of of of infantry mathematics still work in the finite world and which ones break down.
24:32 Dynamics is mathematics of change of time is the study of how the rules of incremental change. How a state evolves from one time to the next all kinds of emergent and interesting behavior which you may not expect from the from the initial rules. So we have found that even very simple rules can generate extremely complicated emergent behavior if you iterate them long enough. Evolution is a good example in biology. You know, you have a bunch of organisms and they reproduce and the fitter ones survive more often than than the weaker ones and the organisms have certain traits that they can pass down to descendants. And these are very simple rules and it turns out that you can create a a massive diversity of species and and predator prey relationships and and and incre incredibly complicated dynamics. If you're on the freeway and and you have all these cars and and each car is just trying to move as fast as it can given the car in front of us, you know, so if there's too many cars in front, you'll slow down. If there's not many cars, you speed up. And each individual car is not doing anything very complicated. It's just trying to optimize its flow. But when you put all the cars together and you see what it does to the whole network, you get these amazing emergent phenomena like traffic waves. You get these waves that can compress and grow. Kind of like a slinky. If you mess around with a slinky, you can have a compression wave and an expansion wave and it leads to phenomena which you wouldn't initially expect like like if there is a traffic shock like there's an accident and then the cars pile up then even when the the traffic accident is removed and there's no obstruction to travel the laws of traffic if you if you iterate the dynamics you can actually do the math and you can see that takes a while for the compression wave to dissipate and so living in Los Angeles this is a a phenomena that I've encountered quite often that sometimes I was I encountered a slowdown in traffic and there was no accident or no immediate cause. It was because many hours ago there was something that caused a slowdown but the dynamics it takes a certain amount of time for the the the wave to dissipate.
26:32 You know once you understand the dynamics really well you can do modeling and simulation and then you can do things like you can you can make predictions like if I add another lane to this this this freeway will the traffic get better in fact sometimes it doesn't these paradoxes where actually sometimes closing off certain lanes of traffic and actually make the traffic the global traffic flow flow faster now some dynamics are predictable sometimes we have equilibria which are states that just stay the same all for all time and sometimes these equilibria are stable Well, if you move a little bit away from from that state, you you come back to that state. Like if you have a pendulum that's going straight down, that's a stable equilibrium. And if you modify it a little bit, you perturb it, it will sort of move a little bit and it'll get back towards a stable equilibrium. But if you make a pendulum upside down, it's balancing on the tip. It could be an equilibrium. Like technically, it can stay in that position forever, but any slight pertabbation will actually cause it to move away from the equilibrium over time. So it's important to know what equilibria are stable and which ones are not. We are now facing a world of climate change where we have lived for a th00and 10,000 years in the climate in a pretty close to an equilibrium state. Make it hotter or colder some years, but it would bounce back to equilibrium. And we're now actually in danger of leaving that equilibrium and to a much less stable dynamics, which is scary, but it needs to be modeled. and we may need to figure out how to adapt and and change our agriculture and all our other practices. Understanding dynamics and which systems are stable, which ones are not, which ones are chaotic, which ones are predictable. It's actually extremely important. There are very mundane things like predicting the weather. You know, we take for granted that that we have accurate weather predictions for the next seven days. This was this is actually an amazing achievement of atmospheric scientists. They they collected lots and lots of data, but they also solved a lot of dynamical systems problems that allowed over the years got the error rate down to a point where we we can actually reliably predict weather. forecasted to say a week in advance. It's still not completely 100% accurate, but it's much more accurate than just guessing.
28:37 Systems which involve a lot of a lot of humans are still very unpredictable. So the dynamics of the stock market or or politics. This this this is well beyond ability of current dynamical systems theory to to model. But but natural systems and some human systems like traffic we can actually model. So it it is a fairly advanced area of mathematics. We often need a lot of computer simulations and we need to solve very advanced differential equations but it can give some very valuable insights. One of the discoveries of dynamical systems is is that most systems exhibit what's called chaos. And this was this came as a surprise in the 17th century. So Newton when he introduces law of gravitation.
29:19 one of the great successes of this theory was that it explained the motion of the moon around the earth and the earth around the sun. He could explain retroactively all these funny laws of Kepler like why planets move in ellipses and things like that. And so he solved what we now call the two body problem that if you have two massive objects like the sun and the earth and you move them around and you govern by a single law of motion Newton's law of inverse square law of universal gravitation he could solve the equations using his newly derived theory of calculus and he could have perfect formulas for for the orbits and they were perfect ellipses ex exactly verifying Kepler's theory it was an amazing achievement once he solved the two-body problem it was it was very natural and many of Newton's successes And I think also Newton himself tried to solve the three-body problem. I think Newton once said that this was the only problem that ever gave him a headache because no matter what he tried, he could not get an exact solution.
30:15 leading societies of the time offered major prizes for anyone who could who could write the solution. This was considered one of the major open problems in mathematics. We still do not have an exact solution for these equations. And the belief now is that there isn't really one that you can write down as a nice neat formula. But when you actually look at the numeric you see that it is not some nice periodic pattern. It often stays periodic for a long period of time but then suddenly it will change to something a little bit different and then it will change yet again. We suspect now that our solar system which currently has what eight planets or something that in the past there were other planets that were in the system and they mostly moved in sort of elliptical orbits like as as according to Kepler. Every so often the the little interactions between the gravitational force of Jupiter, exerted or Mars and so forth would jiggle these these planets a little bit out of their usual orbit and occasionally they would just veer off completely and sometimes two two planets would collide or one would escape the solar system and you know for example there's an asteroid belt which we believe is the remnant of a collision from millions of years ago.
31:20 Even the most stable of systems like the solar system that looks like it it hasn't changed for millennia, there are long-term instabilities in in that system. Once you move beyond the simplest of systems there there's lots of little tiny unpredictable or or very hard to predict deviations that that that occasionally can can pile up just like occasionally a bunch of monkeys can sort of write the works of Shakespeare. Occasionally gravitational perturbations can can set an entire planet off off course.
31:52 So often actually in the most advanced forms of dynamics today, even if you start off with with a completely deterministic system with no unpredictability whatsoever, we find that the best way to model it actually is to to approximate it by a using probability and just assume that there's going to be some random fluctuations back and forth and eventually the your predictions will just get blurriier and blurriier which and that just seems to be a fundamental feature of chaos which many systems have. So these six essentials they don't describe all the mathematics but they do describe six of the great themes that mathematics tries to encapsulate and this is much more precise. This is just a taste of of what what goes on these days.
32:37 Chapter 2 how math solves the problems of science. I view STEM as a whole ecosystem. At the bottom there's basic research like mathematics and some other fundamental sciences where we pursue things mostly driven by curiosity. We see a phenomenon that is crying out for an explanation or further study. It may not be a phenomenon that we urgently need to solve right now for an immediate problem but it's something that looks like it should have an interesting answer. And so mathematics is is is almost entirely curiositydriven like that. There's some pattern in numbers. There's some pattern in shapes.
33:14 People just observed while trying to do something else. And we want to understand it better. At some point, other scientists are able to connect that pattern to something that they're studying and and some mathematical numerical pattern might show up in the behavior of of of insects in a swarm or in a stock market or whatever. and then sometimes once you understand it you you can actually convert it into some useful technology or you can have a company that actually makes some money out of out of somehow some service related to to that exploiting that phenomenon. We often don't we don't see that. I mean that that's that that's much further down the pipeline. What you do need is that you do need the people doing the basic sciences to talk to the people doing applied sciences and they have to talk to people who are doing engineering and they have to talk to people in industry. If you didn't have one of these communities, then you wouldn't have this pipeline of getting from curiositydriven questions to actual, you know, commercial results, you know, you know, like the ability to communicate across the across the planet with almost zero cost. This is part of what Eugene RNA calls the unreasonable effectiveness of mathematics in the physical sciences.
34:22 He observed that mathematicians often discover concepts such as complex numbers or curved space or whatever just because it it seems to be a natural extension of the mathematical objects they're already studying. and then 10 20 50 years later some scientists discover that that these concepts that were introduced for fun or for play were in fact exactly what was almost exactly what was needed to understand some new new type of science. So that's a really amazing phenomenon and we still don't have a good explanation really for for why that actually works. So one historical example of how curiositydriven mathematics led to a really deep scientific advance was the story of the parallel postulate.
35:04 Uklid in like the 3rd century BC introduced the notion of proof of being able to explain complicated results in geometry in this case from simpler axioms. for example that the sum of angles of a triangle always added up to 180 degrees. He was able to explain that in terms of simpler axioms. He laid out five axioms of geometry that he thought he reduced all the other facts he knew about points and angles and lines to these five statements and four of them were very straightforward like if I give you two points there's always a line you can draw between them right things like that these are very straightforward non-controversial axioms but there was there's one axiom which is called the parallel postulate which gave him a lot of grief and in fact his original version was very very complicated it got simplified but even even the simplified version was controversial The simplified version is that if you have a line and you have a point and the point is not on the line then there's exactly one line you can draw through that point which is parallel to the first line by parallel I mean that it never crosses this first line so that was this axiom that you can always draw a parallel line through any other point and there's only one you cannot draw two parallel lines once you have that you can do all kinds of things you can derive what I just said the angles of a triangle at 180° and all the other classic results of uklitian geometry but it was a very ugly axiom compared to the other four which were which were really elegant.
36:27 Eventually they realized that that what they had done was that there actually were multiple geometries beyond uklidian geometry. There's something called spherical geometry where there's no parallel lines at all. so in spherical geometry instead of lines you have great circles like the equator or a longitude. And these great circles on the sphere they always intersect. You can never make two parallel great circles. So there are no parallel lines. And then there's this weirder geometry which is harder to visualize called hyperbolic geometry where lines actually diverge from each other and have lines that start off looking parallel but they they move further and further apart and and in fact now there are actually multiple parallel lines you can draw from one point to a given line and those two geometries are entirely self-consistent and eventually it was just accepted that there was there were more geometries out there than just ucleian geometry. So these were the first two non-ucleian geometries to be discovered spherical geometry and hyperbolic geometry. But once we had sort of freed of our notion that there was only one geometry, this opened all the floodgates and people studied all kinds of other geometries.
37:27 So all kinds of curved spaces, you know, spaces that were shaped like donuts or had had twists in them. There are geometries where you know, if you're right-handed, you can go off explore the universe and come back. And if you start off right-handed, you will come back left-handed. that there there are geometries where where you can change your orientation just by by by travel. So which is very unintuitive or you can come back smaller or larger than than than than what you started with. People developed all these geometries and they developed a very nice language for for describing all all these geometries. it's called Romanian geometry after Bernard Reman and but it was a curiosity. I mean though these abstract curved spaces but you know the the universe we lived in seemed completely flat. But then Einstein when he was trying to understand gravity he eventually came to the conclusion that what gravity was doing was it was bending space and time in a certain way and he needed a language to to describe how space and time could bend in such a way that light rays would would become not straight and sometimes you know they would hit each other or or diverge. He asked his mathematician friend if there was any existing mathematics that would describe this and here he said oh yeah there's this bright chap Bernard Reman who developed this theory but it turned out to be almost exactly the right language to describe the Einstein equations there's a notion of curvature that some space can have positive curvature negative curvature in Romanian geometry and the Einstein equations turned up extremely simple in to state in this language that basically mass and energy create curvature that the curvature of space and time is proportional to how how much mass and energy you have in your system. And that's basically the Einstein equations.
39:05 Now, solving them is a different matter. They're extremely hard to model. Even the question of for how to model two colliding black holes, we can barely do it with modern supercomputers. But but stating the the the equations is actually extremely natural once we had this language. Another example of of how mathematical curiosity led to like really practical developments centuries later is the story of sphere packing. There was some British sailor who was just curious about the question of you know there's a certain number of cannonballs they had to stack in the hold of of their ship and these are these are round cannibals they're not they're not square so when you when you stack them there's a certain amount of wasted space and he was curious what is the most efficient way to pack cannonballs so you can get the most cannonballs into a certain amount of space and so he asked a physician friend who happened to be Johannes Kepler eventually proposed that the most efficient packing should be the same packing that that you see in nowadays in supermarkets when you pack oranges. It's what's called a hexagonal closed packing. You pack layer by layer. Each layer is sort of the triangular grid of cannonballs or oranges and then you stack another triangular grid on top of it just shifted by a little bit and then you stack it back and there's a a regular pattern which is the most natural pattern and it's about 76% efficient. And Kepler thought this was the the best you could do that. there was no clever way to to to squeeze in any any any more space but he couldn't actually prove it. So this became known as the Kepler conjecture. and it was one of the most famous unsolved problems in geometry for for centuries in two dimensions. I think it was solved by about 1900 or something like you're packing discs in a plane that's a simpler problem. That one there's a similar lice of triangular lattice and and that was relatively easy to prove that this was the optimal one. three dimensions are just too many possibilities. There was no way there.
40:57 in fact, we still do not have a nice simple proof of the conjecture that that that humans can completely understand by themselves. The conjecture was eventually solved. It eventually got published I think in 1998, but it required computers. It was one of the first computer assisted proofs. There was a team of referees who I think said that they could not verify all of the computations but they at least believe that the the strategy is correct but there were still lingering doubts.
41:25 It is only much more recently 2014 I think and finally the proof was converted to what is called a a proof assistant language computer language that is specifically designed to check proofs with 100% certainty. So the Kepler conjecture is now formally verified. we are now 100% certain it is true but mathematicians were not content with just that the threedimensional problem so they also asked what happens if you're in four dimensions or five dimensions or six dimensions so here of course there is no practical you know I mean there are no fourdimensional oranges or or cannonballs that that you would like to like to pack but people still ask this question people also ask what happened if instead of a continuous space you have a discrete space in particular computer scientists once computer computer science became developed We realized that in addition to the geometry of regular space where XYZ where coordinates are given by real numbers, we're interested in studying the geometry of of strings of bits. So this is now very divorced sounding from the original sphere packing problem both because now you have many many dimensions like like thousands and thousands of dimensions and space is now discrete rather than continuous but still it is geometry and and many of the techniques to understand sphere packing still work. when you when you're in a setting. And then it turned out that this setting of this problem of of packing spheres as efficiently as possible into on this big huge cube of of bitst strings turns out to be extremely practical. When cell phones became digital, they every signal that you that you send, you know, like an image or a text, whatever is encoded as some bit string, which is then sent over over some wireless network, but there are other people also sending the signals and you don't want your signal to be corrupted by interference and be mistaken for for someone else's signal. So, you want to keep each different signal that that comes out.
43:17 You want to keep them as as separated from each other in this space of of bitstrings as possible. And it turns out mathematically this problem of separating all these signals so that there's no way that one can be confused for another is almost exactly the sphere packing problem except that it's in high dimensions and it's discreet. All the mathematics, well, not all of it, but a lot of the mathematics that was used to understand spear packings could be used to design really efficient sphere packing codes and not just to design codes, but also to to to tell engineers what was the what is the theoretical limit of communication like what what is the maximum number of of bits per second you could possibly hope to to send in a certain wireless spectrum. So that gave really good benchmarks to measure how efficient your your protocol was. They you could price how many billions of dollars you should pay for a certain spectrum wireless spectrum because now you know exactly how much data you can push through that. And so like the entire wireless telecommunication industry is is based on being able to to to pack oranges in really high dimensions. One of the applications of mathematics that I was involved in that I'm most proud of is the story of compressed sensing.
44:27 So I was once at an inter interdisiplinary program at a mass institute here in Los Angeles and I met with a friend of mine who is a statistician and he was working with an electrical engineer and trying to improve imaging medical imaging and specifically MRI scans at the time MRI scans were were quite slow you had to sit in this scanning machine for like 3 minutes so that you could collect enough data from all different angles that the the scan could reconstruct a good image of your body and be able to pick up tumors or cysts or anything else which is sort of medically important to to to resolve. But if you only sat in the machine for a short period of time like say half a minute you will not get enough data. so if you tried the standard reconstruction algorithm from all the data to recover the image using what's called le squares approximation which is the standard technique at the time you would get an image that was so blurry and so res so so low resolution that you could not tell anything useful for diagnosis purposes. So we had to sit in these machines for for minutes and minutes and if you were a kid sometimes they you had to be sedated because the kid would wrigle around and not follow instructions after like the second minute. So they were trying a new technique not least squares something called total variation minimization.
45:45 They had a hunch that that this other method might perform a little bit better but so they they tried it on some test data. They were expecting a slightly sharper image than the least squares approximation but they got perfect resolution that they got back almost exactly the correct image even though they only took a few measurements. It would be like giving someone a cross word puzzle where you had only filled in, you know, 10% of of of of the letters and suddenly they could fill in all the other letters without having to look up the clues. They couldn't explain this and and they showed this to me. And my first instinct was you made a mistake. You could not possibly have done what you said you did. And in fact, I'm going to prove to you that that there was not enough information in the data that you you took to make to make this measurement possible. So I I I went home that night. I tried to write down a proof that that there was no way to just correctly guess the the right image from the small amount of data that they were measuring. And while writing it down, I found one of my steps does not work. And in fact, in fact, it showed the opposite that if a certain measurement matrix had a certain property, then actually what they were doing was actually going to work. And then I checked that what that their measurement metrics actually did seem to obey this property. So I actually understood how how the method worked. And so I went back to them the next day and explained this and they got very excited and we wrote a couple papers and that got everyone else excited. This method that they had stumbled upon was not completely new.
47:16 There were seismologists who had discovered a similar method. So they had a se a different problem where they were trying to understand the fault lines to locate the fault lines of a crust based on on a small amount of seismic data and astronomers had a similar problem that they were trying to to measure the location of stars or something using a very small amount of of observed data and there are a couple other disciplines where a similar problem of trying to extract a high quality image from a very small amount of signal had had occurred.
47:47 each case they had found some ad hoc fix that could kind of squeeze more data out of a more better resolution out of the data they had but they could not explain mathematically why it worked. The seismologist thought, "Oh, here's a trick, but it only works for seismographs." And the astronomers had a trick, but it only worked for astronomy. But once we found the mathematical explanation, we found that this was a general technique that we we now call compressed sensing. And it's useful for MRI, but it's also useful for wireless broadband. It's useful for for certain types of sensor networks. Once we figured out the underlying mathematics, we could see all the other applications that it's useful for. And so now compress sensing is taught in textbooks right next to least squares. So sometimes with least squares it's the right thing to do and sometimes compressed sensing is the right thing to do sometimes neither. It's now a a very very welldeveloped theory. I was quite pleased to be involved at the very beginning of that. It's a fascinating interplay between mathematics and science. So we have this unreasonable effectiveness of of mathematics where mathematical discoveries are often end up being the the best way to explain physical phenomena. Philosophers and his historians have have debated why this is the case. one of my theories is that whenever we learn anything, whether it's math or science or any other subject, the first theories or explanations that we we make are often not the best. We don't understand what is the cleanest way to to to express something. And before you understand something completely, you might have an overly elaborate explanation for why something is true.
49:20 But often the the the true explanation is is is more elegant and shorter and simpler than our initial attempts to describe the the phenomenon. But finding the the short explanation takes time because you have to sort of unlearn certain assumptions that you might have that turn out to be to be incorrect. For example, with Einstein's theory of relativity, one of the key assumptions that people had before Einstein was was that was that time was universal. everyone had the same notion of time. You know that that an owl for me is the same as an owl for you. That mindset really blocks you from from from finding the right way to explain gravity in particular. But once you accept that everyone has their own relative notion of time then you can find the right language to explain things properly. Mathematicians also try to take phenomena that they first understand using very inefficient language and try to condense it. try to find the the the most concise elegant explanation of a mathematical phenomenon. And I think just because there's only so many ways you can you can say things concisely. It just happens that that that often the concise way to describe some mathematical phenomenon is also a concise way to describe a physical phenomenon. So that is that is my theory. Unfortunately, we only have one timeline of science. You only have one history of scientific development and we have maybe a 100 turning points in science and that's a little bit of data and you can make some theories. I would love in the future in the far future maybe we will meet other civilizations and we will see their history of science and we will see whether they also had Kep their version of Kepler and Einstein and Newton and whether it what they they followed a similar track or a completely different track. I don't know the funny thing about doing mathematics is that so you know there's this stereotype that you know we're all geniuses and you know we are stuck at a problem and then you get this eureka insight a light bulb goes on and and like you get this genius idea out of nowhere. I would love for that to happen to me actually. This this does not happen so often to me. What what does happen when I do work on a problem is that I try something and it doesn't work. Okay, I try something else and it kind of works but it gets stuck at a certain point. but now at least I know kind of there's at least one obstacle and I need to find some tool that is will will help me deal that obstacle and so maybe I will now identify a sub problem which has the same type of difficulty but is simpler and I try to solve that one first.
51:52 and if I can do that one then I can try to scale up back to the original problem and I go back and forth often a lot of what you're doing is that you are you are exploring the negative space of the problem like all the the techniques that don't work and eventually if you have enough of the negative space the the path forward becomes clear almost by elimination that there's only sort of one thing to do that could work some sometimes there's nothing that can work and then you give up on the problem but there's this repeated what you might call failure but it really is kind of just just really understanding the limitations of what different approaches can do. And then after weeks or months, you know, of of this, the answer becomes clear. But by that point, it's so internalized what the difficulties are that it doesn't feel amazing to you anymore. It feels natural like of course you had to do this because there's this difficulty. You must go around this this pothole. You must do this step first because you know in five lines you're going to need this hypothesis to to solve this problem. you you just become so attuned to the problem that everything becomes natural. yeah or sometimes you never solve the problem because you never attune and and you and you give up. The the feeling I get is never so much eureka but it's always oh how come I missed this early? I was so stupid. but often it it's it's a constant experimentation and and and failure that really primes you to find to accept the right solution.
53:17 There's this dramatic contrast between the standards we assign to outcomes and the standards we assign to process. So for outcomes, mathematics very famously has a very high standard of correctness. You know, for a given problem, there's a correct answer and lots of incorrect answers and and when we grade our students homework in mathematics, you know, if you didn't if you get your sign wrong or whatever and you got the wrong answer, you got all these all these negative marks. You get criticized if you make math mistakes in your final answer. One consequence of that is that many students who go through you know say a high school level of mathematics become very averse to making any mistakes whatsoever in when they try to approach a math problem.
53:59 Paradoxically the process of arriving at the answer is is almost the complete opposite where you almost have to make mistakes over and over again and you have to try the stupid things first to appreciate why the clever things work. There's a quote by Neil B who's a physicist who said that an expert is someone who has made all the mistakes that can be made in a very narrow field. You don't publish these mistakes. You execute these mistakes in your process in order to locate the correct answer but only after exploring a lot of incorrect answers first. It's it's very important I think to to normalize failure in the process to disclose that behind every successful solution to a problem there are dozens of of of incorrect attempts. and it's not because the people trying these problems were were stupid, but this is often just part of of the learning process.
54:54 >> Chapter 3. How AI is changing math and science forever. >> Science and mathematics has changed a lot over the centuries. Traditionally in science the two major paradigms were theory and experiment. like you would you would create a theory like you know Kepler might create a theory of how planets move or or Newton might create a theory of gravity and then there's this experimental data that you would run an experiment and and see what happens and then you try to see if the theory and the experiment fit. Math was a little different in that it was almost entirely theory. there's there are very very few experiments that you would do purely in mathematics. There were a few for example Gaus famously computed the first 100,000 prime numbers and that was a data set that he used to make predictions. He predicted what we now call the prime number theorem. But science and experiment were the two major forms of science. Then later on simulation came along. but you didn't have to run a a big expensive experiment. Sometimes you could just simulate let's say a hurricane in in a supercomput instead of in real life. And then later on big data came along that rather than than just do a small number of experiments to try to confirm or deny a specific theory, you could take megabytes or pabytes of data and try to discern patterns, try to extract out laws from just massive data sets. and that's a more emerging type of science. But now all these modes of science are being transformed because we now also have AI to to to help us. So so in the past every one of these ways of doing science had to be done by human scientists. You know you had to have someone to perform the experiments or someone to do the theorical calculations or run the simulations or go through the data and you could use computers for some of that but you have to but even then you had to program the the data analysis tool or whatever and you still need a lot of expertise. You know, we have automated labs that can that can perform experiments automatically.
56:53 you can you can get a coding agent to run a simulation for you and you can try to to also run automated data analysis. and increasingly you can also do automated theory. You can take some mathematical problem and ask what are the consequences of these hypotheses and these axioms. What what conclusion should you get? These tools can now be done at scale much faster. he could potentially run many more theoretical analyses than than any one human scientist could. On the other hand, this is not the only thing that we want.
57:23 There's value in doing things a slow way, you know. So, a scientist who spending hours and hours working things out on pen and paper, doing the experiments in the field with with their bare hands and actually debugging the simulations that show up, they often learn a lot of of extra insight beyond just getting the answer that they're trying to seek. They can find they can discover new phenomena. They can see connections. They can see similarities to some some previous thing that's been studied elsewhere in the literature and they can communicate what what they what they are finding to other people.
57:54 So there is this paradox that on the one hand AIS are becoming more powerful and more capable and and making fewer mistakes and they are ostensibly achieving a lot of the goals that we want we think scientists are trying to to to do. They're they're running experiments. They're analyzing data. They're writing papers. But it may be that it comes at the cost of the AI becomes picks up some skill, but but no human scientist gets any better at doing the science. No human can communicate exactly what just happened and and and why this this scientific discovery is interesting, why this proof is new and and and and what features it has and how it connects. We may have to sort of redesign our conception of what science is and what we actually want out of science. What exactly is science for? And and what are we trying to do? And is there a danger that we are optimizing the wrong thing when we are pointing our AI tools at science? One analogy I've given in the past is that science is a little bit like going on a hike. You've heard there's some interesting waterfall, some beautiful waterfall out there. So you decide to to to to go hike with some friends to to to find it, but you need to make a map.
59:03 you got lost. You get lost a little bit, but maybe while getting lost, you discover something else which is interesting and you make a note of it. on the way to to this waterfall, you find an even more spectacular waterfall in the distance. You can't get there yet, but maybe some future hiker will figure out a way to get there, too. And so, there's there's a whole process to get to your goal, which is also very valuable. But these tools, these AI tools, they can be like helicopters that will just fly you directly to this waterfall and you can see it and then you fly back, but you learn nothing about how to get there. You you you may not see any other interesting phenomena than the specific thing that you asked for. And so even though technically you achieve your goal much more efficiently, there may be something that that is lost. Modern AIs are powered by a a type of algorithm known as machine learning, which is trying to predict patterns in data. So a very simple example of machine learning is regression. So if you have some inputs and some outputs like let's say you observe that if you feed some animal more food they get they get bigger right? So you can plot how much food you give various animals and and you plot their weight or something and you get some dots on on a graph and if you're lucky they will they will fit some line and then and that line becomes your prediction. So then if you give this dogs this sort this much food you they would gain this much weight. Now in the real world you don't always get these nice linear relationships. Often there are many many inputs and there's many many outputs and the relationship can be can be really complicated. But sometimes the data has a shape and and there we now have all kinds of clever ways to kind of detect this shape and try to fit curves to to these input output pairs. What large language models which power chat bots and things like that they're just playing the game of naming the next word in a sentence. Roughly speaking, if I say roses are red, violets are blank.
60:54 What is the next word to fill the sentence? You can probably guess the answer is blue. That's an output. You can imagine this giant plot where the inputs are all these incomplete sentences and the outputs are the words you want to complete and you got all these dots in this high dimensional space and you want to fit some curve to it that will try to explain what is the most likely word to come out. Sometimes there's more than one answer.
61:16 Hello, my name is you know there could be many many names you could put after the end of the sentence. So you don't always get a single answer but you could try to get the most plausible answer. People have tried this and and you know the the autocomplete feature on your phone does this you know like you you text something and it will suggest the next word and sometimes it's kind of right sometimes it's silly. Once you have any kind of operation like this, it creates some dynamics and you can just keep you know many people have just played on their phone just press autocomplete over and over again and but you get these these gibberish sentences okay you get monkeys typing on typewriters but the the magic of LLMs is that if you train these LMS on enough data okay so trillions and trillions of data points and you you you really try to fit as good a curve as possible and this takes like millions and millions of dollars of of computing hour and months and months of of time, then suddenly even when you iterate, it stays coherent. It begins to sound not like monkeys, but it actually sounds like a human speaking. And somehow we don't fully understand why that's the case. But what seems to be true is that language like like English or other natural languages contains a lot of hidden patterns that we're not consciously aware of. I mean, we know some of the laws of English, you know, there's laws of grammar and things, but there are there are sort of unspoken, unwritten rules of of language that humans pick up. You know, a human child, even though they're not taught, you know, what a noun is, what a verb is or whatever, they can they can pick up what order English words go in. And just by continual exposure to the language, it seems like you can teach these models to also pick up patterns in language to the point where you can give them math questions. the answer to 2 plus 3 is and they will they will say five. They have been trained to to get the the correct answer to to at least simple math questions. Once you have a little bit of ability to to speak English, you can kind of go in loops and and sort of check your work and make fewer mistakes and you can prompt these models to to to proceed step by step and and and not say something unless it's been double checked and so forth. And so they become a little bit smarter quote unquote to the point where they can solve many many complicated tasks. but they're still just guessing the next word to say. It's not really grounded in any deep understanding of the real world. It's just that they they just have seen the patterns in in the English language or other language that they've absorbed so well that they can mimic people who are speaking in intelligent fashion and they can present as being intelligent long enough that they can fool us but long enough they can actually do useful things. So you know we can now solve certain math problems by asking the LLM to provide a proof and sometimes the proof is complete rubbish but if you loop it enough and you have enough checks you can actually start having a positive success rate. so it's it's a very strange way of solving problems like it is it is like completely orthogonal to the way we normally think of intelligence as being very grounded methodical thinking first principles. You know, it's like having someone who knows a lot and is but is slightly drunk and is sort of throwing out ideas, but with enough guidance, you you can actually extract useful output. It's not the most advanced mathematics out there actually. but you give it a lot of data and a lot of time and a lot of other band-aids and things and it actually works pretty well. I find that some of the debate on AI's role in in science and other disciplines is we often default to a one-dimensional view of thinking like this. there's there's easy tasks and and and hard tasks and very hard tasks and and and humans are can can do tasks up to a certain level and AI can do task to a certain level and so which one is which one is better that's kind of a one-dimensional way of thinking but what I found when when sort of working with AIs and and comparing their way of solving problems to humans way of solving problems is that they are really quite complimentary human experts will will focus on depth you know so like a mathematician which will solve thousands and thousands of problems they can work on. But they will pick one or two problems that they think are are difficult but not so difficult that they're impossible but they're difficult enough that the the exercise of trying to make a bit of progress towards them will reveal all kinds of insights that they can share and maybe their students or some other collaborators or or other people can build upon what they do. When we point the AIS at really difficult problems where none of the standard techniques apply, they are still very very bad at I mean they're just randomly guessing but they excel at breath. So if you if you point them at a thousand problems of various difficulties now some may be just too hard but there will be some which actually are within reach of existing methods and there's some method out there in the literature which will solve your problem or maybe you have to combine together two separate methods and it's just that there there just not enough human experts to look at all these problems and the human experts that do look at these problems they may not realize that there was this obscure paper from a journal in 1970 that actually has the key idea that will solve this problem. They don't have the patience or the time to sort of go through all the different combinations of how which technique might work on which problem. But the AIS, you know, they will somewhat randomly they will take sort of educated guesses as to what techniques might work for a problem and some of them will be stupid and but some of them might work and through all these combinations we're finding that sometimes they can catch they can catch a solution that that the rest of all the humans have missed. Occasionally the consens the conventional wisdom on of of the experts is wrong. You know that we all think that that a problem has a positive answer but actually the has a negative answer and we just didn't look at the negative case too much because we thought everyone thought that the the answer was was true but an AI may not have that preconception. So sometimes the AI just serves as an independent pair of eyes. And so some problems that we thought were very difficult had a surprisingly simple solution which in retrospect we should have as maybe we should have gotten ourselves to. They're beginning to become successful at when you point them at a very broad range of problems and they solve some percentage of them. Like maybe you point them at a thousand problems and they solve 5% of those problems. That's still 50 problems solved. you can already have tools that in some sense outperform humans mathematicians by raw number of problems solved. now the 50 problems that get solved may not be the 50 problems that you most want solved. They could be 50 random problems. but still it it is it is very impressive. What I think we will have to do as a profession is find ways to to incorporate this new capability to to solve some problems at broad scales and somehow figure out how to to to make that mesh with our existing capability to solve a few deep problems very slowly. Kepler's story of how he found his famous laws of motion is a is a fascinating one. It shows how important the process is. Kepler learned of Capernacus's theory of the motion of the planets. And Capernacus had roughly worked out how far the Earth was from the sun, how far Mars was, and so forth.
68:18 And Kepler noticed that the ratios of these orbits looked a little bit like the like certain ratios that showed him in geometry. And so eventually he proposed that actually if you take spheres one sphere for every planet and he had six planets known at the time that he could inscribe five platonic solids you know like a docedron and a cube and a tetrahedrron and so forth between these six spheres and he thought it would get a perfect fit and this explained the shape of the solar system in terms of the five platonic solids.
68:46 This was his beautiful geometric idea. It was only after he managed to get his hands on some really high quality observational data of Taiko Brahi which he had to fight for actually and possibly even steal and he tried to fit it his theory to this this data and he found that it didn't actually quite fit that with the precision that Taiko's data offered he could not quite get these spheres to fit and in fact he discovered from that process that the the orbit of Mars and Earth could not be circles at all that there had to be some other shape. He spent many years figuring out what to do. I think I I don't know how long he held on to this this theory of of the platonic solids and you can see in his writings he tried many other things. He tried to to make the circles off center and at some point he landed on the ellipse and then suddenly everything fit. It does show that there is an interplay between theory and experiment. You know that that you can pose a theory if it doesn't fit the data it's it it may not be a good theory. But it's it's more complicated than that too. Before Kepler, one of the criticisms of Capernacus's theory was that already Capernacus acknowledged that that his measurements were worse than the best predictions available at the time. So the best models were the geocentric models which had been developed by the Greeks and then by the Arabs and Indians. There were many many adjustments and fine-tuning and they had a very very precise model that could predict in a very complicated way where all the planets would be. and Kepler Capernacus's model was worse. Just knowing agreement of data is not necessarily the the only metric. It was only after Kepler found his his revised model where the orbits were not circles but ellipses that the heliocentric model became more accurate than the geocentric model. What this tells you is that is that science is you can't always get instant feedback as to whether you've solved a scientific problem or not. If Kepler and and Copernicus had AIs and they asked them to predict a model for for for the universe, it could be that the AIs that generated the correct kentric model would be discarded because initially their predictions were not as good as as as the geocentric ones. It takes time to really digest all these theories and see how they fit with everything else that we know about planets and motion and and gravity and everything. One concern actually is that AI are too fast.
71:06 there's a danger that these AIs will do what's called overfitting and and create a very complicated model which is not which has nothing to do with what's actually going on but just fits your data extremely extremely well but it doesn't extrapolate beyond that that that data set how we incorporate AI into the scientific discovery process will be a challenge it can certainly accelerate individual steps of of the process you know you can make experimentation faster you can you can write code faster you can write your papers faster but size as a whole may not necessarily accelerate just because every single component gets faster. There's a danger that that we will optimize the wrong thing when we when we point AI at science and we will on paper get all these amazing successes and but find out that science is not actually advancing as it in the way that it used to. But we will find out it's still better to have these tools and not have them. But we're still learning how to use them most efficiently. Part of what we do is is we solve problems and and we we try to find solutions to problems and and and prove things. and proofs go through a certain life cycle. First of all, you have to to to generate a proof for solution. And this used to be quite hard, but but some of the proofs that you generate are incorrect. so then you have to to verify them, check which one check that is correct. and that also used to be quite tedious. But both of these of these tasks are becoming more and more automated. So we we beginning to see more and more proposed solutions to various problems and many of them are actually correct but proofs are also getting longer and when when when they're written by AIs they are often not very pleasant to read an AI generated proof might might spend a lot of time talking about something very trivial and spend very little time talking about the most interesting portion of of of the paper I think because the AI can't distinguish sort of what is hard what is difficult because by brute force some everything takes the same amount of time for them a human who has sort of naturally had to struggle at the most difficult step of a paper would naturally spend a lot of time on that step. And so you need to write out the paper in a proof in a way that it reads well and it can be explained to other people. And then other people have to get excited by it. they they they have to accept it as oh this is really interesting that this this will help me solve my own problems or it really clarifies why this phenomenon was true that it has to be accepted and this is where we we tra traditionally have the peerreview process where we we send papers to referees and if the referees are are excited by this result then the paper get accepted but you know there could be papers that are technically they are correct and and they are readable and fine but but they could answering a question that no one cares about. And then finally, it needs to be sort of completely polished and put into textbooks and and taught to students. And often the first version of a proof is not suitable for writing on textbooks. It often is is done in a very inefficient way and is not the the ordering of steps is not quite logical.
73:59 There's a certain digestion process where where someone has to spend a lot of time thinking very hard to what is completely the the right way to organize to edit the paper to sort of flow in the same way. a little bit like how you would you would edit a documentary or a movie. And so what we're finding is that AI tools are accelerating the early stages of this process but not the late stages. We are now generating many proofs. We are verifying a bunch of them but the pace of understanding them and putting them into the final textbook form is still done by humans. and in fact, we're now experiencing what you might call proof indigestion where suddenly there's lots and lots of pending solutions to problems that should be understood and should be going to textbooks, but they just we're just flooded now with with too many of them.
74:45 and we have to pick and we have to triage. And this is something that has never had to happen before. it used to be that solutions came out so rarely that if a solution to a major problem got got solved. All the experts would sort of drop everything and read it and and try to digest it as quickly as possible because it was so rare and so valuable that it was worth doing. And now we're just getting flooded with with all these possible solutions. I myself, you know, I've had to stop trying to stay current with all the latest developments in in in my in in my field.
75:14 Sometimes there's just so much going on. I I can't promise now to read every single development that that that shows up. I mean, this was already beginning to be a problem even before AI, but AI has really accelerated the sheer volume of of content being generated. And so, we're going to need much better curation and and filtering. It's a good problem to have. I mean, it's like it's better to have a to have too much food, more food than you can eat than than not enough food to eat. But it is still a problem.
75:42 AI have become increasingly capable in mathematics. For people like me who have been following the developments for the last three years, there's been a kind of steady progression, you know, so four years ago they could solve middle school math problems and then they could solve high school math problems and then high school Olympia level problems and then some problems you like graduate student level qualifying exam problems and then they're starting to solve a few of the the minor unsolved problems that maybe someone like Paul Erdish would have proposed but no one really looked at. So it's a lot of low hanging fruit. And then just recently there's been one or two occasions where they they managed to solve some problems that people actually really did try very hard to solve. Somehow collectively the humans all had were taking the wrong turn and the AIS which had a different set of biases had managed to cobble together a solution which was quite clever and has been quite this already has some impact. there's been some nearby problems to the union distance problem for instance which have not also been solved by humans who have adapted the the AI's technique. I found that quite exciting. so I think for some my colleagues it was very concerning especially if they hadn't been following the previous developments and and and not realizing that this was where they were at. If a colleague had only seen say what chat GPT could do in 2023 and if you asked it a difficult math question then it would give you complete rubbish and they they are quite different now. I still fundamentally the same technology but but they have found ways to reduce the error rate and and become genuinely useful. Now it's still unclear how replicatable this is. Many of these achievements they're done by private companies. They they not disclosing how much resources they spent to to use. I mean is it was it $100,000 of commute compute? Was it a million dollars? We don't really know. And we don't know their success rate. was this problem that they solved the only problem that they looked at or did they look at 10 problems? They look at 100 problems. While the results are impressive, we don't have enough data to to really gauge whether this will become a completely regular occurrence going forward or whether it's only if you spend $100,000 over several months with a team of 10 people that you can get results like this. And maybe it's only 1% of all problems that we care about are immunable to this method.
78:01 yeah we don't know but there are efforts to more properly benchmark in a scientific way. The most recent of these challenges is called the first proof challenge. They tested the latest models against a test set of 10 research level questions and the best models could do like five or six out of 10 of these sort of medium level difficulty math problems which already had a solution but the solution was kept secret that I think there's a lot of routine tasks tasks that we we do every day as a as a as in our research that some percentage of those can now be done by by AIS it could be expensive u many of of these tools they require say a couple hundred dollars to run to before they can get the solution and sometimes they fail.
78:41 They they spend all this compute and and they end up with with nothing useful. We are seeing now in programming that many expert programmers are reporting that their ability to write code has increased by a factor of five or 10 or 100 with these with these tools. But they are also they can they can feel themselves learn losing the ability to code by hand and sometimes they cannot review the code that that that comes out of these agents. There is a trade-off you know speed and is not everything. I very much like the collaborative aspect of mathematics. Something that I didn't realize was so important until relatively late in my career that like a lot of the mathematics I've learned I I learned after grad school by by working with with mathematicians and and and scientists in different fields. I teach them what I know, they teach me what what what they know and I become much broader as as a as as a result. When you're working with a collaborator that you've been working with a long time, there's a point where you become almost mentally attuned. Perhaps you've been familiar, like if it's a really close friend or family member that you talk to for a long time, sometimes you can complete each other's sentences. You know what the other person's going to say? And you can sometimes get that when you collaborate. You can you can throw out an idea and before you even finish the sentence, the other person gets it and and can run with it. People have tried to use AIS like this and AIS they are you can't converse with them. Of course they make mistakes. Sometimes they they sometimes will be psychopantic and and only tell you what what you want to hear but also the interactions right now with these tools are kind of personal. Like I've tried to collaborate with people in person and also have an AI present but breaks up the rhythm. these these tools they don't they're not really conversational like they're not as fluid conversation as as as as they are with human collaborators yet. maybe they will get they will get there until recently they don't learn from your conversations you know so with a collaborator when when you you know you can resume the next day you can pick up very quickly and sometimes even for a call that you haven't met for years you can pick up some very old threads AI have a certain amount of context and they can remember some things and and they can record they can make notes and kind of simulate this memory but you can't attune to an AI the same way that you can to a really close collabor better yet. I actually don't use these tools so much for the actual problem solving process. I found to date that these tools are much better at secondary tasks like like doing literature searches or or checking a proof or writing some code, proof reading something that I wrote to see if there's any opportunity to make things a little bit tighter. I find yeah the rhythm of of working with an AI is not quite the the rhythm I I prefer working with a human collaborator but that could just be the the current state of current technology. Maybe future AIs will be much more conversational and and much more human to interact with.
81:37 So we're at a somewhat risky point in in sort of the the structure of of funding the scientific enterprise in general because on the one hand these tools are allowing us to create the outputs of science or what seems to be the outputs of science at at a much accelerated rate but it could come at the cost of of nurturing our seed corn for the next generation of scientists. For example, there there is a real concern that the training problems that we give our graduate students to work on to as as their as their first projects to to get a little bit of of of recognition and and and career training and experience. These are the types of problems which now AIs can they can replicate many of the papers. But if you replace the grad students by by these AIs, you know, the AI will generate these grad student papers, but then we won't get the next generation of of of students. But if we if we don't continue this process of digesting, you know, all this AI output to build the next base of knowledge for the next generation of humans and AIs to to build upon that, we may end up stagnating as as as a a scientific society. You know, that we will be able to optimize everything that we can we can do with our current technology, but we may not actually develop really original new ideas anymore. we will need to to really have a much more open discussion about what basic science is and what is useful for and why it's important to still have curiositydriven re research. why we still need a community of of of of humans to explore things sometimes slowly, sometimes you know in in in ways that are not as efficient as the latest model AIS and also the the insights that that we that we we we gain. I think we should share them more and we should do more outreach to to the general public. I think the general public today, you know, they can see the visible outputs of science. You know, they have a cell phone, they have the internet, they they they have GPS or whatever. Many people, they they don't see the the whole process and and how a basic understanding of math and science actually makes the world around them a lot less scary and just a a lot a lot clearer. I think a lot of people now are just living in in a state of anxiety. The world is so complicated. We haven't emphasized these sort of softer values of science as much as sort of the hard you know like technological outputs and things but science does add a certain amount of clarity to to to one's thinking and you know these are valuable things and they need to be supported.
84:14 Four times a year, we print a magazine worth putting your phone down for. Big Think's premium print magazine is built with custom artwork and expert insight from the world's biggest thinkers. No AI slop, no rage bait, just big ideas for curious people. Support the media you want to see in the world. Go to bigthink.com/membership to join.
Summary
- The six essential pillars of mathematics are numbers, algebra, geometry, probability, analysis, and dynamics.
- Numbers allow for precision in communication and are fundamental to civilization, enabling trade and taxation.
- Algebra abstracts numbers into variables, facilitating the study of operations and their properties.
- Geometry helps measure and understand spatial relationships, crucial for navigation and architecture.
- Probability quantifies uncertainty, essential for decision-making in various fields, including finance and medicine.
- Analysis deals with limits and infinities, addressing inaccuracies in measurements and understanding complex systems.
- Dynamics studies change over time, revealing emergent behaviors in systems like traffic or biological evolution.
- AI is transforming mathematics and science by automating tasks, but it raises concerns about the depth of understanding and the nurturing of future scientists.
Questions Answered
What are the six essential pillars of mathematics?
Terrence Tao introduces six fundamental concepts in mathematics: numbers, algebra, geometry, probability, analysis, and dynamics. He emphasizes that these concepts, while basic, have evolved into sophisticated mathematics over time.
How does probability relate to real-world uncertainty?
Tao discusses the difference between sanitized mathematical problems and real-world unpredictability. He explains that while exact predictions are difficult due to uncertainty, understanding the frequency of outcomes is essential, a realization first made by gamblers.
How has mathematics improved weather prediction?
Tao highlights the achievements of atmospheric scientists in reducing error rates in weather forecasting through the application of dynamical systems. While predictions are not perfect, they are significantly more reliable than mere guessing.
What is the significance of sphere packing in modern technology?
Tao explains how the mathematical problem of sphere packing relates to digital communication. Efficiently packing signals in high-dimensional space prevents interference, making it crucial for technologies like cell phones.
How is AI changing the landscape of scientific research?
Tao discusses the increasing role of AI in automating various aspects of scientific research, from experiments to data analysis. However, he cautions that while AI can accelerate processes, it may also overlook valuable insights gained through slower, manual methods.
What are the potential pitfalls of using AI in scientific research?
Tao warns that AI's speed may lead to overfitting and the creation of overly complex models that do not accurately reflect reality. He emphasizes the need for careful integration of AI into the scientific process to ensure meaningful advancements.