Section Insights
Introduction to the ACS Lecture Series
What is the significance of the American Chemical Society's 150th anniversary?
The ACS's 150th anniversary highlights a century and a half of scientific discovery and innovation, emphasizing the importance of chemistry in improving lives globally. The event serves as a platform to explore emerging frontiers in chemistry and acknowledges the contributions of chemists over generations.
- The ACS has a rich history of scientific contributions since 1876.
- The anniversary is a celebration of past achievements and future possibilities in chemistry.
- Human intelligence remains central to driving innovation in the field.
Historical Context of AI in Chemistry
How has AI been applied in chemistry since the 1940s?
AI applications in chemistry began in the 1940s with projects like Dendril, which assisted in analyzing mass spectra. The 1960s saw further developments with programs like Lassa and foundational work on quantitative structure-activity relationships, establishing a basis for AI's role in chemical research.
- Early AI projects like Dendril and Lassa laid the groundwork for modern computational chemistry.
- Quantitative structure-activity relationships have been pivotal in understanding chemical properties.
- AI's integration into chemistry has evolved significantly since the mid-20th century.
Data-Driven Synthesis Planning
What challenges exist in data-driven synthesis planning for complex chemical reactions?
Data-driven synthesis planning has become a vibrant field, but it faces challenges, particularly with complex target molecules. As the complexity of the input increases, the performance of synthesis planning tools tends to decline, impacting drug discovery and the synthesis of natural products.
- Complexity in chemical structures poses significant challenges for synthesis planning tools.
- The field has seen a rise in open-source contributions and startups focused on synthesis planning.
- Understanding the limitations of current tools is crucial for advancing drug discovery.
Advancements in Predictive Modeling
How are predictive models evolving in medicinal chemistry?
Recent advancements in predictive modeling allow for the mechanistic prediction of reactions with high accuracy. These models build on historical frameworks, such as quantitative structure-activity relationships, to enhance our understanding of structure-property relationships in drug design.
- Predictive models are becoming more sophisticated, enabling better understanding of chemical reactions.
- Historical frameworks continue to influence modern modeling approaches in medicinal chemistry.
- The evolution of these models is essential for improving drug design and discovery.
Exploring New Chemical Scaffolds
What role do computational methods play in discovering new drug scaffolds?
Computational methods are crucial for identifying new chemical scaffolds that may have been overlooked. These methods facilitate the exploration of chemical space, leading to the synthesis of novel compounds with potential medicinal utility.
- Computational pipelines help uncover previously overlooked drug scaffolds.
- The integration of computational methods is transforming the landscape of drug discovery.
- Identifying new structures is key to advancing medicinal chemistry.
Transcript
0:00 The benefits of science will be even more spectacular than we can ever imagine. >> >> I wanted to do something that is of long range. Strange benefit to humankind.
1:19 >> >> Practically everything we touch in our daily lives has been developed and improved through basic research.
2:21 >> >> Heat. Heat. Heat. Heat.
3:00 >> >> Science is taking us on an ever accelerating journey into the future. Please join me in welcoming to the stage 2026 ACS President Rigabberto Hernandez. >> >> Goosebumps. Every time I see the video, I just get goosebumps.
3:49 It's amazing the inspiration of what the Cavi of Cavi did and what the Cavi Foundation continues to do. Let's give them a round of applause. Good evening everyone. I am Rioberto Ernnandez, president of the the American Chemical Society, and it is my sincere pleasure to welcome you to the ACS lecture series sponsored by the Cavi Foundation.
4:22 As we celebrate the American Chemical Society's 150th anniversary, we are reminded of a century and a half of scientific discovery, innovation, and service to society. Since its founding in 1876, ACS has brought together generations of chemists whose work has expanded our understanding of the world and improved lives around the globe. It is especially fitting that during this milestone year, our sesy centennial, we gather to explore emerging frontiers that will help shape the next era of chemistry.
5:06 Whether the tools for discovery involve experiment, theory, simulation, or artificial intelligence, the one thing that will remain central is human intelligence. It is because of every one of you that chemistry is needed and you are the ones that are driving innovation. Our fellowship through the ACS has brought us together and will continue to do so for the next 150 years. That's what we celebrate today.
5:41 Let's applaud each other. ACS is deeply grateful to the Cavi Foundation for enabling us to highlight the achievements of outstanding researchers in the chemical sciences through this lecture series. This programming reflects goals that ACS and the Gav Cavly Foundation share. Advancing science for the benefit of humanity, promoting public understanding of scientific research, and supporting scientists in their important work. The cavly emerging leader in chemistry lecture spotlights exceptional early career scientists under 40 who have already made outstanding contributions to scientific or engineering research.
6:32 This evening we have the privilege of hearing from one such outstanding leader. It is my honor to introduce Connor W. Kohley, associate professor at the Massachusetts Institute of Technology. Professor Kohley is an associate professor in MIT's department of chemical engineering and department of electrical engineering and computer science. His research group develops computational methods for molecular discovery at the interface of chemistry and artificial intelligence with applications in synthetic, medicinal, and analytical chemistry. He received his bachelor's in science in chemical engineering from Caltech and his doctoral degree from MIT followed by a post-doal training at the Broad Institute. His contributions have been recognized through honors including it's a long list. Chemical and engineering news talented 12, the NSF career award, the Bayer Early Excellence in Science Award, the Camille Drifus Teacher Scholar Award, selection as a Schmidt AI 2050 early career fellow, and Samsung AI researcher of the year.
7:48 He currently serves as an associate editor for the journal of the American Chemical Society. Tonight, Professor Kohley will share his lecture AI and the expanding computational toolbox for chemistry. His presentation will explore how advances in artificial intelligence, chemiratics, and optimization are expanding the scale and scope of problems that chemists can approach computationally with examples in synthesis planning and therapeutic discovery. He will also reflect on the importance of critically evaluating new technologies in a field shaped by substantial hype but meaningful advances.
8:32 As ACS marks its 150th anniversary, his lecture offers a timely perspective on how the continuing integration of chemistry and computation is expanding the way chemists formulate and investigate scientific questions. Please join me in welcoming Connor W. Coley small to the lies we're shaping every bond, every reaction. Science moving us into actions.
9:10 All >> right. Good evening everybody. First want to thank ACS and the Cavley Foundation for this recognition and for the opportunity to share with you some of our work today. I've been coming to ACS national meetings for a decade now and I've always enjoyed these Monday and Tuesday evening lectures sponsored by the Cabley Foundation. So, it's fun to be on the other side of it for the next 45 minutes or so. Before we get into the science, I'd like to acknowledge the current and former group members that have of course made all of our work possible. very fortunate at MIT to have a group of students and postocs who span computer science, chemistry, material science, biology, chemical engineering disciplines to work together on challenges at the interface of all of these techniques in the chemical sciences.
9:58 We also have a very creative and energizing set of collaborators who I've listed on this title slide and I'll try to highlight them throughout the talk as well. So the subject of today is artificial intelligence and its role in the computational toolbox for chemistry. As was mentioned, there's a lot of hype about artificial intelligence. It's sort of a loaded phrase these days, but there's also a lot of substance as I hope you'll be convinced by at the end of the talk.
10:25 There's a lot that can be said and very little of it can be said within 45 minutes. So, I'm going to try to talk about the brief history of AI and chemistry and highlight some of where we've tried to contribute to these different topics in recent years. We are here on the 150th anniversary of the American Chemical Society, but it's been more than 150 years since the field started reasoning quantitatively and formally about molecular structure and their function.
10:55 So in 1874, Kaye and their contemporaries recognized that just by rearranging a familiar set of elements subject to veance constraints, you can access huge numbers of structures and became very fashionable to count these structures. Much of this work formed the basis for later developments in graph theory. But beyond just counting possible molecules even before this in 1868, Crumb Brown and Frasier made an observation that the properties of closely related chemical structures tended to also be closely related.
11:30 Because it was the 1860s, the way they were studying this was by counting the number of grains of substances it took to poison rabbits. But irrespective of what they were doing, it led them to this very succinct mathematical observation or analogy that if we think about the constitution of a molecule as some C and its physiological function as some fi, we can think of this as a formal mathematical relationship or we can write some f of c to represent that structure property relationship.
11:58 So these two articles and others of the time right predating ACS certainly predating any form of digital computer laid the foundation for much of the field in chemiratics and computer aided molecular design. If we fast forward to the 1960s we get to quite an important decades in AI for chemistry. Artificial intelligence itself had its foundations and formalization in the 1940s and50s and its application to the sciences and to chemistry in particular came very soon thereafter.
12:32 So one program that emerged from this decade was Dendril. This was a NASA sponsored project coming out of Stanford and the idea was to assist in the analysis of mass spectra observe peaks in electron impact mass spec as it was called at the time make inferences about what fragments could be present and assemble those into plausible structure hypotheses for what those compounds could be. Another example from the 1960s that you may be more familiar with is Lassa.
13:03 So alongside EJ Corey's formalization of retroynthesis as a conceptual framework was the formalization as a computational framework. If you take a computer program and you teach it these handcoded rules of chemistry, it can explore the vast search space of multi-step synthetic pathways. Then in 1964 we also saw sort of a landmark paper from Honch and Fuja that laid the groundwork for quantitative structure activity relationship modeling. So mathematically trying to relate the substructures or substituents of a molecular structure to the properties it exhibited. They also worked on toxicity and in this case actually used descriptors that came three decades prior from Hammet relationships.
13:49 So all three of these research threads really kicked off the field of AI for chemistry in the 60s and these questions sort of persist to this day and they persist in my research group at MIT. And so for this talk, I want to step us through these three different sub fields, talk about some of our recent contributions to these spaces. We're going to start off with analytical chemistry in this long-standing challenge of structural lucidation. We'll move on to synthetic organic chemistry which encompasses synthesis planning like retroynthesis but some related tasks as we'll see and then move on to medicinal chemistry and understanding how we can design new molecular structures to exhibit the kinds of properties we find most useful in therapeutic applications.
14:34 My hope is that by having sort of a few different vignettes spanning these different areas, there'll be a little nugget of something for everybody because in this room we span many many different subfields and ACS of course spans many different subfields. As an aside before we jump into it, I'll argue for the purposes of this talk that AI models statistical learning models do one of two main things. The first category would be these forward models which we might equate to property prediction models or discriminative models. We can write these as Crumb Brown and Frasier would have and say that we're trying to learn some mathematical relationship f ofx where x may be a molecular structure represented in a number of ways.
15:20 If we're trying to represent this in a way that accounts for uncertainty, we might record this as a probability distribution and write it in that alternate formats, but the difference doesn't matter so much for today. And aside from these forward models, we're going to have inverse models as well. And so people call these also design models or generative models. And the learning objective for these models is different. It's to learn a complex distribution over some structured object. So it could be as Kaye was doing in 1874 just learning what normal molecular structures look like in an unconditional manner. But it could also involve learning conditional probability distributions where we'd like to specify the properties we want our structures to exhibit and we want to sample then new molecular structures that exhibit those properties.
16:13 On a sort of footnote here, maybe that's large language models feel a bit qualitatively different than most statistical learning techniques people may have seen in the past. But even large language models are just another instantiation of generative models, right? Conditional token prediction for text. So starting off with analytical chemistry, I mentioned Dendril as one of the earliest applications and Dendril is often referred to as the first expert system not just for chemistry but for all of science.
16:44 And so what Dendril was doing as I briefly introduced before is it perceived a mass spectrum in this case an EI mass spectrum. It made some inferences about what fragments could be present. It did a combinatorial enumeration of different ways those could be recombined. It then predicted how those candidates would fragment themselves and it would compare the predicted spectrum to the query spectrum. And the first versions in the in the 1960s used handcoded expert rules and then a decade later in the 1970s this was relaxed in metadril to allow them to be inferred directly from data.
17:20 It's this inference from data that we pick up on. We're going to change massspec modalities though from electron ionization massspec to what's traditionally used with LC tandanda massspec and collision induced association. And we're going to revisit this problem with the benefits of large standard reference libraries that contain tens of thousands of chemical structures and their fragmentation patterns acquired under relatively standardized conditions. And this data will be the basis for our models.
17:51 And so the setting that we're most interested in here is one where we have a complex mixture. We analyze it through LC tandom mass spec. So we get a first separation on the basis of polarity, a second separation on the basis of the precursor mass. And then we have a second stage of mass spec where we fragment the molecule and we observe these fragmentation patterns. And in principle these have a correspondence to small molecules that led to these fragmentation patterns.
18:19 In essence, what we're trying to learn and solve is this birectional mapping between the fragmentation pattern and substructures of the molecule that accounts for them. And although this has been a topic of research for 60 plus years at this point, it's a very hard problem. If you were to take a sample of human plasma or another boflu and analyze it through such an instrument, more than 85% of the features that you see would be completely unidentified.
18:46 This is in very stark contrast to the field's ability to sequence nucleic acids and understand what DNA and RNA is present in our bodies. So the way that we approach this is by focusing on the forward problem. We want to try to train models that learn how to predict the spectrum of a molecule from its chemical structure. But if we're asking machine learning models to predict spectra, we have to be very precise with what a spectrum even is.
19:14 Computers are used to dealing with integers, floats, lists, and strings. They're not used to dealing with spectra as structured data objects. So on one end of the spectrum, we can consider a spectrum to be just a discretized set of values. So we can represent a spectrum as a vector. We can define intervals for our mass to charge ratio. And then this is a nice wellposed numerical task. We can recognize instead that actually the mass to charge ratios at which we observe these peaks should corresponds to additions of elemental masses. So we can attribute each of these peaks to a specific chemical formula or naturally one step further recognize that what's actually observed are charged molecular fragments and we can account for these peaks by thinking explicitly about the structures that would have led to that particular M over Z value.
20:10 Now, as you move from left to right on this slide, you can think of this as a greater incorporation of chemistry knowledge or at the very least a greater alignment between the computational formulation in the physical process that we're trying to model. So, we train various types of geometric deep learning models to analyze molecular structures and predict this fragmentation recursively. We build up these fragmentation trees and we can learn how these map onto actual mass spectra. We try to account for you know adducts collision energies and some of the other subtleties associated with this data modality.
20:46 The end result is that we have a model where we can take in a new molecular structure that's never been seen and we can predict how it would look in our spectrometer. This mirror plot shows the prediction on the top and an experimental reference on the bottom. And because our model is grounded in the physical fragmentation process, we can attribute each of its predicted peaks to an actual substructure. Now, quantitatively, this grounding does also help us. If we compare what I'll loosely describe as this chemistryaware approach with a blackbox deep learning approach, we do see sort of marketked benefits in terms of elucidation accuracy or performance on some standard benchmark tasks.
21:26 So, we have this newfound capability to predict the fragmentation of molecules with accuracy that we had not had access to previously. And we can go back to some of the long-standing elucidation challenges that the field faces in a few different areas. So with Clary Kish at the Broad Institute, we're working on the identification of novel biomarkers of diseases and have confirmed experimentally several of these from clinical cohort studies with Desire Plots at MIT. We're looking at solucidating new degradants of pesticides that might accumulate under dark conditions and aren't prone to hydrarolysis.
21:59 We've tried to revise reaction screening protocols where pulled reaction screens are typically restricted to nonisobaric mixtures of products. But with tandem aspect, we can disambiguate and elucidate more complex pools. And with the Jin Wong at Nor Eastern, we're thinking about biosynthetic pathway elucidation where these types of predictive tools can help us pinpoint where the site of oxidation might be when we have SIPs acting on complex natural products. And all of these capabilities and all these applications stem from this core ability to mimic the physical fragmentation process trained through machine learning models.
22:40 So dendril started out as a search over chemical space conditioned on a spectrum. With analogy to that, MASA could be thought of as a search over multi-step synthetic pathways subject to information about singlestep transformations. Expert encoded synthesis planning tools rely on definitions of reaction templates that look a lot like the schemes that you find in textbooks of named reactions. These are general patterns that define what transformations are possible as well as what the exceptions to the rules are. So apply this rule unless rp prime equals this rp prime equals that.
23:18 And this notion of expertly curating these singlestep reaction templates to apply in a large search is something that barowski has really championed in the past couple of decades. Combining well-curated rules with prioritization strategies led to very successful deployment and commercialization of these ideas in Cynthia. And there's a really strong analogy between expert synthesis planning tools and chess programs. In chess, if the rules are clear, all you need is a search. And that's the strategy that Deep Blue took 1997 to beat the world's champion at the time.
23:56 Chemistry is a little bit different in that we don't have rigid rules that we know. We have approximations. That's why they require curation when you follow this expert approach. But instead of following this expert encoded approach, we wanted to learn these from data as others before us had done. And if we look at databases like SciFinder and Reactis, we do have access to tens of millions of records of reactions that have been run before and published. We have information about what reactants were used, what conditions were used, what the major outcomes were. And it's based on this information that we can start to train models.
24:33 Models that can do retroynthesis. They can take in a product and recommend precursors that could produce that product in one step. We can use the same data to solve the above the arrow problem and recommend reaction conditions. We can also try to predict the products of chemical reactions by flipping the task. We'll revisit that soon. And so this line of datadriven synthesis planning research is not one that we started, but it's one that's become quite a lively field in the past decade.
25:01 And there are now many many different open source contributions, many startups and many very interesting contributions to this arena using reaction databases to learn how to do synthesis planning and presenting these tools to chemists for ideiation and ultimately for selecting what routes might be worth taking into the lab. This is just a video showing our own ASOS program that pulls together a lot of the different techniques that we've developed at MIT over the years. Now, there is a problem with these datadriven synthesis planning tools that we run into that everyone else working on these tools runs into, which is that when you actually start to look at the performance of these methods as a function of the complexity of the target that you're using as input, you see this very clear correlation.
25:51 And it doesn't really matter how you define complexity and it doesn't really matter how you define accuracy. you see this steep drop off where the more complex the input is, the more these programs tend to struggle. This might obviously lead to some conclusions that perhaps for complex secondary metabolites, things that we would consider natural products, this is a problem, but it's also a major problem for routine small molecule drug discovery. So this plot is just mapping the complexity of structures over a 50-year period of newly launched drugs as well as structures reported in Jedcam.
26:27 And there's this nice monotonic increase in the complexity of drugs over time. Nothing in the past 17 years has reversed this trend since it was published. So we have to learn to cope with the complexity of increasingly tricky targets when we do datadriven synthesis planning. One of the ways that we've started to approach this is by reformulating what it is we ask our models to reason about. So instead of operating on actual molecular structures, we can think about teaching our models to reason about pseudomolecules.
26:57 And these pseudo molecules are not quite synths, but they're related. They take after Dave Evans charge affinity patterns. And they let our models operate at this higher level. We abstract away details of leaving groups which one might consider tactics and we focus exclusively on strategy and so many different specific instantiations of a cross coupling for example could be abstracted into a single type of strategic disconnection. Now, the reason to do this is because when we then put in a target that we actually want to plan a synthetic pathway for, the pathways that we're producing are far simpler and shorter than if we hadn't done this abstraction.
27:40 This lets this be a much better position algorithms and increases the success rate for complex targets. Now, this higher level pathway is not something that you can readily take into the laboratory. It's something that requires refinements. And so in this particular work, we collaborated with Sarah Reeseman and Richmond Sarpong and in particular Logan from Richmond's group and Omar from Sarah's group to elaborate this higher level strategy into a fully specified synthetic pathway with ideas of the conditions that could be used at each step. So choosing specific leaving groups, making specific decisions about strategy beyond strategy into the tactics that we had not considered algorithmically.
28:24 And so at this stage, this work is very much a collaboration between human experts and the computational tools where human experts are filling in these details based on their intuition and experience about what makes a feasible pathway. The consideration of feasibility is part of a much larger set of considerations for selecting synthetic pathways. So there are many different ways to judge whether a synthetic pathway is worth pursuing and whether it might be optimal for a given use case. If you're a process chemist, the list of your considerations will be far longer than if you're a discovery chemist.
29:02 I'd argue that the simplest place to start is just feasibility. And one way of phrasing feasibility is that if I try to carry out this pathway as proposed, would I make any appreciable isolable quantity of the product I want? And so this needs to computationally check the feasibility of our recommendations motivates in retrospect the task of forward reaction prediction predicting the products of chemical reactions. And this too has deep roots in the fields. There's been many different programs developed for this since the 1970s.
29:35 Dunji and Ugi in the 1970s began formalizing reactions as transformations of bond electron matrices. Johnny Gastiger had his aeros program which tried to perceive different physicochemical properties and combined that with expert rules. Bill Jorgensson's cameo program behaved similarly acting at the mechanistic level to perceive pKa's felicities and so on. Then a little bit later Sedon Fatu developed some of the first datadriven approaches to learning from reaction databases to predict reaction products.
30:09 So before we actually talk about some of the models that we've been training to solve this forward task, it's worth thinking about what is the data that we're going to be using to do this learning. Because of course datadriven tools will inherit the biases and constraints of their training data. If you just learn about synthetic organic chemistry by looking at the published literature, you will get a very rosy view of how well things work. Most reactions seem to make their desired product in high yields. Also, for some reason, 90% yield is much more likely than 89% yield. And this, of course, does not reflect the reality of reactions run in the laboratory. Right?
30:49 There's far more of a balance between successful and unsuccessful reactions. But if you just look at literature data, that's not something that you can really pick up with these models. There are both technical and social barriers to why these biases exist. For a few years now, me and others have tried to work on some of the technical obstacles through the open reaction database, which I'll be happy to talk about afterwards offline if if there's interest from folks.
31:15 But instead of looking at the broad literature, another approach that many have followed is by looking at smaller sort of in-house data sets often obtained through high experimentation. Right? Lab notebooks contain plenty of failures in my experience. So we've had collaborations over the years with Ying Wong and ABV working on a few different topics and one of those was analyzing some of the Suzuki coupling data from their internal high throughput experimentation group in a parallel med setting. In the histogram of yields that I'm showing should appear as pretty stark contrast to the literature distribution of yields. So in this data set only 1 of the entries had isolated yields above 30%. And onethird of the entries had yields of exactly 0%.
31:59 So these HGE data sets or other in-house data sets provide perhaps a better way of learning between successful and unsuccessful outcomes. But of course if you train a model on these data that model only knows about Suzuki couplings and so its domain of applicability is quite narrow and it doesn't fill the same role as some of our literature trained models do. So even if we can't successfully train literature models to understand the subtleties of yields, we can at least train models to understand the products of chemical reactions.
32:33 And so we can train any number of deep learning architectures to take in reactant structures and the conditions of the reaction and generate the structures of products that they might form. In these models, if you look at the benchmarks, they look quite good, but the cracks start to show when we evaluate them in slightly more stringent settings. In this case, we're looking at a time split where we train the models up through a certain publication year in the patent literature and then we test them on reactions published afterwards, sometimes decades afterwards.
33:06 And here you can see this precipitous drop off in accuracy. And this becomes even worse as you might expect when we hold out entire types of reactions. So we challenge the models to generalize to new types of reactions that they've never seen. That does not tends to work terribly well. So our models can fit patterns in our data, but we started to question if they're really learning the underlying chemistry because they're not seeming to generalize the way that we want them to.
33:33 That prompted us to go back to some of the earlier ideas from the 1970s on how we approached reaction prediction which was with a more mechanistic lens thinking about elementary steps and the feasibility of them that way rather than overall functional group transformations. The unfortunate thing is that there is no sci-finder or reactis for mechanistic data sets and so we had to come up with some way of getting them ourselves. And so Junyong Jung, who's now a faculty member at Kumman University, curated mechanistic reaction templates that we could use to imputee elementary steps between experimentally reported reactants and products.
34:14 Mechanisms are all theories to an extent. And so we were trying here just to mimic generally accepted textbook level mechanistic definitions. There are going to be some rough edges and some errors, but we're trying to have the least objectionable definitions as we can for imputed mechanistic data. So, experimentally reported reactants and product and we're filling in the middle. And with that data, we can then train deep learning models to predict mechanistic pathways. This again looks fine on standard benchmarks until we start to test the models in circumstances that aren't quite what they were designed for. So one example is that if we take an off-the-shelf transformer which has been revolutionary for deep learning since 2017, we train it to predict the products of chemical reactions one elementary step at a time. We feed in these two reactants. The model very helpfully decides that it would love to use palladium and so it gives itself palladium and then it decides it would really like a hydroxide and so it gives itself a hydroxide.
35:18 And statistically this is not a bad strategy if you're trying to just mimic what's in the published patent literature. There are a whole bunch of palladium catalyze crossouplings. But this is not what we have asked our model to solve. We've asked it to predict the reactions between these two reactants. And we do not like the fact that it's hallucinating catalysts just so we can make a reaction work. There are a few ways to overcome this hallucination. And the way that we decided to pursue was to go back to the 1970s actually to Evore Ugi's formalization as bond electron matrices.
35:52 So I did not define it fully before and so here I'll just mention that bond electron matrices are a way of tracking where the veence electrons are in a molecular system. So along the diagonal we're essentially keeping track of electrons in lone pairs. Off diagonal we're assuming coalent bonds have perfectly shared electrons. We do have to overlook stereochemistry in this representation. But the advantage here is that we can train a generative model to predict the product of reactions elementary step by elementary step with precise mass conservation.
36:28 We do this by training the generative model to not change the sum of this matrix. If all we're doing is redistributing electrons in this matrix, we're not touching the borders which are the atoms. We're not changing the sum which is the total electron count. We can achieve strict mass conservation and overcome that hallucination. And achieving mass conservation I think is nice because it is elegant. That's the motivating thing for me. But there's also practical benefit in that we can go back to some physics-based methods and use these mass conserving hypothetical reactions and evaluate them now in terms of thermmochemistry.
37:05 And this is working in collaboration with Alexander Isab and Dylan Einstein using machine learned interatomic potentials to predict the evolution of reactions mechanistically and with thousand times acceleration over DFT. We can put in structures like this complex polyen and predicts with full sort of predictions of the stereos selectivity this famous cascade andric acid C just by entering this structure. The coverage of this workflow is limited by the coverage of the potential itself.
37:37 But as that potential expands and as it's able to describe a greater number of reaction systems, so too will this workflow in letting us predict with some fidelity the outcomes of these elementary reactions. Now our third and final chapter is going to be one of medicinal chemistry and molecular design. Right? I mentioned Hanin Fuja formalizing quantitative structure activity relationships in 1964. Actually that same year free and Wilson published their sort of substructure group contribution methods. And what predated Han and Fuja were Hammond relationships which to be honest were kind of the exact same thing but just applied to rate constants instead of toxicity and other sort of physiological properties. hunt for you to repurpose the same substituent scalar definitions that were used by Hammet three decades prior.
38:34 But both of these are in effect property prediction models. They're structure property relationships. They fall into that first category of AI methods. And so the types of structure property relationships that we use these days look a little bit different in their architecture, but the idea is the same. We try to train models to understand the relationship between structure and function. We'll have some hypothetical sort of cartoonified landscape here as we change the structure. What happens to the properties of a molecule?
39:03 If we're trying to maximize this property, the goal will be to take this understanding and try to find the structures that exist at the maxima. And so with structure property relationships, we can apply virtual screening approaches where maybe we have a funnel that starts with hundreds or starts with trillions of candidate structures and we use our model to predict the performance of each of them. we funnel it down into a small set that we can take into the laboratory experimentally.
39:31 But alongside virtual screening has been generative modeling. And generative modeling is very alluring in principle because the hope is that we can take these models and we can directly explore chemical space in a much less constrained manner. not being limited to fixed virtual libraries, but instead exploring freely the entire space of hypothetical structures as Kaye might have done in 1874. We have already had generative modeling for chemistry for decades, but 2016 is when everything was revisited in the modern deep learning era. And so around that time, people recognized that you can take models to operate on sequences that had been developed and matured in the context of natural language processing and image captioning. You could repurpose them for chemistry by using string representations of chemical structures like smile strings.
40:23 And the hope was that we could use these language models to directly generate the structures of the next blockbuster drug. And that of course is a very alluring prospect. But the past decade since has shown us that the reality is a bit different. There are many challenges in applying these models. One of those is synthesizability which we'll come back to in a few minutes. We also have challenges with efficiency and efficacy. So sometimes these models require a lot of trial and error before they find something that's a good solution.
40:54 They're wandering off in chemical space unproductively for for quite a long time before. But the first limitation I want to talk about is that of liant quality. which is sometimes the molecules coming out of these models just simply aren't useful in the context of medicinal chemistry and the specific dimensions along which I'm talking about when I say quality are the existence of good selective non-coovalent interactions so the coarsest sense of protein lian binding we'd like to achieve some shape complimentarity between the ligan and the protein we'd like some electrostatic complimentarity but then there's this whole rich vocabulary of phicophoric interactions and nonco coalent interactions that really leads to good potent selective binding. These are things like hydrogen bonds, pi pi stacking, cation pi interactions and the field has this vocabulary for discussing these and for engineering them manually.
41:47 But these are not considerations that have really made their way into our generative models. So we wanted to elevate the consideration of these kinds of interactions from being something that are applied after the fact just to critique models to having them be something that is inherent to the generative process. So we can think about a molecular structure as existing in a few different representations or a few different views. We have a threedimensional confirmation. We can then think of the solvent accessible surface area as a representation of its shape. We can think of an electrostatic profile as the effective surface charge and we can think about the potential non-coovalent interactions or hydrogen bonding for example that liant could achieve.
42:30 We train a generative model to learn a joint distribution over these views and these representations. We teach a model about the relationship between all these different aspects. We then borrow a technique from computer vision known as inpainting. And in painting, you take an image, you mask out part of it, and you ask your model to generate diverse samples that are consistent with the part that you're holding constant. In this case, we're keeping, you know, the eyes, the hair, the background constant and coming up with diverse facial expressions.
43:02 Well, in our case, we're not dealing with RGB values of pixels. We're dealing with structured objects. So, we're going to condition on the interactions we want the molecule to achieve, and we're going to request that the model come up with a specific threedimensional molecular structure that recapitulates those interactions. And so, the specific implementation that we follow here is an SC3 equarant denoising diffusion model. And what that means is we start with this big messy ball of atoms with nonsense atomic identities. and we iteratively remove noise and adjust them until the model prediction settles onto a reasonably valid not strained threedimensional confirmation.
43:42 And so this is a stochastic process. So we sample it once and we're left with one recommendation. We superimpose that with the query and we can find that it does recapitulate the interactions that we want it to. We can sample it 100 times and find many many different structurally diverse solutions that achieve the interactions we've requested. And this lets us now decouple conditioning to achieve good protein engagement from specific chemotypes or scaffolds and do things like scaffold hopping and mobility hopping.
44:14 But this creativity when you talk about generative AI turns out to be a double-edged sword. So creativity sometimes goes off the rails. And this is something that we've known for a while in the field, which that sometimes when you request models to invent new structures, they come up with solutions that do not look quite right. They do not look like things you want to take into the laboratory and ask your synthetic chemistry colleague to make for you.
44:40 And not to be too defensive, but this is not coming from one of our models, but this is coming from a model that does well on quantitative benchmarks in the field. And so the reason for this discrepancy and the reason for this failure mode is because most generative models have been taught to operate on atoms or on strings or maybe on fragments. But this structure by structure modification is not really how we make molecules in practice. We make molecules through synthesis.
45:10 And so we can try to constrain what it is our models are able to explore by teaching them what they should have access to. We can pull in stock collections of building blocks. We can define couple hundred expert reaction rules like Lassa did. And we can then constrain our generative models to operate within this constrained space. So here we're really teaching our generative model to explore multi-step synthetic pathways or incidentally those pathways lead to a new molecule of interest. But every sample that is generated is a pathway. So we get a molecule and its recipe at the same time.
45:49 If you're familiar with Edamine Real or Wooshi Galaxy or similar make on demand virtual libraries, this is essentially just a generative version of that. We don't have to enumerate trillions of structures. We can simply teach our model the rules and it can explore in a much more flexible way. And the way that we build this model, we can use it to clean up designs of other structures. So if we do have other generative models that have some quirks, we can tidy those up. In this case, replacing a four-membered ring with a tetrazole or replacing a seven membered chyro building block that doesn't exist in our database with one that does. But we don't just have to find synthesizable analoges. We can use this directly for optimization of properties. And so with Brian Shortcut and John Urban, we've been applying this to structure based drug design, looking specifically at the muopioid receptor. And by letting our model propose structures get feedback from docking, we can run this about a million times and find several submicroar hits that validate experimentally with distinct chemotypes and scaffolds and what's been known before. And the potencies of these structures rival what you get with a multi-billion member virtual screen. And so we are achieving this improvement in efficiency that was promised with generative models.
47:01 Now the synthesis awareness is helpful but it does constrain us to what I describe as breadandbut medicinal chemistry but what is synthetically accessible is a moving target. It is not something that is truly static in time. We can take an example of this dazospyron nonane where the first synthesis of this was reported in 1991. Couple decades later, maybe coincidentally coinciding with the escape from flatland paper, we see this rise in the fraction of patents for which this structure is reported. That's what the y-axis here shows.
47:35 We then see it in investigational new drugs. We then see it in approved drugs. And so this is a scaffold that would not have been considered synthetically tractable pre-91 until the first demonstration of synthetic enablement. This is not unique to this one scaffold. is a trend that exists across many different interesting heterosycycles and other scaffolds. You have the synthesis advancements, the demonstration of initial medcam utility and then its appearance as a clinical candidate if all goes well or at least its use in screening collections.
48:07 And so we wanted to ask the question of whether we could use computation to accelerate the process of essentially reproducing these kinds of trajectories, find interesting scaffolds that had not been reported before, enable their synthesis and see if they have interesting utility. And so for this we built a computational pipeline. We start with 10 million hypothetical scaffolds. This is from the GDB project from Jeanlu Ramons. This is a modern instantiation of something that looks similar to what Kaylee did. Just brute force enumerate hypothetical molecular structures that provides the top of our funnel. We then filter out any scaffolds that we've seen before. We use machine learnic potentials to assess stability.
48:52 We impose our own synthetic filters so that we are left with scaffolds that are synthetically tractable. And then we filter out some things that are a little bit less interesting to us in a very subjective sense. So fused ring systems with the fused ring is an epoxide or spycyclopropane. They're just too many of those for us to look through. But we're still left with thousands and thousands and thousands of new hypothetical scaffolds that have not been reported.
49:19 And so I'm just showing 14 of them here. These are substructures that don't appear anywhere in PubCam, Kemble, Reactis, SciFinder, Enamine. But if you look at these and if you have sort of any sort of med eye, these look like completely plausible and heterosycles that could very well be parts of drugs. They're simply scaffolds that were overlooked in the vast sea of chemical space until we started with this computational pipeline to systematically pick out these opportunities.
49:48 And so now this is work that with Masha Elen and Yonas Rin at MIT. We're sort of partway through the synthetic realization of these and hopefully will ultimately demonstrate some medicinal utility and especially saturated scaffolds when achieve some interesting threedimensionality that could be attractive. So there is this really interesting rich toolbox of computational methods that we've tried to work in. I've tried to highlight different applications across sub fields of chemistry. So again, we're looking at structural lucidation in the context of analytical chemistry. For synthetic organic chemistry, thinking about not just retroynthesis, but the related tasks of mechanistic reaction prediction, ultimately reaction discovery. We hope medicinal chemistry and molecular design let us identify new structures that achieve specific non-coovalent interactions or find synthesizable analoges or find scaffolds that have been overlooked by traditional means.
50:49 And for many of these tasks, the data, the representation, and the methods that we're using have changed substantially since the 1960s. But sometimes fundamentally the questions being posed have not. And so I just want to close with sort of a few final thoughts which is that in parallel to all of the different types of models that I've talked about which focused on statistical learning approaches and more bespoke models there's been a wealth of progress in large language models and agentic workflows. These have dramatically changed the convenience of using, orchestrating and developing different sorts of computational tools. So the barrier for applying these methods is lower than it ever has been before. And the barrier might now be moving a little bit more towards having the data and information to act on. This is where organizations like ACS have the chance to be really leaders in the field of enabling this kind of datadriven chemistry research.
51:47 I appreciate that in the introduction this sort of already came up but it is very much still the role of subject matter experts and chemistry experts right everybody in the room to help the field understand what's important to work on and how we can formulate these challenges in terms that a computer can understand without oversimplifying the problems that we're dealing with. We also have to know how to measure success. And this is not just traditional benchmarking, but it's truly understanding when these approaches have made a real difference on the way the field approaches discovery.
52:21 And lastly, people and especially computer scientists really like to rally around grand challenges. And I would argue that the grand challenges in AI for chemistry are no different than the grand challenges in chemistry itself. We should be working on the problems that will outlive the methods that we're using today. Just like the original problems that Dendril and Lassa and Honchin Fuja were working on have far outlived the specific methods that were being used in the 1960s.
52:53 So I'll close by thanking again the group that I get to work with every day at MIT, all of our generous funders that have made our work possible over the years. I'll highlight in particular the NSF Center for Computer Assisted Synthesis, which is a phase 2 CCI that has enabled a lot of these conversations at the intersection of computer science and chemistry that we benefit a great deal from being part of. And let me again just thank ACS and the Cabley Foundation for this opportunity and for all of you for your time. And I would be glad to take any questions. Thank you.
53:31 Please join me in welcoming back to the stage 2026 ACS President Riabberto Hernandez. Thank you, Connor. Amazing. he has graciously offered to take some questions. There are mics at the at least to my right and to my left. So please gather in front. We will not we probably be able to take all the questions if there are too many but if there are only a few we'll definitely get to you. And I see one here on my left so please ask your question.
54:11 >> Great talk. So >> and introduce yourself by the way. >> Okay. I'm Pavl. So >> thank you Pavl. >> Okay. >> There may be 10 other Pavl's in the room. >> Okay. Pavvel rehack. >> Thank you. >> Okay. So, so recently I've been using AI as a computational chemist to set up systems and to work more efficiently. But it but for setting up certain quantum calculations or MD simulations, where do you see AI being useful or do you see do you see other applications?
54:54 So I think that that application of using AI to remove a lot of the tedium of setting up jobs whether they're sort of DFT or MD simulations is is a very effective one. One of the things that we've been exploring a lot is using AI to serve as guard rails. Right? So there's a there's a nice article in CNN several years ago about how DFT was becoming more commoditized and ran the risk of becoming sort of a push button piece of software. But that has its dangers because you can have sort of the wrong choices for how you set up your system and have the wrong then interpretation thereof. So I do think actually using AI to help make the right decisions for setting up calculations and the right interpretation of them is a very good application of these methods.
55:36 So we found quite a lot of success in these sorts of workflows. we have a lot of success with sort of orchestration of multiple pipelines if you wanted to piece together different sort of pre-processing strategies and other filters. and so I think what you described is really a tremendous use case. So I'm very glad to hear that you're already working on that. >> Okay. Thank you, Pavle. and Connor, I'll take the question from Chris Maduro here on the microphone to my right.
56:02 >> Hi, Connor. Thank you for such an educational lecture. It was very, very interesting. I noticed that you cleverly diplomatically avoided mentioning any of the commercial AI tools. So just curious, are you guys using those? Are you using your own AI tools? How do you see the commercial products interacting with the research you're doing? >> So we do use plenty of commercial tools. I buy Ever in the group subscriptions to to claude for example. I find for orchestration many of the tools work similarly well which is to say they all work quite well. we do also use a lot of open source open weight models and that's actually something that I find preferable when we can get away with it because that leads to better reproducibility. You don't run the risk of publishing a workflow or a paper where six months later that version is deprecated by a company and you can no longer actually reproduce what it is you were working on. So, we do prefer working with openw weight models when possible. But for for day-to-day tasks, just helping our research, we absolutely will use the the commercial frontier lab solutions.
57:12 >> In the interest of time, we will hold the the questions to the last two that are at the mic here. So, please Yeah, you win and say your name. >> hi, I'm Annie Ait from Abby, actually. and I just wanted to know what it looks like to train one of these large commercially available AI models on a specific data set for a specific use case like from end to end. What does that look like for you?
57:39 >> Yeah. So, the the way that we a lot of our models that we actually train, you know, fully from scratch are not they don't look quite like the large language models that sort of frontier companies are developing and training. So for us a lot of our models are small enough that we will use you know a handful of GPUs for days or a few weeks to train them from scratch. Many of these are geometric deep learning models working on graphs and other structured objects.
58:03 But when we take openweight models and we try to apply them in some other workflows. The way that we do that sort of post-training can look different each time. It might look like giving a pre-trained model access to smaller chemistry specific data sets. there's sort of ever evolving set of techniques that are used to to do that. But I will say for basically I think everything that I actually showed from from the group here those tend to be trained a little bit more conventionally than existing large language models. But we can chat offline about that too.
58:35 >> Thank you. >> Hi. I'm Maria. I'm a PhD student at the University of Chicago and I was wondering about the synthetic scaffolds that you're predicting. How well does your model take into account biological stability as well as synthetic stability of these new scaffolds for example under acidic conditions or heat things of the sort and how well these models are able to predict this at the moment.
59:11 >> So the answer is not as well as we would so we we're not right now accounting for metabolic stability. we're not thinking explicitly about stability under different sort of aquous or acidic conditions. We do end up getting a lot of acetals coming out of that pipeline and many of those would be you know prone to hydrarolysis and so I do think it is something that's important to layer in especially with the goal of moving into medcam applications but that's sort of a a whole other can of worms predicting metabolic stability that we're also interested in. but for now this is sort of pure thermodynamic stability in gas phase that we've been using as a filter. Do you think that's something that can happen in the near future?
59:51 >> I do think it's tractable. There there are interesting data sets of you know oxidation by you know cytochrome P450s other sorts of primary metabolism enzymes. So I do think it is something that makes sense to integrate in and a lot of the properties may depend quite a lot on the substituents that we use to decorate the scaffold. So it could be the case that instead of applying those filters to one scaffold, we use that as a filter to understand how we'd like to decorate it before we bring it into the lab.
60:16 >> Amazing. Thank you so much. >> So Connor, I do have one question. Well, I have two but first question is you mentioned as you were doing your generative prediction that you use a sort of bottom up or theoretical chemistry techniques computational techniques to do some of the pruning on the back end that made you discount a few of the structures. Have you to what extent could we imagine these bottom up with techniques coming in in the front in the prediction itself not just AI but AI integrated is that a direction that you see there is potential >> absolutely so I think it's it's hard to know exactly where the right marriage of the sort of physics base and AI methods should take place it's easiest to use them as filters I will say it's most convenient but I do think that there probably ways to more meaningfully bring them together and and ensure that Those kinds of constraints and considerations are really baked in from from the very beginning.
61:12 >> Cool. So other Nobel Prize winners have been on the stage have been asked this important question. So I asked you this important question to conclude. What is your favorite ice cream and what is your AI's favorite ice cream? Favorite ice cream would be there's a black raspberry chocolate chip that I grew up with, which is a good one. And I'll have to go back to the lab and ask to find out the second answer. All right, let's thank Connor one more time.
61:45 And with that, this concludes today's Cavly lecture, but come back tomorrow. We'll have a second one.
Summary
- AI is accelerating advancements in chemistry, enhancing research and discovery processes.
- The American Chemical Society has played a pivotal role in scientific innovation for 150 years.
- Kohley’s research focuses on integrating AI with chemistry for applications in synthetic, medicinal, and analytical chemistry.
- Historical context is provided, tracing AI's roots in chemistry back to the 1960s with programs like Dendril and Lassa.
- Kohley discusses the importance of data-driven synthesis planning and the challenges posed by increasing molecular complexity.
- The lecture highlights the need for collaboration between computational methods and human expertise in chemistry.
- Future AI applications in chemistry must account for biological and synthetic stability to enhance drug discovery.
- The integration of physics-based methods with AI can improve the predictive capabilities of generative models in chemistry.
Questions Answered
What is the significance of the American Chemical Society's 150th anniversary?
The ACS's 150th anniversary highlights a century and a half of scientific discovery and innovation, emphasizing the importance of chemistry in improving lives globally. The event serves as a platform to explore emerging frontiers in chemistry and acknowledges the contributions of chemists over generations.
How has AI been applied in chemistry since the 1940s?
AI applications in chemistry began in the 1940s with projects like Dendril, which assisted in analyzing mass spectra. The 1960s saw further developments with programs like Lassa and foundational work on quantitative structure-activity relationships, establishing a basis for AI's role in chemical research.
What challenges exist in data-driven synthesis planning for complex chemical reactions?
Data-driven synthesis planning has become a vibrant field, but it faces challenges, particularly with complex target molecules. As the complexity of the input increases, the performance of synthesis planning tools tends to decline, impacting drug discovery and the synthesis of natural products.
How are predictive models evolving in medicinal chemistry?
Recent advancements in predictive modeling allow for the mechanistic prediction of reactions with high accuracy. These models build on historical frameworks, such as quantitative structure-activity relationships, to enhance our understanding of structure-property relationships in drug design.
What role do computational methods play in discovering new drug scaffolds?
Computational methods are crucial for identifying new chemical scaffolds that may have been overlooked. These methods facilitate the exploration of chemical space, leading to the synthesis of novel compounds with potential medicinal utility.