Episode 139 · 17 August 2024 · 00:35:26
There are no big data battery papers
Sam Cooper on AI and Big Data in Battery Research
Imperial's Sam Cooper on why battery machine learning is data limited rather than compute limited, what Polaron learns straight from microstructure images, and why he would throw away 99% of battery papers.
Read the article: There are no big data battery papers

Sam Cooper
Senior Lecturer, Imperial College London
Cooper is also chief scientist at Polaron, a spin-out applying machine learning to the relationship between manufacturing parameters, microstructure and performance. He co-created Imperial's Mathematics for Machine Learning course on Coursera.
Recorded in London
What this episode covers
Cooper opens with a claim that sets up everything after it: there are no big data battery papers, only medium data ones. Academics cannot get at industrial data because it is a genuine trade secret, which limits the kinds of models anyone outside a manufacturer can build. He says a few corners of industry are starting to share, and calls it an exciting breakthrough, but the shortage still shapes what the field can learn. His own position is blunt. He has more than enough compute on the cluster at Imperial. What he needs is data, and he uses the episode to ask manufacturers for it.
The part he cares most about is microstructure, the middle of a stack that runs from atomic arrangements up through cell design, pack design and fleet charging behaviour. Making an electrode means mixing material in a bucket, spreading it on a foil, drying off the solvents and pressing it down. Simulating that from physics is slow, expensive and, in his words, currently just not very accurate: the state of the art amounts to dropping imaginary spheres in a box, rumbling them around and pressing them down. Polaron skips the physics and learns the relationships directly from images of microstructures produced under known manufacturing conditions.
The payoff he describes is engineering rather than chemistry. A factory might have on the order of a thousand interacting parameters, and tuning them properly could yield something like 20% more accessible energy at a given discharge rate without touching the chemistry or the machines. He points to CATL and BYD as the clearest examples of that approach paying off, with every knob squeezed as tight as it will go rather than a radical new chemistry invented from scratch, to the point where LFP packs are energy dense enough for a perfectly good city car. That, he says, was not expected ten years ago.
He is sharply critical of how battery experiments are done and reported. In an academic lab a PhD student cuts, stacks, drops electrolyte on and crimps each cell by hand, then hopes for repeatable results, usually without time for repeats. Robotic labs are starting to fix this, including one coming online at Imperial in Magda Titirici's group, an automated cell assembly line inside a glove box. His larger point is that methods should be code submitted to a robot rather than prose in a paper, so the same experiment in another country gives the same answer. As it stands he thinks 99% of battery papers could be discarded without society losing information.
On large language models he is practical rather than starry-eyed. They may not be genius oracles, but they are tireless workers, and the drudgery of designing an experimental campaign is exactly the sort of work they can absorb: reading several hundred pages of machine documentation and pulling out what matters for a given test plan. He is honest that reproducibility is unsolved when the model is stochastic and the vendor keeps changing it, and suggests reporting the exact release much as you would with a commercial physics solver. Materials discovery by brute search, meanwhile, he calls a stubborn problem, because a promising crystal still has to clear interface stability, internal stability and economics before it means anything.
The last stretch covers Polaron itself and where he thinks the hard problems are. Two of his PhD students, Isaac Squires and Steve Kench, are now CEO and CTO, funded, with an office in Shoreditch, and Cooper visits one day a week. The tool is deliberately not battery specific, because alloys, concrete, additive manufacture and pharmaceutical coatings share the same processing-microstructure-performance problem. He found industry far easier to engage as a company than as a university, where contracting, IP and the pressure to publish get in the way. On the future he is most exercised by grid storage, and repeats the line that dirt cheap batteries have to be made with dirt.
Questions from this episode
- Why is it so hard to do big data battery research in academia?
- Because the data sits inside manufacturers and stays there. Cooper says industrial battery data is a real trade secret, so academics cannot access it, and that limits the kinds of models they can develop. He describes the field as having no big data papers at all, only medium data ones, with consequences for what can be learned from them. Some corners of industry are now starting to share, which he calls an exciting breakthrough, but the constraint has not gone away. His summary of the position: he is not compute limited, he is data limited.
- What does Polaron actually do?
- It learns the relationship between manufacturing parameters, microstructure and performance directly from data instead of simulating it. Physics-based models of electrode processing are expensive, slow and not very accurate, so Polaron feeds a machine learning model images of microstructures generated under different manufacturing conditions and lets it work out the relationships. The result is an algorithm that can search that space quickly and identify the best microstructure for a given application. Cooper says the tool is not yet something you can click and buy: companies work with Polaron directly, and the near-term plan is to partner with a few firms in a few key industries.
- How much can AI improve a battery without changing the chemistry?
- Cooper's example is a factory with roughly a thousand parameters that all interact. Tune them correctly and you might get something like 20% more accessible energy at a particular discharge rate, with no new chemistry and no new machines. He sees this as what CATL and BYD have done: rather than inventing a radical chemistry and reinventing every process around it, they finessed the engineering until conventional chemistries performed remarkably well. LFP packs energy dense enough to power an ordinary city car are the result, an outcome he says was not expected ten years ago.
- Why does Cooper say most battery papers could be thrown away?
- Repeatability. In a typical academic lab, cells are assembled by hand inside a glove box: pieces cut out individually, dropped on top of each other, electrolyte added by hand, then crimped, with everyone hoping the alignment and quantities matched last time. Often they do not, and there is rarely time for repeats, which leaves researchers in an awkward position when reporting findings. Cooper accepts the point is cynical, but argues you could discard 99% of battery papers ever written without society losing information, because the results are so hard to reproduce.
- What could robotic labs change about battery research?
- They make the experiment itself reproducible. Cooper points to a robotic lab coming online in Magda Titirici's group at Imperial, essentially an automated cell assembly line inside a glove box, replacing the manual cutting, stacking and crimping done by PhD students. His broader ambition is for experimental methods to be written as code submitted to a robot, rather than described in prose, in the same way simulations and data analysis already are. Run that code here or in another country and you get the same result, because the process is tightly controlled.
- Where are large language models genuinely useful in battery science?
- In removing the drudgery, not in producing insight. Cooper's example is designing an experimental campaign, where you have to reconcile budgets, timelines, available chemicals, funder interests and the capabilities of each machine, sometimes documented in several hundred pages per instrument. A model with a long enough context window can read that documentation and extract what matters for a specific test plan. Asking a PhD student to do the same is a lot of work for a small result. He is more sceptical about using them to search chemical databases, and honest that reproducibility with stochastic, frequently updated models is unresolved.
Listen to this episodeWatch on YouTube
Part of The next chemistry
Transcript
About this transcript. Generated automatically from the recording, then corrected against a glossary of company and guest names. It has not been checked line by line. Machine transcription mis-hears technical terms, numbers and names, so treat any figure here as a prompt to check the recording rather than a quotation of record. Spotted something wrong? Tell us.
0:00Introduction
Dr Simon Engelke
0:00Welcome everyone and thank you so much for joining us for the Battery Insiders podcast today live from London. I'm really excited to have you all with me. My name is Simon Engelke and founder and chair of Battery Associates and I've got a fascinating guest with me today, Sam Cooper, who's the senior lecturer at Imperial College London and chief scientist Polaron. So, thanks for being here. You're very welcome, really nice to be here. Which is really, really good and I think I still highly appreciate it by some of the new people joining. But also you have been involved in the topic of big data and battery research. I know actually some of your earlier work, even I think during my PhD I was looking at, remember these days, this is when I was in Cambridge, it was good to see and some of the some cool tools you developed actually on looking at data. But yeah, today we want to talk a bit about like big data, battery research, generative AI, and some of the approaches you have been taking in that space. Awesome. Well, I'm delighted to be
Sam Cooper
1:01here and it's obviously a topic close to my heart. But one of the first things to say about it is that it's very difficult to do any big data battery research as an academic because there's not very much data around. And I think it's a discussion that kind of rumbles on in the community that we, on the academic side, really can't access the data from industry because of course it's secret, because these are important industrial secrets. Which means that it really limits the kind of models that we can develop. And I think there are now starting to be some exciting breakthroughs where various corners of industry are going to share their data and that's great. But it still remains a challenge that I think, I would say, there are no big data battery papers. There are only sort of medium data battery papers, which has implications for the kind of things that you can learn.
Dr Simon Engelke
1:48Fascinating. And I think also maybe one to kind of start us off also on a bit of the approach you're taking right now, right? So I think you have been in academia, then you're very successful also in
1:57Teaching machine learning at scale, and AI at every length scale
Dr Simon Engelke
1:58courses. I know you're also very successful course over course, which is I think maybe many other people might know you from. How many people are you now? So we're over 600,000 students through that
Sam Cooper
2:10course. And that's called Mathematics and Machine Learning for those who are interested. But it launched seven years ago, I think. And this was Coursera looking to build a partnership with Imperial College about education. And when we launched the course, obviously you have no idea that these things are going to be successful, much like the BatteryMBA, you don't know, you just got to go for it. And then it turns out, perhaps unsurprisingly in retrospect, that there's a big appetite for learning about machine learning. And there's a huge appetite to learn about batteries as well.
Dr Simon Engelke
2:37Now you definitely hit two of the really trendy topics there, which is really well done. And yeah, maybe we can talk a bit about, you know, yeah, like how maybe AI is kind of shaping,
Sam Cooper
2:46I think, the battery field. Yeah, absolutely. So batteries are a very broad topic, right? You've got everything from understanding the arrangement of a few atoms that can really influence the performance of your battery, all the way up through microstructure, cell design, pack design, and then the understanding of the dynamics of your fleet and charging network. And to some extent, even including the grid itself, because you want to understand how to effectively, cost effectively charge your vehicles. So there's so much in there, and it requires the combination of skills from so many different disciplines at the same time, chemists, physicists, engineers, all need to play together. And just like everything else in science and engineering, AI is starting to sort of creep into all these different spaces, but in quite distinct ways. So for the atoms people, some of the computations that you have to do from first principles are very expensive. And people have found ways to make efficient versions of those. So you can call them surrogate models that represent some of this complex physics very cheaply. And then you can do much bigger simulations, and you can understand bigger crystal arrangements and these kind of things.
3:53Microstructure, and learning the relationships from data
Sam Cooper
3:54Then at the fleet level, there is seemingly what feels like about 100 papers a week, or sorry, at the pack level, about 100 papers a week to try and understand the relationship between what you've asked your battery to do, so the charge discharge cycle, and how much charge you've got left, or how efficient you are, or how much remaining useful life of your cell you've got. And those are radically different things. But the interesting thing for machine learning is often it requires the same core skills. So behind that is going to be someone who probably knows how to code in Python, understands a little bit about how to send their work off to clusters to compute, and understands a bit about cleaning their data and making it tell you the story that you need to hear. Right in the middle of that is microstructure. So that's the topic that I'm particularly passionate about. And there's a big problem there, which can be characterized by the idea that the relationship between all of the processing parameters in your factory. So when you make a battery, you need to mix together some stuff in a bucket, spread it on a foil, let the solvents dry off, and then crush it down to make it a bit denser. And to simulate those processes in a physics-based model is insanely expensive and slow, and also currently just not very accurate. So the state of the art for this is kind of dropping a bunch of imaginary spheres in a box and rumbling them around a bit and pressing them down a bit and hoping that it looks a little bit like the real thing. And of course it doesn't look anything like the real thing. What we're trying to do at Polaron is to, instead of doing all that complicated expensive physics, just learn directly the relationships from data. So you've got microstructure that's been generated from a particular set of manufacturing parameters, and you've maybe got an image of that microstructure. And then you just say to your machine learning algorithm, here's lots of different images from lots of different manufacturing conditions. Learn those relationships. And in a sense like, I don't care how you learn them, just learn them. And now I've got an algorithm that will be able to really quickly search around that
5:51A thousand knobs in the factory
Sam Cooper
5:52space and understand what's the best microstructure for my particular application. And that's very exciting and similar to all the other length scales, where you say, I actually don't know the physics to model all of this stuff, so I'm just going to learn it from a big data set that must in some sense implicitly represent the physics inside it. Fascinating. What would be the best outcome from that?
Dr Simon Engelke
6:14Right? Like, you know, how can this improve the battery if someone is curious? Right. So,
Sam Cooper
6:19I think at the micro scale, again, where you want to know, you've got a factory, and inside that factory, you've probably got a thousand, for the sake of argument, a thousand different parameters that you need to optimize. And the relationships are really complicated. They all interact with each other. And so, if you can tune them just right, you might be able to get a battery with, you know, 20% more accessible energy at a particular discharge rate, for example. That's a huge win, because you haven't changed anything about your chemistry. You haven't had to change your machines. You haven't had to do anything except tweak all the little knobs until they're in just the right setting to get your best performance. So, I think that's something that we've seen particularly coming out of CATL and BYD, Chinese manufacturers, where the engineering side, rather than inventing a radical new chemistry, and having to reinvent all your processes, instead, if you could just finesse the engineering, get every little bit squeezed as tight as it'll go, then you can actually get amazing performance from relatively conventional chemistries. And that's one of the things we're trying to do at Polaron.
Dr Simon Engelke
7:22That's fascinating, because I just read an interesting press announcement actually recently about CATL, where they actually spoke about that they have a really new innovative way about grading, which makes them much more competitive now, and also more on NFP, energy dense, and things like that. So, grading to compressing. So, I think, as you say, even now in the announcements, you kind of see this at least, you know, whatever they're going to be able to share, but I mean, at least highlighting also this kind of, I guess, microstructure engineering. Absolutely. There's a really interesting
7:48Why NASA grades cells rather than building perfect ones
Sam Cooper
7:49story about grading from a guy called Eric Darcy, who was head of batteries at NASA's Jet Propulsion Laboratory. And he came over to give a talk, and he's done some amazing work where you put deliberate errors into your battery and see how they fail. And he said that NASA spent a long time trying to make perfect batteries so they could put them in space. Obviously, I don't mean perfect, but I mean the really, really good ones. Can they get all the machines just right with the power of NASA behind you? It turns out they couldn't do it, because you have to operate at huge scale to make good batteries, and there's just no way around it. So, in the end, the solution is to buy a very large number of cells from whoever's the biggest supplier at the time and grade them. And then you just take the best, you know, 1% of those cells, and those are the ones that go in your spaceship. And it seems that there's just no way around that, even still to this day. You really can't make a small number of excellent batteries, no matter how much money you throw at it. Fascinating. Great. Maybe also then some of the, you know,
Dr Simon Engelke
8:45key other challenges we kind of see being overcome, you know, with machine learning, AI in the battery
Sam Cooper
8:50world, some of the other things you could approach. Yeah, so there's a lot of work being done in the space of materials discovery. So, you know, the periodic table is what it is, and the crust of the earth is what it is. So we've got to work with the materials that are actually available to us. And that means some materials like cobalt, which is a very commonly found element in batteries, is very uncommon in the earth's crust and only available from very specific places. One of them is the the Democratic Republic of Congo. And that means that we really should be trying to move away from this chemical, this element, not just because it's rare in the earth's crust, but because the conditions with which it's being extracted from the ground are generally very unethical. There's an attempt to try and find the next generation of materials based on machine learning driven, massive search through the space of all possible combinations of materials in the in the broadest sense. But that's turning out to be
9:48Materials discovery as an intersectional problem
Sam Cooper
9:49very stubborn, a very stubborn problem, because a battery is not just about which material, which crystal will incorporate the most lithium or incorporate lithium at the best voltage. It's such a intersectional problem. It needs to do those things, but it also needs to have very stable interfaces with regard to the electrolyte that's touching it. It also needs to be economically viable. It also needs to be stable internally for a long period of time. Some of these materials, their crystal structures degrade. So there are so many factors that interact with your design decision that mean that just because you've found an amazing crystal structure in a sort of perfect idealized condition, it might turn out to never be a useful battery, because it really has to go through so many additional thresholds before it can become technologically relevant. Not to mention the development time, which is going from a material that you found in a computer to something that you're going to find in your car. As it stands, that's sort of a decades type timescale. And hopefully that's coming down, but it still is a long journey.
Dr Simon Engelke
10:55One thing, I listened to another podcast more recently about, you know, like essentially developments of AI and kind of neural nets not being that modern, like not novel. A lot of these algorithms actually existed there for a long time. What really enabled them now was, you know, Nvidia and all of these, you know, GPOs and the costs coming down. So now you can do this really intense computations, which were not economical before. So really, sometimes technology, you know, I mean, the knowledge was there in the early days, but just the computational work wasn't there.
Sam Cooper
11:24Are there similar topics maybe in... I would say that there are lots of, you know, distinct kingdoms of machine learning where the ratio of compute limitation to data limitation is different. And my gut feeling is in the battery world, we are not compute limited. There's plenty of compute, we're data limited. And that may not be true if you are inside CATL and you've been testing batteries
11:47Data limited, not compute limited
Sam Cooper
11:47for a very long time. But from an academic perspective, currently, I've got more than enough computers with my cluster at Imperial College. I just need more data. So that is, you know, that is a message to all battery manufacturers out there. If you're willing to share data, we'd love to have it. Yeah, but I think, you know, we have clearly benefited from the advancements in the hardware behind machine learning. And as you say, many of the algorithms are very old, they've been around for decades. I think what's interesting is that even the most sort of state-of-the-art thing that's just emerged, which you could say was transformers or diffusion or something like this, it's still essentially just a neural network with a very particular structure and a very particular training strategy. But the fundamental pieces are just multiplying and adding. And I think that's really exciting, or it is from my perspective, where electrochemistry is really hard. And if you want to be good at electrochemistry, it takes many, many years. If you want to do anything with electrochemistry just to get started and have the confidence to write down some equations and use them, that can take a long time as well. But with machine learning, you can get up and running, you know, in a week you can train a model that does a little thing very, very quickly. Of course, to do it well takes a lot of practice, but you can get started faster. And I think that's actually really exciting and it's allowed a lot of people to get involved and have a go across the whole discipline. And I think that's probably going to be really good for us.
Dr Simon Engelke
13:10Yeah. Another, I mean, I know battery data is a topic we're both quite passionate about. I know we also connected with Battery Dev, which was our hackathon. We were running about data and it still was the battery cycle, right? We have this ambition to kind of also open up more data. So we are, one thing I definitely can appreciate now, hardware is hard. So really for anybody waiting out there, you know, for the battery cycle, we still work on it. So it's progressing, there's exciting
Sam Cooper
13:33developments, but... Excited to hear. Can I ask you, what's the like limiting step? What's the current
Dr Simon Engelke
13:37tricky bit? I think, I mean, many, like always I feel, a lot of learnings and everything. I think,
13:44Hardware iteration and the arrival of robotic labs
Dr Simon Engelke
13:44you know, now we're in the certification stage and then also with the software integration. And I think, I mean, in general, right, like always this balance between like, there's always something else you would love to integrate. Yeah. And you have to just kind of say, no, that's going to be the next version. But, and then just like, you know, revision cycles, I think hardware is still also the stage we're in just revision, you know, producing hardware. It's just the, for software, it's just like so much faster to kind of iterate this hardware. Also with us being a distributed team in different countries, shipping out equipment, having them, you know, equipment manufactured, updated, modified. Right. So this is just all things, which I think we're getting fast every time. So it's just, that's gets me personally excited because I can see the potential of quick iterations. And we have been quick and iterating. But yeah, this is just things which are, as someone who's spent more time on software before,
Sam Cooper
14:37can definitely appreciate that. That's, I think this software hardware, uh, twin relationship is also coming along in lots of parts of experimental chemistry. So for a long time, the people working in their software were having a great time, as you say, because you can really specify things very exactly. And you can version control your code and you can really be confident when you've made some progress in a particular way. In the experimental world, you can do the same experiment twice on two different days and get a totally different number. And that's, that's a real problem. Uh, what's been exciting over the last couple of years in particular is to see the emergence of robotic labs. Uh, and so, uh, Professor Magda Titirici at Imperial, she's just about to get a robotic lab on, come online, which is basically an automated cell assembly line inside a glove box. So when you make your cell, you've got to have a controlled gas environment around it. Otherwise, you'll, uh, rot the surface of the materials. And the usual way to make a cell in an academic context is to have
15:39Methods as code instead of prose
Sam Cooper
15:40a PhD student or a postdoc, uh, and they will manually drop each piece on top of the other, having cut them out hand by hand, and then hope that they will align well and hope that they've dropped the same amount of the electrolyte on top and then crimp all these things together and hope that everything has come together close enough that you can get repeatable results. And the frustrating reality is you often don't get repeatable results, but you often don't have the time to do more repeats. And so you're left in a very awkward position in terms of how you report your findings with a robot. Hopefully you can automate a lot of that, uh, and speed up the process. And I think then you can imagine that currently methods in academic papers are reported in prose. So you write a sentence, we mix these things using this thing and that's fine. But what would be much better is if the methods of your experiment were also code, just like your simulations and data analysis are code, your experiment should also be code that is submitted to a robot. And if you do it here, or if you do it in a different country, you get exactly the same results because it's being done by a very highly controlled process. And I think that that is happening and coming online and will have a huge impact on battery research. This hopefully doesn't sound too cynical. I mean, it is quite cynical, but I think you could probably throw away 99% of battery papers that have ever been written, and we wouldn't have lost any information as a society because it's just very hard to get repeated results. Um, and that's, yeah, we're working on it. Yeah. Totally. And I, I was just in California
Dr Simon Engelke
17:10recently in Silicon Valley, especially there's quite a few startups now thinking about these robotic topics and should they integrate it, but of course costs and it hasn't really been done so much at maybe a larger scale, some of these topics or maybe industry is much faster to move. So I think everyone is holding off a bit. And I'm sure in academia, then you might have a grad student who's spending a lot of time integrating it off. So you still might need them there. Um, I want to actually come to startups in a moment, but one thing before, just because about repeating, I think there's one topic also
17:37Reproducibility when the model underneath keeps changing
Dr Simon Engelke
17:38with AI, right? Like the black box and it keeps iterating. So maybe today you develop your protocol and that's, that's everything I put, let's say really simple and chat GPT or that's what I put in, but then actually the underlying tool changes without my control. So you might can have a great protocol. This is all my prompts or whatever, but then actually the platform itself, the algorithm were iterated on. Yeah. How do we do this like in, you know, in the topic, but if you want to use AI and be using third-party services, so they keep developing, how can you actually create repeatability in that?
Sam Cooper
18:09I mean, particularly in the context of large language models, that's a very difficult problem. There are of course some open source options where you know what's going on inside the machine, but inevitably they are somewhat behind the closed source ones where you're using API calls. I don't know is the honest answer to that. I guess much like you would with, if you used a commercial physics solver like COMSOL, ideally you would state which exact release of the software you were using as you would with the language models. The thing with the language models is that they're also stochastic, right? So they're not giving you the same answer every time by design. And so how do you report results in that context? I think it's one of the things that we are hoping to use large language models for in a proposal that we've just submitted to the European Union, is around pulling the boring and drudgery components of science out of the mix. So currently I would say that one of the major obstacles to doing high quality science is how boring it is to be very rigorous about each of the steps. So if you imagine I want to design a experimental campaign and someone's come to me and said you've got two years and you've got you know two million pounds and I need you to answer this question, there will be of course better and
19:33Language models as tireless workers
Sam Cooper
19:34worse campaign designs in order to answer that question. And in order to design that campaign you need to integrate all of the resources that you have, all of the constraints that you have, things like which chemicals are available to you and which ones the maybe the funder is interested in, but also things like what capabilities each of your machines has. And each of your machines comes with a very detailed book that describes all of these capabilities, maybe several hundred pages of notes on this particular machine. And in order to synthesize all of that information into a plan, currently we rely on a person with experience, usually you know the PI of the project, to say we should probably do this and probably do this, and we've done this before and it worked and we should do this. And that seems to work okay. But I think what's really exciting is large language models, although they may not be genius oracles yet, what they are is tireless workers. So you can say to it, if it's got a long enough context window, read this book, what's the important information I need from this guidebook on how to use a cycler if I want to design a campaign that tests cells of this type? And it will extract that information for you. And you could ask a PhD student to do that, but that's very unkind because it's a huge amount of work for a very small thing. And so I think, yeah, there's already some really exciting applications that are somewhat resilient to the stochastic nature of the LLMs and that's quite cool. Of course people are using them for some other weird and wonderful things, things like searching chemical databases, but I think it's yet to be seen whether those results are going to be important for the community or just a kind of a curiosity.
Dr Simon Engelke
21:17Yeah, absolutely. And then if you talk about curiosity and then also this, you know, still with Imperial College, but also chief scientist of Polaron. So what's kind of your thinking maybe to also do something more startup or like, you know, industrial related?
21:30From the research group to Polaron
Sam Cooper
21:32Yeah. So I have been an academic for, this is my eighth year. And it's certainly, I've been aware, by being an Imperial, there's a lot of spin outs from Imperial. I've been aware of this being a thing that would be beneficial both for my career and just for my interest in the topic, to try and bring the impact, to try and generate impact in the world, right? It's fun to publish papers and I love to do it, but how can I turn all the understanding that we've generated into the group into something a bit more meaningful? And so over the last year, two of my fantastic PhD students, Isaac Squires and Steve Kensch, they are now the CEO and CTO of Polaron respectively. And they've raised some funds and won some prizes. And they've got an office in Shoreditch and they're going for it. And so I come and see them one day a week to go and find out how things are getting on and also to try and contribute in various ways. But it's incredibly exciting to see what you've been working on for seven years turn into a tool and a product that industry can use to make a difference. And what we've been focusing on as a group, as you indicated at the start, is AI tools for understanding the relationship between microstructure and performance and microstructure and manufacturing. And what we realized quite early on with Polaron is that although we come from the battery space and that's where we're very comfortable, there are very, very many applications that need to understand the relationship between processing parameters, microstructure and performance. And so this could be concretes, it could be alloys, it could be additive manufacture. All of these things really, really have a very strong, very significant relationship between manufacturing parameters, microstructure and performance that is still hard to understand. In fact, batteries might be one of the hardest applications for
23:27Alloys, concrete and coatings beyond batteries
Sam Cooper
23:28this because they are a pain in so many ways. And so Polaron is focused on just generating a tool that can be applicable across industry more broadly rather than just batteries.
Dr Simon Engelke
23:39So what would happen now if someone wants to, can people already use the tool?
Sam Cooper
23:42Yeah. Yes and no. So they would have to do it through Polaron. You can't just click and buy just yet. But the plan, roughly speaking, is to work with a few particular companies in a few key industries and try and work with them to make sure the tool meets their needs.
Dr Simon Engelke
23:57Yeah. So it could be like a use case, maybe I've seen anything you can share without, I guess, customers, but like more what will people want to do?
Sam Cooper
24:04Well, the ones I obviously can't be too specific, but the ones I mentioned a second ago. So people like alloys manufacturers who really care about their relationship between their crystal morphology and the mechanical performance of their system. Or another interesting one is coatings. So often for certain applications around medical, so medical applications, you can see I'm struggling. So just for context, as somebody who spent eight years being an academic, I've always been in the habit of saying everything that I know. And I've now had to learn this skill of not saying everything. Some of it has to remain private or secret. Yeah. So coatings are sort of a polymeric substance on top of a pill, for example. And if you want it to be an effective medication, you might need that coating to have very particular properties. And so that's another area where this kind of microstructural optimization might be really handy. And of course, batteries as well. So we're looking into all these spaces.
Dr Simon Engelke
25:05Fascinating. So you're kind of venturing out a bit more maybe then, because I guess in your academic,
Sam Cooper
25:09you're very focused on batteries, right? Yeah, yeah, yeah. Absolutely. For the past seven years, batteries have definitely been the major focus. Although it's always been a microstructure in general type toolbox, because I started off actually working on fuel cells. So solid oxide fuel cells were
25:25Fuel cells, and lessons from spinning out
Sam Cooper
25:26much more fashionable when I started my PhD. And now they're going down a little bit, and batteries have obviously come a long way. But a lot of the problems are the same, where you're trying to co-optimize various things, including, you know, transport of things through your poor network, and reaction kinetics on the surfaces of things. And the optimization of all these things at the same
Dr Simon Engelke
25:46time is just really tricky. Fascinating. And then you already mentioned now you have to be a bit more careful with what you say. Any other learnings kind of, you know, through like, you know, being involved in
Sam Cooper
25:55the startup. Right. And you have seen. Yeah, I think having conversations very early on, from an academic perspective, having conversations early on with your institution's commercialization team, it's maybe it sounds almost too obvious to say, but I think it's beneficial because you, it's just such a complicated and distinct space from academia, that there will be some relatively obvious things that you hadn't thought of. But then the thing that we found to be such a joy is how incredibly helpful and welcoming the startup, well certainly the battery startup community is. So there are various companies about energy and breathe batteries, both of which are spin outs from Imperial, who have just been so generous with their time to say, oh, don't do that, do this. And that has really saved us from falling in a lot of holes. And it's also just, it's just a real joy. And you might have thought that academia, because the stakes are in many ways quite low, will be much more collegiate. And you might have thought that startups would be somewhat more adversarial, because there's only so much money around. But I would say it's actually quite the reverse, and that we've been found an incredibly welcoming environment in the startup community, whereas academia can be a bit barbed sometimes.
Dr Simon Engelke
27:08I'm brilliant to hear. I agree on that. So I think, and I guess maybe then one is just the next, right? Because do you feel now it's easier to work with, because you mentioned data, right? Like how you would wish to get more data. So you did that even ask our audience. So I know quite a few executives from
27:23Why industry engages more easily with a startup
Dr Simon Engelke
27:23Automotors are listening to this podcast. So if you feel like it, you're welcome to also get back to Sam and myself. You are so welcome to do that. Do you feel now it's easier to interact with industry? Absolutely. Being a startup versus being an academia? Yeah, absolutely. And one of the things is,
Sam Cooper
27:40understandably perhaps, that if you want to do a project with Imperial College, then you need to go through a tricky contracting process and IP arrangements where the college might wish to have some stake in what comes out of it. Not always, but often. Whereas as a startup, you would very often provide a service for a company. And if you're not, the arrangement is still relatively straightforward to do because you're not trying to publish in the public domain what comes out. That's the important distinction. Whereas universities, that's one of their key performance metrics. So yeah, we've also found industry astonishingly keen to engage. And we've been so delighted by how much inbound traffic there's been. Of people saying, oh, we heard about this on the grapevine. We'd love to hear more. Amazing. We were
Dr Simon Engelke
28:24not expecting that. So it's really been a delight. I can attest to that. I agree on that as well. And I think it's a pity, right? Because I think maybe there's also ways for academia then to find maybe other effective mechanism to this industry. On the other end, of course, there's a lot of, especially here, remember from my PhD days, government funding, from my joint research grants, which always need an academic as well as an industrial, maybe also an SME partner. So I guess then you kind of come back in
Sam Cooper
28:47this in these funny situations. Yeah. And there are, I'm sure there will be scenarios where Polaron ends up interacting with academia again. And that will be great. But it has been a relief to see just how much easier it can be to interact with industry once you've started a company. So if you're an academic and you are frustrated by how difficult it is to interact with industry, just start up your own company, is my advice. Much, much easier and much more likely to have immediate impact, which has been
Dr Simon Engelke
29:15really exciting. Brilliant. Maybe I guess the final point is kind of look a bit into the future,
29:20The gap in grid storage
Dr Simon Engelke
29:20some of your hopes in this space of machine learning, AI, structural understanding, etc. Where do you think we might be going? Because I love these utopian views of where we might end up,
Sam Cooper
29:30but you can go as wild as you want. Wow. Well, there are some chemistries that have been worked on for a while. For grid storage, I think is a particularly exciting thing. When you see the Tesla Master Plan V3, so they make the effort of publishing a document that talks about how are you going to electrify the world based on the best knowledge of their in-house guys. And one of the things that's shocking is how much grid storage is going to be needed and how big the gap is between the technologies we currently have and the technologies that we're going to need. And so, you know, there's still discussion of things like cavern hydrogen storage as a major source of stored energy. And these technologies are possible but like unproven at the scale that we need, in my view. And so it's going to be fascinating to see which very large scale, very, very cheap technology is going to fill that gap. And there's, you know, there has been various projects like redox flow batteries, which have been in people's minds for decades, but have never quite made it because electrochemistry is very hard and degradation happens and all kinds of expensive components are required. Similarly, there's a guy at MIT, Professor Don Sadoway, who does an amazing job of selling his liquid metal batteries, but they turn out to again be harder than they were expected to be. But the sales pitch that he gives, I think, taps into the core thing, which is in the end, if you want dirt cheap batteries, you've got to make them with dirt. And I think, you know, as I said earlier on, there's no way that we're going to be powering the grid off batteries that contain things like cobble, because there's just not enough of it. So what technology it's going to be
31:16What LFP made possible, and the limits for aviation
Sam Cooper
31:17that ends up storing gigawatt hours is, I think, yet to be seen. And I don't think that can be just finessed with engineering in the way that particularly BYD and CATL have finessed the lithium-ion phosphate chemistry such that the energy density of their packs is high enough to fuel a perfectly good car. Right? That has been amazing to watch and I think was not expected 10 years ago. It was thought that you needed to have these higher energy materials. You don't. You only need those high energy materials if you want to drive silly cars like a cyber truck. But if you want to drive a car in a city at a reasonable price, you can do it with lithium, iron and phosphorus. And that's absolutely astonishing. So I'm really excited to see that. In terms of powering things like aeroplanes, I simply don't know. And it's not at all clear that you will ever be able to power large long-haul planes with just electrochemistry. It might end up being that you need synthetic fuels to do these things or some other kind of liquid fuel. Maybe you've got a different opinion, but I don't think we've yet got a battery in the build-up that could really directly replace the amount of energy that's needed to fly across the world with 400 people in your cabin. I mean, I just saw an announcement today from
Dr Simon Engelke
32:34CATL which announced that they have a prototype flying now tons of weight of airplane. Okay. On tons scale. So...
Sam Cooper
32:39That's exciting. Yeah, yeah. I think one of the things that I think is often quite cheeky in battery, in the battery space, I suppose this is true in any space, but I get exposed to it in batteries. Sometimes I will be approached by companies to come and review another, someone else's technology for them. So I get special access to read all their documents. And that's always really exciting. And one of the things that comes up quite a lot is people reporting various metrics separately. So what you... what I want is a plane that can fly this fast, this far, and for this long. But what the
33:13Cherry-picked metrics and reporting in good faith
Sam Cooper
33:14company might cheekily report is a plane that can fly this fast for one minute. It can fly this far, but you have to do it at 10 miles an hour. And it will last this long if you never use it. And that is, of course, not at all what we need. So you have to be very cautious when you see people reporting that they've made past these milestones because you don't know what other things they've compromised in order to do that. And yeah, there have been... you know, it's a bubbly space, just like machine learning. So you see companies suddenly explode into prominence and be incredibly important, and then sort of gradually fade away again, as it turns out that their technology was amazing in one metric and didn't deliver in any of the others. So yeah, it makes it very exciting. But I hope that everyone's doing it in good faith. Because I think every dollar that gets invested into a dud is a dollar that could have been invested into something more meaningful. Yeah, absolutely. And in the past episodes, we spoke
Dr Simon Engelke
34:08about A samples, B samples, C samples, and all the challenges, as you just mentioned, optimizing and then getting closer to the real product. Yeah. But yeah, I mean, I will definitely keep following you as Polaron and see the developments there as well, how this is developing. And definitely, we're going to stay in touch. I mean, you know, again, you know the Battery MBA community. I'm a big fan of what you're
Sam Cooper
34:27doing as well. Yeah, it's been great. I mean, one of the nice things about the MBA is because I gave one of your early lectures, I've had so many people from across the battery sector, both academics and from industry, reach out to me and say, ah, what's your lecture? Can we have a chat? That gives me an excuse to say, oh, I've started a company, you should have a look. And that's all good. I think that's very healthy for everyone to have all these connections. So thank you for that. Absolutely. Pleasure. And yeah,
Dr Simon Engelke
34:48again, also big thanks to you, Sam, for sharing insights to get today with our audience. And also want to thank you as the Battery Insiders audience to listen in today. You know, if you're interested in hearing more conversations like these, you can either follow us on YouTube, on the Battery Insiders as part of the Battery Associates channel, as well as also on Spotify, Apple Podcasts, or anywhere else you listen to your podcasts. And you can also go on batteryinsiders.com to get notified about future episodes as well. And with this, thank you so much, everyone, for listening. And thank you again, Sam, for being with us today. Thank you so much, Sam. And good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good good