Episode 126: Why AI Excels at Some Things and Sucks at Others

Download MP3

[00:00:00] Dr Genevieve Hayes: Hello, and welcome back to Value-Driven Data Science, where data professionals become strategic experts. I'm Dr. Genevieve Hayes, and I'm here again with Lauren Pearl, a business strategist, three-time founder, and CFO advisor, who also teaches financial modeling to founders at NYU Stern's Berkeley Center for Entrepreneurship.
[00:00:25] Last week, I interviewed Lauren about why building something valuable with AI doesn't necessarily make you valuable. Today, we're going to swap roles,, with Lauren taking over the hosting duties. Welcome back, Lauren.
[00:00:39] Lauren Pearl: Yay, I'm so excited to swap roles. Now I get to pepper you with questions.
[00:00:43] Dr Genevieve Hayes: gosh, I'm not sure if I made the right call in doing this, but let's see how it goes
[00:00:48] Lauren Pearl: Well, it'll be fine. I'm just going to tell you all my AI problems, and then you can become the representative voice for all of AI. It should be very easy.
[00:00:56] Dr Genevieve Hayes: Okay, so A few weeks back, Lauren and I were talking, about why large language models or LLMs excel at some tasks while struggling badly with others, and what that means for anyone currently using AI in their work.
[00:01:10] We only scratched the surface in that conversation, so we decided to pick up where we left off here. So now, over to you, Lauren
[00:01:19] Lauren Pearl: Yes. Okay, so I need to start this conversation off with a story because this question comes from great pain on my part. So I do a lot of modeling and teaching people modeling, and so I've been trying to use AI to help me when there's permission to. And I have this really interesting experience with it, which is AI really sucks at Excel.
[00:01:42] So I've tried all the different platforms. I've tried doing it in a desktop app i've used Claude, I've used ChatGPT, I tried Copilot. I think the best one I've experienced was using the Claude plugin in Excel.
[00:01:56] That one seems so far the best. But when I say it sucks at Excel, it sucks in this very frustrating way, which is that it doesn't totally suck. So AI gives you the same amazing feeling, I think, as it does when you're vibe coding with it, where it just makes you feel like you're flying.
[00:02:11] You can ask it to build you a model of something, and suddenly a model that would've taken me hours and hours appears on my screen in Excel, which is really cool And it looks beautiful. It can actually talk to you about design.
[00:02:25] And then you get into it, and it is terrifying because when you actually check the formulas, they're wrong, and they're quietly wrong. And it hides a gajillion mistakes that if you used it in a work product could get you fired, I think one of the best examples of how it will get things wrong is whenever you're using it to build tests.
[00:02:46] So in Excel, you have this thing called a test where basically you will check that a string of formulas is doing its job by seeing if two values are equal to each other, two values that if added together should equal to see if they are equals. So if you ask Claude plugin for Excel to do this for you, it will create tests, but often it will use the wrong values, so it won't actually test anything.
[00:03:11] It'll just pick two values that inherently will equal each other so that the test passes. Now, this doesn't happen all the time, but quite often it does. And something I was curious about was, is this just wrong because the AI isn't good enough yet, or is there something foundational going on that will always make it struggle at problems like this?
[00:03:33] And there were two examples that made me think, you know what? Maybe it's actually the domain. Maybe it's just not the right problem. The two things were, one, I listened to, I think I was actually on LinkedIn, and I was scrolling through, and I heard an influencer give me the quote around what if AI is really just good at programming?
[00:03:53] There's apparently been some talking about this online of someone 20 years in the future, what if we all look back and discover we were trying to use AI to do all of these things, and what if the big deal really was just that it could c- So that was, like, kind of thing one of maybe we're trying it on these problem spaces that really aren't a good fit when all we should do is code with.
[00:04:11] And the second was an engineer friend shared this article with me around the difference between Shannon complexity and Kolmogorov complexity. So there was a Columbia article by a professor there called Shannon Got AI This Far, Kolmogorov Shows Us Where It Stops. This idea that there are certain problems beyond pattern matching that just AI will just never quite be.
[00:04:31] And so all of this preamble to ask you, Genevieve, will AI just always suck at Excel, and what's going on here?
[00:04:42] Dr Genevieve Hayes: Okay. Firstly, will AI always suck at Excel? I don't know, because there'll probably be a point in the future where it doesn't, so I'm not gonna commit to that one. But I can tell you some of the reasons that might be leading to the weaknesses that you've identified. So firstly, when you're working with generative AI or LLMs, what you're really working with is a machine learning model.
[00:05:05] And one of the golden rules of machine learning is a model is only as good as the data that it's trained on. Now, one thing you've gotta keep in mind when you're dealing with generative AI models, they're trained predominantly using text to produce text.
[00:05:22] There might be some Excel in that training mix, but the goal is to produce text, so they're always going to be naturally weaker when they're not dealing with text.
[00:05:33] Secondly, your problem is that LLMs are stochastic. So that basically means that if you ask the same question twice, you can get different answers. Now, most processes you wanna build using Excel spreadsheets are deterministic, so if you've got random variation, you're gonna have problems, and that might end up manifesting itself in, out by one row errors in your spreadsheet there are ways you can fix that but you can get those errors in it. And the third thing is that LLMs work best when they have sufficient context around a problem but not so much that they lose track of what matters. And with spreadsheets, you've got massive amounts of context encoded into that spreadsheet. Things like, the color of this cell tells you something.
[00:06:28] The fact that, the labels tell you that B6 really is an interest rate and that C6 is the principal inflated by that interest rate. You need those labels in order to encode it. And you can have a massive web of context that goes in all different directions that's enhanced by the visualizations.
[00:06:51] The AI has to basically take all that and then turn that into a text description, which it then has to keep in its context window, and that can end up being too much for it if you have some of those massive financial model spreadsheets that some financial professionals have. All of those are gonna cause AI to struggle with Excel.
[00:07:14] In the case that you actually mentioned, where you've got it giving you the correct number but for the wrong reason I would say that's a focus problem. Because AI's really good at optimizing for things and focusing on things. So Did you hear that account of the AI during some sort of security test, the AI broke out of its testing sandbox and hacked-
[00:07:39] Lauren Pearl: Oh, yes. This has happened actually because I'm doing a security newsletter. I know this has actually happened a couple times recently with Hugging Face and then with another company.
[00:07:48] Dr Genevieve Hayes: Yeah, exactly. I was thinking of the Hugging Face one. So the OpenAI agent breaks its way out of its little security box, hacks into Hugging Face in order to steal the answers to this cybersecurity test.
[00:08:00] Lauren Pearl: Yes.
[00:08:01] Dr Genevieve Hayes: And so the problem there was it was just really focused on optimizing for what it was told to optimize for.
[00:08:08] What you've got there is AI is focused on getting you the correct outcome. It knows that your goal is for the number in this cell to equal the number in that cell. So it's focused on getting that correct, and it's optimized it in the wrong way. , So it's optimization without taking into account the constraints because You've got implicit constraints in your head, which are, "I want you to optimize this subject to it actually being correct."
[00:08:39] But you never actually explicitly said, "Subject to it being right." Just like OpenAI never said, "I'd like you to pass this security test subject to you not committing a felony in the process." So yeah, we as
[00:08:53] Lauren Pearl: Let's add that to the skill. Don't commit felonies.
[00:08:56] Dr Genevieve Hayes: Yeah, no felonies. So what you've got, I think, is an optimization problem.
[00:09:00] AI got too good at optimizing without taking into account the constraints which were implicitly there in your head, you know, subject to it being right. And yeah
[00:09:09] Lauren Pearl: That's really interesting. So I'm curious, as someone who's like playing around with this and building some models with this, given what you know about LLMs, I know that you can't know for sure if AI will always suck at Excel, but knowing what you know, what could be done on a technical basis from your understanding as a data scientist?
[00:09:31] 'Cause I recognize you are not an engineer working at OpenAI or at Anthropic. But from your understanding, you mentioned, changing the constraints around the optimization is like one way of potentially tuning a model to get better results. What are the other mechanisms that you know about that could potentially make a model better at solving a problem in this-
[00:09:54] Dr Genevieve Hayes: The most obvious one is train it on more Excel models. So if you... train an LLM on a dataset that's made up of lots and lots of Excel spreadsheets. Then it will naturally get better at working with Excel spreadsheets.
[00:10:10] Secondly, those constraints as a user, you can put those in yourself. But there are actually commands that the AI companies put in behind the scenes. And if you actually talk to something like Claude you can actually get some understanding of what those instructions are.
[00:10:31] They won't give you them exactly, but they'll give you some idea. So you could have someone who's designing one of these just put in a requirement, hard coding those constraints. Those are the two big things that I can think of. Yeah, basically those
[00:10:48] Lauren Pearl: Okay, so giving it more spreadsheets and then building in strengths, which I presume is happening, 'cause I will say that it's been very interesting playing around with Excel as the models have been changing. So I recall using the Claude plugin before we got to Opus 4.7 or 4.8 or something, and it switched and suddenly, wow, it's more usable, fewer problems.
[00:11:12] Recently, it switched to Opus 5, I think we're at, and again, there was, like, a definite improvement of I ask for something, I look through the output, and there's fewer mistakes. So clearly they are optimizing. So interesting. It seems like maybe they're feeding it more models and probably trying to think about the constraints around what someone who's prompting a model might want but not say explicitly to get it to behave better.
[00:11:39] As I'm encountering this problem I'm putting it in the buckets of the many problems that work with AI and realizing that maybe I'm adding up to a taxonomy. And I'm curious what you think about it. We've spoken I think in the past about deterministic versus probabilistic.
[00:11:55] Or you mentioned it as one of the examples of things that AI doesn't deterministic problems because it's stochastic. Did I say that word correctly? Because it's a stochastic problem solver. And that's something I feel like is really well known in the industry, that AI can't do things.
[00:12:11] And that wasn't known for a while, but now it's like one of the sort of things we keep in mind when we work with AI, that it's actually a statistical tool. And so if you need a deterministic answer, it is more likely that you might end up with it is not your right output. But I'm starting to build a list of some other sort of rules of the road of if you're asking a question like X, you might expect a funky output.
[00:12:34] If you're asking deterministic. And the other one I'm doing here is potentially high context problems maybe. I've noticed that AI struggles in problem solve nuance where there's like a lot of context, maybe not even on a technical basis or on the file size basis, but just in terms of what it would need to know exact answer for an exact person in an exact scene.
[00:12:58] And then maybe I think this kind of leads to thinking about the article with Shannon versus Kolmogorov. Is there like another type of problem we need to add to the list? Like maybe where it's not a pattern matching problem,
[00:13:12] What rules can I add to my list
[00:13:14] Dr Genevieve Hayes: So as you said, LLMs, are fundamentally pattern-matching tools. They are correlation-based learners, and so therefore, if you're gonna ask it a question, it needs to find something similar to draw on in order to answer it.
[00:13:29] So if you ask it something like, "What's the capital city of France?" It's fine because it knows that. Or If you ask it, "How do I create a computer program to do something that millions of people have done before?" It's good because that information is all over the internet, so it can find the answer and use that to create an answer for your problem.
[00:13:54] But if you ask it to do something that's never been done before for example "Claude help me to design an engine that pulls static electricity from the atmosphere in order to run this engine without needing any fuel whatsoever that would literally involve creating an entire new branch of physics
[00:14:22] Lauren Pearl: Do that yet?
[00:14:23] Dr Genevieve Hayes: No, you can't do that. AI cannot create something that, doesn't exist from scratch because it's not there. No one's ever done this, so therefore any answer it gets, it'll just be recombining what it's already got, and it'll probably just take the plot of a couple of science fiction novels and mix it with some real science and come up with something that sounds very convincing.
[00:14:48] But because it can't create something from scratch, then it's going to fail
[00:14:54] Lauren Pearl: It's really wild because I encounter this problem, I think, quite often because I do a lot of writing. And actually, as a funny example, I even encountered it in prepping for this interview. Again, I use AI all the time in my work. And often what I'll do when I'm trying to get started is I'll take whatever notes , and use that kind of as a basis to get some ideas of where we can take conversation.
[00:15:14] This is especially true, I think, in normal human conversation. We go all over the place, especially I certainly go all over the place, as you know So I'll maybe use AI and ask it "How can we pull all these themes together?" It's awful at answering that question. It can tell me all of the different topics that were discussed, but if I ask it to pull it together in the way that I might if I were reflecting, like in a post on The Daily CFO where I pull , like what do chicken nuggets and AI and the resource-based view and, building skills have to do with each other?
[00:15:47] And my brain can just do that, and AI absolutely cannot. So that totally makes sense
[00:15:52] Dr Genevieve Hayes: Yeah, exactly. If no one's ever written an article about comparing AI and chicken nuggets, it will come up with something really lame-o because that involves creativity
[00:16:03] Lauren Pearl: So creatives are safe?
[00:16:04] Dr Genevieve Hayes: Creatives are safe, but it's how creative are you? Because let's take novels for example. Your mystery novels, how many of those are rip-offs of Agatha Christie?
[00:16:15] They-- You could probably just put in the work of Agatha Christie and ask, "Can you write me a new novel?" And given how popular Agatha Christie knock-offs are, it'd probably still sell pretty well
[00:16:27] Lauren Pearl: Now it's funny. I'm wishing we were back in our first episode where you were the host asking about rare skills. I'd add creativity to that list. We've discovered another one that AI can't do
[00:16:37] Dr Genevieve Hayes: But that's it. There'll be some creative person one day who comes up with some brilliant physics invention that's equivalent to pulling static electricity from the air. AI can't do that at the moment. I don't know if they will ever be able to do that. This is getting into artificial generalized intelligence, and there's debate over whether that's possible or not.
[00:16:58] But based on the fact that LLMs are correlation-based learners, I have my doubts about this being the path to artificial generalized intelligence, 'cause I would say that artificial generalized intelligence requires creativity, and you can't do creativity with just correlation-based learners
[00:17:17] Lauren Pearl: Interesting. It also makes me think about on some kinds of analysis, there's a lot of ways where things tend to go wrong where it could easily be solved with, pattern matching because we've been here before.
[00:17:28] But I would imagine that be if you're got a much more novel business model or something really funky or a situation where what's going wrong is a chain of different things that are really nuanced that AI might have a challenge giving you an analysis that actually points at the proper cause.
[00:17:46] Is that true? Is it also true at like a technical context?
[00:17:49] Dr Genevieve Hayes: Think about it. You deal with unusual startups as opposed to your average established business. If the AI has taken in a whole bunch of information to learn from about established businesses, and then you ask it about a edge case, which is your startup that you're advising, it's going to draw on its knowledge about your average generic established business.
[00:18:16] Because it has far more information in its training data about those standard cases. Every model is weak when it comes to edge cases. I hope that answers that question.
[00:18:28] Lauren Pearl: It does. It totally does. Okay, I have my final question for you, which is a bit of a absurd question, but hopefully that'll be a fun one to end on. Just going back to my tirade at the beginning of me searching for reasons why AI sucks at Excel, I wanna end us on the idea that I heard around what if we look back in 20 years and the whole big deal about AI was that it can program?
[00:18:53] As a data scientist, what do you think about that? How seriously should we take that idea?
[00:18:58] Dr Genevieve Hayes: I think, from the point of view of not so much a data scientist as a human being, it's what do you want to achieve with your life? And can you use AI to enable you to achieve that? If your goal is to build a solo business, for example, and AI enables you to do that, does it matter that the best use case for AI was writing better code?
[00:19:23] Because for you personally, that has helped. And I think that gets to it. It's a global versus local example. AI has helped me in developing my own business and in my data science work. I don't care if globally the best use case for it is writing better code. It's helping me, so therefore, yeah, I'm happy with what AI's done.
[00:19:47] I hope it doesn't destroy the world while doing all of this. I hope we don't end up in a Terminator situation
[00:19:54] Lauren Pearl: Hard same.
[00:19:55] Dr Genevieve Hayes: yeah or in The Matrix. But assuming we don't destroy ourselves in the process, use AI to help achieve whatever goals you're trying to achieve, and let everyone else worry about what globally the best use case for AI is.
[00:20:12] Lauren Pearl: I love it. I think that's such a good answer. And that's all the questions I think I have for you. That's not true. I could ask you so many more questions, but I will release you now because I feel so much more knowledgeable on the extent of AI knowledge and reinvigorated to continue to bang my head against a wall using it in Excel and other areas.
[00:20:33] So thank you so much for letting me reverse interview you. This has been super fun, as always,
[00:20:39] Dr Genevieve Hayes: Yes, it was fun. And thank you for being a wonderful host, Lauren
[00:20:44] Lauren Pearl: I tried my best. Thanks so much.
[00:20:47] Dr Genevieve Hayes: And that's it for today's conversation with Lauren. If you haven't already, listen to our previous episode where we discussed how data scientists can develop rare and valuable skills. And for those in the audience, thanks for listening. I'm Dr. Genevieve Hayes, and this has been Value-Driven Data Science.

Episode 126: Why AI Excels at Some Things and Sucks at Others
Broadcast by