Episode 116: [Value Boost] What Data Scientists Need to Know Before the AI Free Ride Ends

Download MP3

[00:00:00] Dr Genevieve Hayes: Hello, and welcome back to Value-Driven Data Science, where data professionals become strategic experts. I'm Dr. Genevieve Hayes, and I'm here again with Nicholas Kelly, co-founder and chief AI architect of Delivering Data Analytics and author of The AI-Driven Data Team. Last week, Nick and I discussed how data professionals can evolve their careers for the AI era.
[00:00:29] Today, in this Value Boost episode, we're exploring why organizations shouldn't be building their AI future entirely on rented intelligence and why understanding this puts you ahead as a data professional. Welcome back, Nick.
[00:00:46] Nicholas Kelly: Thanks for having me again, Dr. Hayes. It's always a pleasure
[00:00:49] Dr Genevieve Hayes: We are currently living through a period of subsidized intelligence. This year, for example, the creator of ChatGPT, is projected to lose $14 billion and isn't expected to reach positive cash flow until around 2029 or 2030. But the AI companies aren't giving the world access to their models out of the goodness of their hearts.
[00:01:15] Both OpenAI and its rival, Anthropic, are expected to launch IPOs within the next year. Once there are public shareholders to answer to, it's only a matter of time before the cost of using AI models starts going up, and any organizations or data professionals developing AI initiatives based on today's pricing, are going to be in for a nasty shock.
[00:01:42] Nick, you've argued that organizations building their AI future entirely on cloud AI frontier models are making a strategic error for the reasons that I've just described. Are there any other reasons why you believe relying on these models is a bad idea?
[00:02:00] Nicholas Kelly: Yeah, I think you hit the nail on the head. I would almost categorize it as a catastrophic mistake that they're making right now. And yes, it's cheap, it's easy, and for all the reasons you said, it's intended to be cheap and easy to get on. And you're putting all of your trust and faith into a fixed cost or let's say a managed OPEX that you're going to hope is gonna persist at that price, and it 100% is not gonna do that.
[00:02:29] But there are definitely other reasons. So we can look at HIPAA data. I think it's only really a matter of time before some of the more regulated industries will start requiring on-premise if not already. Some of them do certainly finance. But over time, obviously we're training the frontier models on all your data.
[00:02:49] Even if it's in a secure cloud, it's got all the good requirements and safeguards around it. Maybe it's not being exposed to other users, but it is informing the model. So we have to be very careful.
[00:03:03] So There's two really good reasons, and I think Dr. Hayes, you mentioned this in the previous episode, but the cost is always the one that it's ultimately gonna boil down to, that's going to drive behavior. And when they change the pricing, which they will, it's too late.
[00:03:19] It is already late. Everyone is gonna be scrambling to go on-premise or at least hybrid, and you should probably be looking to do hybrid right now, you should probably start looking there. But I would say cost is the biggest driver for why we should be looking at it right now.
[00:03:34] Dr Genevieve Hayes: I've already seen jokes on the internet about people who are accidentally spending, five-digit sums of money on their Claude bill or a CEO who accidentally spent $100,000 while building an app this is happening already and it's only gonna get worse.
[00:03:53] Nicholas Kelly: Yeah, it's only gonna get worse. One thing that woke me up to this was about a year ago, I was first looking at buying my own GPU rack and GPUs, and within the space of about three months, the GPUs, which were already high price, were just going up, and they're just going up and up in price.
[00:04:11] And why is that? It really bothered me. Seven, eight years ago, you could buy a gaming GPU, and it was, four or 500 bucks, and it's good, like top of the line. To buy the top of the line gaming card now you're talking two, three grand. And, a lot of that was driven by cryptocurrency mining and all of that,
[00:04:28] but then it migrated over to compute, so you can use these for compute, and it's been going up and up. Now, the thing though that I realized is buying a GPU right now is still massively undervalued, and the reason I think that I'll give you an example.
[00:04:44] There's a very capable GPU, NVIDIA, with 96 gigabytes of RAM that you could buy for your workstation, and it is very capable, and you can run local models on that of a reasonable size. It's around 10 grand, and that seems like a lot for a GPU. But if you can put that into production, run a bunch of AI agents off it,
[00:05:06] you could be 10X-ing the productivity of your employees, so for 10 grand, that's really cheap. The thing I struggled with is the price of the hardware. What we should be looking at is what's the outcome? And it's almost like in consulting when you talk about hourly rates versus being paid for, give me a percentage of the uplift that I bring your company,
[00:05:28] so value-based billing versus hourly billing. I think that's the case we're looking at with GPUs currently, is we're looking at them from the hourly rate. It's "Wow, that's a lot for a GPU." And we're not really looking at how they're pricing it, 'cause NVIDIA and the other GPU organizations are partly pricing in value of what it can do, but that's gonna go way up.
[00:05:52] So once you're talking about those IPOs that's also gonna drive up the value-based cost of buying that infrastructure.
[00:06:00] Dr Genevieve Hayes: What you should be looking at is the total cost of the hardware plus what you would spend if you didn't have the hardware using the frontier models and so it'd be a hardware plus software. And if you don't have the hardware at all, you're 100% dependent on the software, and if you have the hardware, you're gonna reduce your software bill.
[00:06:20] The complication here is some organizations will have separate budgets for capital expenditure and software expenditure, so some of them will be actually happier to have a higher software cost. But excluding that situation I can see the argument that you're making.
[00:06:39] Nicholas Kelly: Yeah, 'cause right now the frontier models are cheaper and the cost of GPUs, while they seem very high, are cheap. But once, yeah, once those IPOs happen, once they start building in more of the value into the cost, It'll be cost prohibitive. I think it'll be, just won't make sense. But at that point, the cost of the infrastructure is also going to be prohibitive.
[00:07:00] So as to how much it costs, Dr. Hayes it totally depends. For experimentation, we can get one of these little NUCs, like there's NVIDIA DGX Spark. They're like five grand, and it's 128 gigs of shared memory.
[00:07:13] That's awesome. That's great for experimenting. I think most BI teams, data teams should have at least one of those. There. Start there. Start with that. Experiment. That's the most important thing right now, Having a platform to experiment on where you've already incurred most of the cost.
[00:07:29] After that, you can either get more of those or build a rack It really depends though. There's a wide range of cost. But right now we're looking at maybe anywhere between 50k and $200,000 to have something decent that you could put into production within your own organization
[00:07:45] Dr Genevieve Hayes: And you're not saying that people should 100% work on this GPU, you're advocating for a hybrid approach, is that right?
[00:07:53] Nicholas Kelly: Right now hybrid with a mix of you can put on some of the open models. Obviously, you need to review them, make sure they satisfy your own requirements, your own internal security protocols and everything. But if only just for running agents, you can start there.
[00:08:08] You don't necessarily need to run inference locally. You can still run that off the frontier models. But I think ultimately you want to be setting yourself up to be completely independent. And if you can still use hybrid in the future, hey, that's awesome. But you don't have the same risk exposure anymore.
[00:08:25] It's insane. When you just stand back and look at this Dr. Hayes, you might know better than me, but if you remember that transition from on-premise to the cloud, that was a big thing. «Hey, we have to get on the cloud», and for whatever reason, we thought that was a good idea, and in some ways it is a good idea, but some companies went exclusively into the cloud.
[00:08:42] And now we're looking at a pullback, or we will be looking at a pullback as actually, that's a little bit risky. We need to bring that back. And at least have a hybrid model. And you want to do it before things go completely bananas. 'Cause it looked like it went bananas already. But I really believe that we're not seeing the value priced into the infrastructure at all yet
[00:09:04] Dr Genevieve Hayes: So it sounds like what you're describing is you use the on-prem hardware with something like Ollama or some sort of open model loaded onto it for your development environment, or if you have any highly sensitive applications, so HIPAA data financial data, things like that. And then if privacy considerations permit you would then take what you developed in this development environment and then potentially put it into production using Frontier models so that you can have the benefits of the latest and greatest technology in that area.
[00:09:43] Nicholas Kelly: Absolutely. That's a really good approach. So the local model you use is driven by the size of the video RAM that you have. And so you're generally gonna be talking about what's feasible is like a 80 billion parameter model. So like a Qwen model some of the Google models,
[00:09:57] there's a bunch out there. The ones that are really good are like, Qimi 2.6 or some of the larger parameters like GLM 5.1. They're not on a par, but they're up there with the frontier models. You could say they're like one or two versions behind, maybe two versions behind in capabilities of the frontier models.
[00:10:14] But that's not gonna matter in six months. The capabilities they're already so good. The capability are gonna be so much better in that amount of time. Like for our own business, we still use the frontier models for coding. We're in for that, and right now the price makes sense.
[00:10:29] But we're also set up that we can also use our own infrastructure and the coding models that are available that we can just download and use for free, with the caveat that you just have to be careful that you're not being overly influenced by however that model's trained,
[00:10:43] right?
[00:10:44] Dr Genevieve Hayes: Yeah, Fair enough. So most data professionals have grown up in a world where IT infrastructure decisions were someone else's problem. But at the same time, as some of the biggest users of AI, stakeholders are going to be looking to them for advice when making decisions around AI models and infrastructure.
[00:11:02] This happened to me when cloud started becoming a big thing. For a data professional who wants to be the strategic voice in this conversation inside their organization, what are the most important things they need to know?
[00:11:17] Nicholas Kelly: Yeah. You need to know about agent orchestration. So I would advise get OpenClaw, start there. If you can, I would love you to do this, get like a Mac Mini or there's the AMD Strix, I believe, is recently released. So the lowest cost way to do inference right now is on devices that have shared RAM with VRAM.
[00:11:41] That's your lowest point to get into these things, so you're probably talking like two or three grand. But if you want to get into the space, table stakes, do it. So OpenClaw. Obsidian. The next problem you're going to run into is context, saving context and having it portable, so things like Obsidian.
[00:12:01] Obsidian's not the only thing, but it's open source. You can start using it right away. Then you need to figure out the models, so that's the next one. And a really good place to start is either the Google models or the Qwen models. You want something small, so you're going to start learning about the number of parameters in a model, and you're going to hear these terms like the density of the model, and the parameter probably being the larger driver of what can you actually run on your hardware?
[00:12:29] You're going to start thinking in terms of video RAM. What can I run? Which models am I going to use for what thing? You're going to start wanting to build agents. The rack we have here, there's over 120 agents running right now. And, made loads of mistakes, piles of mistakes.
[00:12:47] That's why you want the hardware, you want your own hardware, and your laptop might have a GPU on it. Maybe, just check it out. Maybe it's already enough. But I would start there. They're the main things 'cause they're like the table stakes. You just have to know this stuff.
[00:12:59] You have to figure it out. And You can read about it for sure. You can read about it and learn a lot, but you're just going to learn more by doing it. And it's like I described in the last episode, you want to do it on something that's personally interesting for you, and you will be just blown away if you haven't used it before with OpenClaw.
[00:13:19] It can just automate a whole lot of tasks for you in your day-to-day, so that's where I would recommend starting
[00:13:25] Dr Genevieve Hayes: That's it for today's episode with Nick. If you haven't already, listen to our previous episode where we discussed what the AI era means for your career as a data professional. Thanks for joining me again, Nick
[00:13:39] Nicholas Kelly: Thanks for having me, Dr. Hayes. I look forward to the next time
[00:13:42] Dr Genevieve Hayes: I will be contacting you when you write your fourth book, and we'll be inviting you back then.
[00:13:48] Nicholas Kelly: Thank you kindly
[00:13:49] Dr Genevieve Hayes: And for those in the audience, thanks for listening. I'm Dr. Genevieve Hayes, and this has been Value-Driven Data Science

Episode 116: [Value Boost] What Data Scientists Need to Know Before the AI Free Ride Ends
Broadcast by