Deep Geeks

The Revenge of the Middle: Managers, Memory, and the Real AI Bottleneck

Episode Summary

AI spending is rising faster than anyone budgeted for, and the unit price isn't the reason. Dr. Serena Huang explores the hidden costs of enterprise AI with Val Bercovici of WEKA.

Episode Notes

AI spending is outpacing budgets, and most organizations don't know why. Dr. Serena Huang sits down with Val Bercovici of WEKA to unpack why enterprise AI costs keep rising despite falling model prices. They dig into the three tiers of enterprise token consumption, the memory shortage reshaping AI infrastructure, and how signal maxing can be a better metric for enterprise AI ROI.

Timestamps:

1:30 - Meet Val Bercovici

3:44 - Jevons paradox: 100x cheaper per token, 100x higher net bill

6:52 - The real cost of agentic AI

8:31 - Three tiers of enterprise token consumption

11:16 - How agents learned to cut their own costs

16:48 - Understanding AI Pricing

18:25 - Signal maxing vs. token maxing

24:39 - AI Memory: The cost driver no one budgeted for

35:14 - What happens when AI subsidies end

47:34 - Revenge of middle management

Links: 

Connect with Serena

Connect with Val

Episode Transcription

[00:00:00] Val Bercovici: There's literally not enough rare earth minerals in the world and not enough manufacturing fabrications, labs to be able to actually produce the memory required today. And so this, this memory wall is literally us hitting a wall as an industry of not being able to satisfy the overwhelming demand for tokens and needing much more clever approach

[00:00:30] Serena Huang: Welcome to the Deep Geeks podcast. I am here today with Val Bercovici, who is the Chief AI Officer of WEKA, and I can't wait to dive into today's conversation because we had such a great prep. We almost, um, couldn't stop talking about AI. So um, Val, welcome to the show. 

[00:00:51] Val Bercovici: Great to be on again, Serena. 

[00:00:54] Serena Huang: Well, I just came back from a conference where every company there was showing off their dashboard of token usage by team, by function, and, you know, for a leaderboard even.

Everyone was so proud. And there I was thinking, gosh, when we all get the bill in a quarter, you're probably not gonna be as happy about it. And then as I got back, I also heard that Uber's COO just announced that they weren't getting the ROI they were hoping from AI, and it's getting so expensive. That token usage isn't translating into features for customers.

Val, are you surprised by any of this? 

[00:01:39] Val Bercovici: Not at all. I mean, uh, I guess I'm surprised how long it took, to be honest. But I, uh... My, my memory kinda goes back to the early cloud days when- For sure ... a couple of developers, there wasn't a lot of IT people back then, but a lot of developers and a couple of the executives close to developers were super excited about what they could do.

The agility of developing and deploying applications in the cloud was super exciting. Uh, and then of course, the bills came in a few months later, and, uh, the sticker shock happened, and this whole category of, uh, of FinOps, right? Just financially- Sure ... engineering applications to not break the bank from a cloud bill perspective became a real formal discipline not long after the rise of cloud.

We're seeing the same thing play out probably faster now, but, you know, I, I kind of was expecting people to have the sticker shock, and I was surprised by the whole token maxing trend, right, that we saw in the spring where everyone is literally just almost inflating token numbers just to game leaderboards, and obviously, uh, everybody wants to be at the top of a leaderboard, but we knew that couldn't last.

[00:02:42] Serena Huang: Yes. Well, I think a lot of companies are surprised, um, in so many ways that they were perhaps not giving the right incentives to users, and then also on the other hand, the cost of AI is surprising in a very unpleasant way, too. So, um, I'd love to just talk a little bit more about the economics of AI and token, tokenomics.

Is that a word, Val? 

[00:03:10] Val Bercovici: Absolutely. Okay. It, it actually comes... Some people in the crypto world have their definition of it, but I think- Ooh ... like many things, the AI industry has kind of co-opted it and taken it over to mean its own thing. 

[00:03:21] Serena Huang: Yeah, and I think people are stuck in this old narrative that AI is going to get cheaper, and we will use more of it, you know, cost goes down, demand goes up, Econ 101, and then cost eventually somehow normalizes.

I don't think that's actually happening, and I'm curious if you can share with us just what's driving that cost and what's underneath. 

[00:03:45] Val Bercovici: Yeah, it's easy to look at just one side of a coin, just one narrative here, 'cause, 'cause you could factually claim and be correct that the unit cost of inference is plunging, right?

The actual per token cost, and I think DeepSeek recently made some pretty big headlines, not just with- Mm-hmm ... the V4 model announcement, V4 Pro and Flash, but their shockingly low pricing- Right ... particularly for cash read tokens. So you could argue, hey, look at the price of inference. It's plunging. What is happening on the other side of the same coin, though, is Jevons paradox, is every time a valuable commodity decreases in price, we're seeing the consumption dwarf that reduction in price.

And the math is like for every, you know, 100X reduction, we're seeing like a 10,000X increase in consumption. So that net difference that, you know, the, the final, the final cost of inference is almost 100X more than we projected it right now, which is why we're seeing the bills that people are beginning to annualize come in, the- Mm

5, 10, $15,000 a day bills- Yes ... for people that token maxed, doing the math at $5 million annualized cost per user just doesn't math. 

[00:04:59] Serena Huang: Wow. Right. Wait, so 15,000, not even per month, you're talking per day just on AI? 

[00:05:06] Val Bercovici: The, yeah, the biggest token maxing boasters, if you will, were at that figure because, hey, it was a great figure.

[00:05:13] Serena Huang: Wow, this is, um, this is going to be shocking. I think, I'm thinking about the CFOs that I've talked to who are, of course, having to think about budget on a much longer term basis, and now on a quarterly basis they might even be surprised because the presentations have been saying the unit cost is falling, but no one was preparing for the forecast of that 100x or even 1,000x usage, and that bill is going to surprise us in, in a year, um, that no one is prepared for.

I think the other part of this conversation that is a little uncomfortable is a lot of companies are making decisions about their talent, their workforce, using this incorrect assumption that AI is getting cheaper- Yeah ... and AI is cheaper than humans, and we are not really seeing that, are we? 

[00:06:07] Val Bercovici: No. In fact, uh, one of the most popular new commands of all the agents, whether it's Cloud Code or OpenClaw or Hermes, is a /goal command.

Ooh. And that's where hopefully you specify something really well in terms of an ambitious project or app that you want, uh, and it will go off. The agents swarm. The agent will spawn sub-agents. Yeah. And it'll go off for not just, you know, seconds or minutes the way we're used to chats, but hours and even days, and iterate and test and improve, but ultimately deliver a highly functional app with great built-in security, with enterprise authentication and auditing, everything you'd expect, but that token cost, you know, for that project will be tens of thousands of dollars.

And so what people are realizing is, yeah, you've got to figure out, you know- This token consumption, uh, it's, it's not as efficient as humans in many cases. And, uh, and it's, it's costing more than humans when applied incorrectly. 

[00:07:06] Serena Huang: Right. And that's, that's a really good reminder 'cause, um, I think we, we were definitely not prepared for that.

And our brains like to simplify things because, well, ambiguity is tough on our brain. So I feel like there's a narrative about AI cost that is a little bit too simple. Like there's one AI, and then it costs the same price. But when we were preparing for this conversation, you mentioned there are actually different tiers that, yes, as someone who uses AI, I might be aware of it, but, um, ultimately at the leadership level, I'm not sure everyone's paying attention to that.

So could you walk us through the different tiers of AI cost that we talked about? 

[00:07:53] Val Bercovici: Yeah. And in fact, that phenomenon you described actually has a name over the last few weeks, which is called AI psychosis. Oh my gosh. And that's like, you know, the, the further you are, yeah, away from, from actually hands-on with these tools, the simpler the mental model you have.

And I think one of the traps that people, especially leaders, can fall into if they're not hands-on, is applying chat mentality- 

[00:08:16] Serena Huang: Mm-hmm ... 

[00:08:16] Val Bercovici: to what agents actually are and how they work. They're fundamentally different in almost every way. We think of chats as a, you know, we have a chat with a model And if we don't like it, we kind of start a new chat with a better model, or if we think the price is too high, we start a- another chat with a, with a lower quality model.

Uh, agents, you know, don't work that way. They don't just work with one model, and they don't just have a prompt and response. 

[00:08:39] Serena Huang: Right. 

[00:08:39] Val Bercovici: They basically have multiple turns, dozens, hundreds of prompts and response turns. They often include lots of context and not just one PDF, but entire libraries or corpus of, of material, entire- Right

code bases and so forth, entire, you know, Wikipedia-sized knowledge bases. That, that is a very, very different experience, which means that we're-- we need to have segmentation. It's almost like an old MBA thing. As any market matures, it segments. So at the- Yep ... very least, we're seeing these three common segments emerge, and, and Anthropic has done a really good job of, like, branding these segments.

You've got- 

[00:09:16] Serena Huang: Mm-hmm ... 

[00:09:16] Val Bercovici: Opus as the high-end premium model with premium token pricing. You've got Sonnet as sort of the, the, the budget model for just, you know, higher volume of tokens, uh, but at, at a slightly lower quality. And then Haiku, which is like the, the budget, the budget b- bargain bin where you have- Yes

really, really high volume bulk tokens, but you don't expect deep thinking there. You just expect execution of very explicit tasks. And the combination of all three are what developers are using right now for latency reasons, if you don't have- Mm-hmm ... or if you have an unlimited budget, but then obviously with real budget, uh, you know, requirements in hand for budget reasons, financial budget reasons- 

[00:09:58] Serena Huang: Right

[00:09:58] Val Bercovici: uh, we're, we're mixing and matching these models because right tool for the right job applies in AI as well. 

[00:10:04] Serena Huang: Yeah. I, I like that a lot. I think that's a great reminder for people, right tool for the right job, and right tool means different tiers. I'm curious if you have even seen, um, let's not call it incorrect usage, but maybe using the wrong tier.

Like, the job doesn't really need an agent, but somehow, you know, may- maybe a regular chat would do or a project, but instead an agent was created. Yeah. Have you seen anything in real life? 

[00:10:31] Val Bercovici: Yeah, like one of the, one of the, I think, really positive impacts of, of OpenClaw coming on the scene in December and going viral is it made the technology much more accessible- Mm-hmm

not just to developers that love hacking on it, but to a lot of tinkerers and hobbyists in home labs- Right ... that go to work and then do knowledge work, right, 9:00 to 5:00. And everyone quickly realized, you know, they did their own math very, very quickly, realized they were blowing through their personal token budgets at home, and just implemented policies, just literally told the agents in plain English, "Reduce my costs."

And the agents quickly figured out, "Well, how do we do that? We use more than one model. We use, you know, some models obviously for simpler work." And there's this, uh, part of these agent frameworks harnesses which just does a regular scheduled task in the morning, like check the weather or check- Yeah ... a stock price.

Uh, developers know that as a cron job, right? Just a regularly occurring task. And, and, and the OpenClaw hackers quickly realized, "I don't need a model to run this cron job," right? "I can just run a cron job." Like, I can just- Right ... use the, the tool directly instead of actually having an agent call a tool, uh, and, and add an artificial layer of latency, but especially cost.

Yeah. And that's what I love, is these agents themselves going out and researching common best practices instantly developed internal policy agents within, within themselves to say, "Okay, here's the complexity of this particular task. I will use a premium model. But here's a lower complexity, I will use a budget model.

Here's just simple execution, I'll use a bargain basement bulk model." And that's the built-in hierarchy that agents learned autonomously- Right ... and are deploying independently. 

[00:12:16] Serena Huang: Wow, I love that. That's so smart. So now there was another, another recent headline about Microsoft where they've moved away from Claude Code because of cost, and they're asking people to go back to GitHub Copilot, which I've used and it's a great tool, but the cost was kind of the driver for that decision, it sounded like.

Do you think they are making a tier decision, or is this more just overall cost cutting? What, what are your thoughts on that? 

[00:12:51] Val Bercovici: So that, that's a really important headline. I, I noticed that the other day myself. I'm gonna speculate that there's at least two aspects to it. We can't ignore the fact that, you know, when one company has a product and their internal developers, users are using the competitor's product-

then that doesn't look good, right? So you do need to be dogfooding for all sorts of reasons, your own product. So I'm sure that contributed to it, but no, I think it's representative of what I'm seeing at other startups, other software companies, and some, you know, some, some dev heavys, uh, enterprises, which is the fact that the token costs are, are really coming home to roost right now.

They're real. To- we, we like to say people are, are moving from token maxing to signal maxing more right now. Mm. Trying to extract more ROI, more signal from the noise of all the tokens, and in doing so, it, it's clear now that, um, we have not just Anthropic or OpenAI, but the open weights models. You know, whether they're the, the resurgent sort of US open source models, but especially the Chinese models that are amazing.

Right. And, uh, and we're seeing a, a natural hierarchy of, okay, maybe there's only a few frontier models that are the premium ones. Mm. But now instead of paying, you know, OpenAI or Anthropic for their middle tier, their medium tier, I will go to DeepSeek, I will go to Kimi, I will go to GLM- Right ... I will go to Minimax, and I will get very much the equivalent sort of sonnet mid-tier functionality, but I'll get it at, you know, 10% of the price.

Yeah. And that enables my agents to be so much more powerful and think and validate more. 

[00:14:24] Serena Huang: Right. Absolutely. And I, I remember even being impressed by DeepSeek, um, two versions ago. Yeah. And it's just incredible now. 

[00:14:36] Val Bercovici: They continue to deliver, right? They kind of went dark for, I think the, the figure- Right ... was 484 days.

And, and th- that's a dangerous- Oh, killing. I know, it's crazy. And it's a dangerous thing to do, right? Because their competitors out of China were just shipping almost every week. 

[00:14:51] Serena Huang: Right. 

[00:14:51] Val Bercovici: Newer and newer models leapfrogging each other. It was kind of breathtaking to observe. Yeah. And then they came out like almost, you know, over a year later and just wham.

Right. This other fundamentally, you know, powerful game-changing release. And cool thing about DeepSeek, not to be too much of a fanboy, but, uh, people in the industry love them 'cause they publish so much. Yes. It's not just about shipping a model, it's about these amazing white papers they publish. Uh, and, and they went ahead and said, "It's no longer that important to compete at the top of the benchmark leader boards," 'cause the models- Mm-hmm

are converging in some sense. Yeah. They're certainly all converging on benchmarks, 'cause that, that's what benchmarks do, is they kind of force you to, to, you know, train for the benchmark. But, uh, the cost, right? Back to the cost again, and seg- segmenting, you know, the tiers, uh, DeepSeek really innovated on the cost, and, and they were able to introduce scalable, workable compression- 

[00:15:43] Serena Huang: Right

[00:15:43] Val Bercovici: that really compressed this memory consumption from KV cache by 90%, where most of their competitors really couldn't compress and, and, and n- and retain quality, uh, having, you know, b- best of both worlds. So that's the magic of DeepSeek, is that's why their prices can be so low, because they use less memory and they haven't sacrificed the quality.

[00:16:02] Serena Huang: Yeah. I- I'm a fan as well. But let me just pause for a little bit and give our audience a quick summary, um, on what we just talked about. For anyone who's listening, before you sign or renew an AI contract, ask which tier your workflows actually require in your organization. Because if you're paying the premium prices for basic work, that's very expensive, and you have a big procurement problem on your hand.

[00:16:33] Val Bercovici: I think, yeah, you know, it's even first principles now. Uh, most procurement teams are used to negotiating with SaaS vendors. 

[00:16:40] Serena Huang: Yes. 

[00:16:40] Val Bercovici: And so it's very much like, let's budget on a per seat basis. And, and what we're quickly realizing, and we're in the middle of this right now, so I think it's messy for the industry, it's confusing for customers, uh, is there's no such thing as per seat production AI pricing.

There might be some floor- Mm-hmm ... some per seat entry, but everyone is moving towards consumption-based pricing. Right. So thinking of software as a utility is very new to procurement teams- Yes ... because it's, it's today be- between hard and impossible to budget for. And of course the, the, the, the prediction tools, the measurement metrics, the measurement tools will improve, and we'll be able to project consumption-based costs better over time.

Right. But we're really not there yet as an industry. So to your point, uh, you really have to do a lot of due diligence now in signing up to these AI contracts, because you are writing a blank check in some sense. 

[00:17:34] Serena Huang: Yes. Um, that will keep any CFO up at night, a blank check on AI. And let's, uh, let's talk a little bit about token maxing.

I know we started the conversation around dashboards that I saw about usage, and I also see from the people side where organizations are now incentivizing AI usage. It's almost like show... You know, use AI, just use it more, show us the numbers. And right now I get it, right? Adoption is hard, and especially for large organizations, it might take time to get to ROI, and the board is pushing hard on show us the numbers, show us AI is making a difference.

Um, but what, what is... I guess, tell us what's wrong with token maxing, and what would you recommend instead? 

[00:18:25] Val Bercovici: Yeah. I, I like to think of these kinds of trends in terms of bell curves. 

[00:18:28] Serena Huang: Hmm. 

[00:18:29] Val Bercovici: So at the front of the bell curve are people that have gone through, you know, all of the early adoption pains and lessons and, and now are on the other end of it and have figured out how to sort of get true ROI to signal max.

But a lot of the industry, sort of the meat of the bell curve today across the Global 2000 worldwide, is still really in the how do I apply AI at work phase- Mm-hmm ... and where's the business value of AI. Right. I am observing, you know, uh, I guess maybe it's anecdotal 'cause I'm not a professional analyst, but there's enough day-to-day evidence I have in talking to customers and partners all day long that there's no question about the individual productivity benefits of AI.

You know- Mm-hmm ... it's, it's hard to find any internal slide deck or internal spreadsheet or internal dashboard that isn't AI-generated today. Uh, you know, it's, it's the group productivity- Right ... that people are still struggling with. 

[00:19:20] Serena Huang: Hmm. 

[00:19:20] Val Bercovici: But, uh, we're still seeing a lot of people, a lot of executives, uh, if, if they're not headline chasing, clout chasing, or certainly their earnings reports a couple of quarters ago, it was very important to say that my company was not behind this trend, that my company was AI savvy, and investors wanted to hear that you were leveraging the productivity benefits.

So we're still in that, uh, in between zone right now of leadership encouraging, you know, staff to use AI aggressively. 

[00:19:48] Serena Huang: Mm-hmm. 

[00:19:49] Val Bercovici: Uh, incentivizing in some cases, where we saw the token maxing trend emerge, like some people, you know, um, literally got bonuses if they were at the top of these leader boards. Right.

There's certainly a lot of internal recognition, promotions- Yeah ... and clout, and that drove, again, to your point, that was unaligned behavior, right? Because the incentives were really wrong there. 

[00:20:08] Serena Huang: Yeah. 

[00:20:08] Val Bercovici: It's not just about, you know, how much gas can you burn. It's about, you know, what kind of mileage are you getting, right?

Ooh, yes. And so that, that's where we are today is I think people are realizing now that, yes, we need to go beyond just burning a lot of gas. We need to look at our mileage. 

[00:20:23] Serena Huang: Yes. I like that. And, and I think this is why, um, I love analogies, as you know. And, um, incentives really matter, so in signal ma- maxing instead, what should a leadership team think about and incentivize differently?

What, what is a signal? Can you break it down for us? 

[00:20:47] Val Bercovici: Yeah, there's a couple of different ways to define it. You know, engineers have had this term for a while called goodput, which is literally, you know, it's output that's useful as opposed to just output, output that keeps the system busy. 

[00:20:58] Serena Huang: Hmm. 

[00:20:58] Val Bercovici: Uh, you know, we even in the AI industry for training, we used to measure, you know, how busy can you keep the GPUs- Yep

'cause the GPUs are so expensive that if they're idle, it's a huge waste of money. So people just figured out how to just run a lot of inefficient training scenarios, right? 

[00:21:14] Serena Huang: Right. 

[00:21:14] Val Bercovici: So what did that, what did that do? On the one hand, it definitely kept your GPUs busy and you felt like you were getting your money's worth.

On the other hand, you weren't really getting model checkpoints, model versions shipping any faster. So the efficient model training used to be one important AI metric of how quickly can you checkpoint, how short are these epoch times that you have. That was a very, very important metric. Today, I think it's really all about tokens, for example.

So, uh, when you're, when you're signal maxing, you're actually getting a lot of verified useful output, either an accurate spreadsheet, a relatively bug-free or functional application- Right ... a dashboard that, that works and is useful and- Yeah ... has a good user experience, and you're not consuming a billion tokens for that, which is, has happened.

Right. But you're consuming maybe a million, maybe only 100 million tokens instead of the billion. So it definitely is, to, to your point earlier on About not just using the most expensive model for everything. It's about understanding that some models are better at planning, others are better at delegating, others are better at doing the work that's delegated.

Others are better at just verifying that the work is accurate because there's clear yes/no criteria, reward functions, if you will, as to what's a, what's a successful outcome or not. Yeah. And it's the engineering of these together that results in great signal-to-noise ratios and good signal maxing and, and less and less token maxing over time.

[00:22:41] Serena Huang: Yeah. I like that. Um, signal, signal per token instead- Exactly ... is, um, is a much better measure. And, and, and I think this is going to be so important for everyone who is listening, thinking about setting goals for even their next year's, um, you know, for, for their team, and how do we continue to reinforce AI as important, as life-changing, and, and has lots of potential without incentivizing the wrong behaviors.

Um, so, um, yes, AI adoption, and, and I like, I really like the, I'll call them business metrics that you mentioned too, right? Is it actually not just, um, not just output that is useful, but impacting the eventual customers. You created a dashboard, great. How's the user experience? Do people actually use it? Is it leading to better decisions?

And that's the KPI that we have always measured in successful businesses, and not so much on the number of hours that you worked, um, or in this case, how many tokens you have burned. 

[00:23:48] Val Bercovici: Exactly. It's the old adage, you can't manage what you can't measure, so. 

[00:23:52] Serena Huang: That's right. That's right. Okay. I'm gonna shift towards the memory, um, and really give our listeners a, a flavor, a bit of what's really driving the cost, because a lot of people I talk to, they, they have heard of GPUs, but they have never...

And they're not going to touch one, right? Mm-hmm. And, and the cost may be really confusing because it seems like in SaaS days, it is per seat, and actually it should go down. Why is it going up? Um, so I want to set this up as a, um, just a chance for us to speak to the VP of AI or a CFO, um, or someone who's negotiating an AI contract.

How should they think about the bill, and where does memory play a role? 

[00:24:42] Val Bercovici: This is where the stock market's actually helping educate a lot of leaders, because if you just take a look at the stock prices of the, the memory suppliers and the, the storage suppliers, uh, the flash storage suppliers, non-volatile memory, it just, you know, they're like meme stocks.

They're just astronomically high right now, only over the last three to six months. And that's a direct reflection of the fact that as the AI industry is maturing and really dramatically shifting from a focus on just training to now monetizing these models that we spend so many billions training- Right

inference, those are just fundamentally different workload profiles. You can pretty much build different data centers for training and inference. Mm-hmm. Uh, and we're actually seeing a lot of vendors come to market, and Nvidia themselves made this high-profile acqui-hire of a company called Groq with a Q back in December- Right

of last year, and that was an inference-only basically accelerator. It's not even a GPU anymore, it's an LPU, a language or a linear processing unit. The Cerebras IPO recently, right? 

[00:25:42] Serena Huang: Right. 

[00:25:42] Val Bercovici: The other week was so popular and oversubscribed because they're focused on inference. They're not really focused on training.

So people are realizing now that Inference is dominating the infrastructure discussion, and inference, when you speak to the actual technologists, the researchers, the inference operators, inference is more memory-centric. Mm-hmm. The math, the simple math, round number math is for every, m- you know, 100,000 tokens, which is a very, very typical sort of prompt nowadays for an agent, that 100,000 of tokens used to translate to, which is about, you know, a, a megabyte of actual data, used to translate to about 30 to 50 gigabytes, gigabytes, not megabytes anymore, of memory.

Actual, expensive, high bandwidth memory. Some of the, you know, most precious real estate in the world, as we like to joke from a silicon square millimeter perspective. Yes. And, uh, and that's just not tenable, and that's one of the innovations that DeepSeek was able to bring to market, and many others have also sort of followed, is let's reduce that giant memory amplification from one megabyte to instead of 30 to 50 gigabytes, only three to five gigabytes.

Wow. But as we spoke earlier on with Jevons paradox, as we reduce the size of that memory per unit, the amount of units now explodes 100x, well, 10,000x, but 100 net x. And what that means is we're just memory bounding, and there's not enough memory. Yeah. And when you take a look at the actual supply chain, the way analysts I speak with- Mm-hmm

have taken a look at it, there's literally not enough rare earth minerals in the world and not enough manufacturing fabrications, labs to be able to actually produce the memory required today. 

[00:27:27] Serena Huang: Wow. 

[00:27:27] Val Bercovici: And so this, this memory wall is literally us hitting a wall as an industry of not being able to satisfy the overwhelming demand for tokens and needing much more clever approach.

DeepSeek is one of them. They optimize the KV cache size in one dimension. One of the things I've been working on at WEKA, the, the personal reason why I joined, is the ability for WEKA to actually bring abundance to this environment, because- Mm-hmm ... when you can actually take the cost basis of storage, but make storage behave honestly like memory, not like storage, so much faster, much lower latency, you're now adding abundance to this very scarce infrastructure reality.

And that combination of better per unit efficiency on the DeepSeek side and abundant memory infrastructure from the WEKA side is what the industry needs right now to scale the memory wall. 

[00:28:17] Serena Huang: Um, this is probably the best use of the word abundance I have heard in a really long time, along with scarcity.

Exactly. Um, I like that. I like that. Um, I, I think I'm gonna have to turn that into a new hashtag at some point. And, and I think what you described there is so important, and I want to give our listeners a, a picture, if you will, that they can take home. So, um, and what I'm hearing is that GPU in the analogy maybe is like a stove in the kitchen, right?

And so you can have a really good stove, top of the line, but then the memory is more like the prep chef or the prep table next to it. So you can have the best stove that is cooking amazing food really quickly, but if your prep chef or your prep station cannot keep up, you can still not produce as, as many fabulous dishes as you would like, right?

How d- how does that analogy land with you? 

[00:29:23] Val Bercovici: If you're a fan of that TV show The Bear, then you're... it lands perfectly, right? Because whether you watch, you know, kitchen reality shows or whether you watch movies or, or, or, or TV series about this, you see it, you know, uh, as plain as day in that, uh, there's real bottlenecks in the kitchen- Mm-hmm

and, uh, and there's obviously very high-priced equipment, like you said, in the kitchen, but if you don't have an efficient system, then you're just leaving those expensive resources like that wonderful $20,000 stove- Yes ... sitting idle 99% of the time, and not only is that a waste, you have very, uh, hangry, unhappy customers right at the end of it also.

[00:30:02] Serena Huang: Yes, and we don't want any hangry customers, for 

[00:30:04] Val Bercovici: sure. Exactly. No hangry token users. 

[00:30:07] Serena Huang: Yes, and, and I think that's another metric that companies might need to think about as well, is that, that utilization. If your GPU cluster is running only, you know, 30% utilization, um, you are paying the bill. You are still paying for the full bill, right, Val?

[00:30:25] Val Bercovici: Oh, yeah. No, the bill always comes home. So yeah. 

[00:30:30] Serena Huang: Um, if I were to summarize this, um, around the memory constraint that a lot of companies and leaders are not thinking about, the costs are not falling because, um... well, uh, they're not falling in the way that the presentations showed you, uh, or promised, and it is, it is now, you know, kind of like, um, maybe your cell phone bill.

Remember, you know, like when- Exactly. Exactly ... texts used to cost money- Yeah, yeah ... per it was like, a per text cost, and now you just pay the full bill, and you can even send images. You can send videos to your friends, and, um, and the cost of each text is almost zero. But guess what? I don't know about you, my bill didn't go down.

It actually went up. Yeah. And I'm using it so much more now. So, um, so that's the same for AI and, um, you know, depending on which analogy you like, the phone bill or the kitchen, if you like to eat. Um, it's something to keep in mind, especially for long-term budget planning purposes. And, um, and then also maybe not make your workforce and talent decisions so quickly based on what you think the cost is going to be, um, 'cause you could be in for a very unpleasant surprise 

[00:31:46] Val Bercovici: The bill, I think in addition to the kitchen's a perfect analogy, because I still, you know, maybe not all of us remember complex mobile phone bills in the early mobile phone days.

[00:31:54] Serena Huang: Right. 

[00:31:55] Val Bercovici: But they were pages and pages. These envelopes were thick- Yeah ... every month, 'cause everything was itemized, line items, and then the bottom line was, like, always too high. And we're, we're, from a procurement perspective, we're very much in that exact same era for AI today, is we're getting very, very complex bills, very utility-oriented- That's true, yeah

high consumption bills. Bottom line isn't what we like or expect, and we have to go through this maturity of AI FinOps to be able to understand, measure, and then engineer these bills down, and we're, we're definitely not there yet, and there needs to be a focus. The other thing I'd emphasize is that a lot of technology leaders kind of have a, an, or a legacy cloud.

You know, I never thought I'd say those two words in the same sentence, but like a legacy cloud mentality to AI clouds, where you can provision a big CPU server and big CPU instance, and then optionally add more memory to it and optionally add storage and so forth. That is not how AI is built and not how it works.

The performance requirements are so high, the latency tolerances are so low- 

[00:33:01] Serena Huang: Right ... 

[00:33:01] Val Bercovici: that AI has to, for example, package GPUs and GPU memory, the high bandwidth memory, in the same package. 

[00:33:08] Serena Huang: Right. 

[00:33:09] Val Bercovici: You can only provision it together, which means that, again, if you're training and inferring, you almost don't want the same GPU, right?

You want one kind of GPU for training, because it comes with certain memory, but you want much more memory and less GPU for inference, and you can't even buy and deploy systems yet that way, and that's what's coming. That's where the abundance from WEKA is coming from, is we're really the first to disaggregate the compute from the memory for inference, for example, which is more of a mental model match for technical leaders in the cloud era, and that's happening finally.

You know, I think 2026, we'll see it next year even more so. We'll finally get back to that familiar concept even for AI data centers. 

[00:33:52] Serena Huang: Yes. And, uh, I, I, I think there are lots of companies who are maybe starting to become aware that will be ready, and then others that will be very unpleasantly surprised. Um, I remember, Val, you've written about this era of, uh, what you call artificially low AI prices right now- 

[00:34:14] Val Bercovici: Yeah

[00:34:14] Serena Huang: where, um, inference providers essentially subsidizing the cost to capture market share, to drive adoption. What happens when that subsidies end? What, what does AI pricing actually look like in 12 to 18 months? 

[00:34:30] Val Bercovici: Yeah. Fans of analogies, right, as we are, I think this is where we apply the Uber analogy, right?

Where the buying market share and getting consumers effectively addicted to a certain way of working, a certain way of just operating with technology, was the, the, the VC subsidized era, the, the, the immediate post-ChatGPT era of AI. But we're seeing very much right now there's... It's not even in, in pricing.

It was really felt in token rate limits. Again, the fact that there's just- Yeah ... not enough raw materials to satisfy the token demand today. Uh, so this notion of surge token pricing definitely became a thing late last year and, and this year, where we're seeing huge prices mar- more market reality prices for final token bills, final token consumption, not just that low unit, you know, uh, unit cost.

And, and we're seeing Providers, for example, famously like Anthropic, just having to artificially impose low strict rate limits on their best customers, their highest paying users, because there's just not enough token supply to, to allocate to the overwhelming demand. So rate limits, even before actual huge bills, were the first wall, the first pain points that people were hitting.

Uh, and, and I think it's being reflected now in, in just even higher pricing to try and not hit rate limits because you're not even gonna consider using, for example, Opus 4.7 fast mode, which has some astronomical multi-hundred dollar per million token output price point that we've never seen before. I think we need to get used to, with Mythos now and new higher quality tiers of models appearing on the market, they will create new ceilings, shattering glass ceilings of pricing tiers, and, and we will start to see just like when Uber had to go public and, you know, had to become a kind of a real company and investors demanded, there's the ROI term again, actual returns on their investments.

We had to see more real market pricing being reflected, which we're all paying as, as rideshare consumers right now- 

[00:36:38] Serena Huang: Yep ... '

[00:36:38] Val Bercovici: cause we like the convenience. Same thing is gonna happen, and we have to figure out can we afford AI? Right. Can we afford one model AI, or do we really have to financially engineer multiple models to the best solution?

[00:36:51] Serena Huang: Wow. So are you saying that If we follow this trend and the, as the subsidies end, this whole AI is cheaper than people thesis is not going to survive? 

[00:37:04] Val Bercovici: Definitely not in the simplest narrative, right? Of like one model being applied for real work, which means in an agent swarm with, you know- 

[00:37:11] Serena Huang: Right ... 

[00:37:11] Val Bercovici: multiple concurrent agents, long context, multiple turns.

Uh, doing that all with one model is already proven to just be broken, right? It's, it's unsustainable, unaffordable. Even some of the best benchmarks in the world right now, like the ARC-AGI benchmark and so forth- 

[00:37:27] Serena Huang: Yeah ... 

[00:37:28] Val Bercovici: uh, are starting to reflect the fact that you have to factor in the actual number of tokens- Mm-hmm

and, you know, what's the cost per solution for the tokens consumed, not just the individual token costs. Artificial intelligence is measuring that now and so forth. So the industry's maturing and realizing that efficiency matters, right? Uh, you know, you can have, you can have revenue, but if you don't have profit, then what's the point, right?

If your, if your costs are higher than your revenues, then it's pointless. So, uh, the industry's maturing, it's segmenting. We're seeing, interestingly enough, as I mentioned with OpenClaw, not just good engineering from a human architecture designer-developer perspective, but the agents themselves have enough autonomy and awareness to engineer some of these segmentations and multiple model integrations themselves- 

[00:38:16] Serena Huang: Right

[00:38:17] Val Bercovici: just by our policy, our English or, you know, natural language policy. So it's, um, it's fun to see this evolution- Right ... but it's becoming an industry. It's becoming mature, it's becoming segmented, it's becoming intricate and, uh, complex for the uninitiated. 

[00:38:32] Serena Huang: Yeah, absolutely. And, and we see companies learning too really quickly, sometimes the hard lessons, right, where they had maybe laid off 40% of their workforce and said AI agents can do it, and then realizing, ooh, not so m- not so fast, and then they bring them back.

There's a public apology and all that. And, and I feel like there is a conversation around the cost realization that actually doesn't make the headlines. So yes, there's a capability gap for sure between AI and some of the agents. We get it, but, but no one's really talking out loud that, "You know, um, actually there's a cost component that really surprised us, um, and we are going to change how we think about the workforce strategy go- going forward."

[00:39:19] Val Bercovici: So that's so important right now. There's almost sort of two realities. One is that zero-sum mentality where, yes, you know, I was expecting to just reduce my cost by replacing humans, and that either hasn't happened because I, you know, the, the quality's not there and so I'm, I'm getting lower revenue, or yes, you know, I figured out that I can get, I can match the quality with AI- Mm-hmm

but it's just, uh, you know, swapping one for the other. There re- really isn't a benefit to doing that. 

[00:39:48] Serena Huang: Yes. 

[00:39:48] Val Bercovici: The, the more interesting conversations actually, and not enough companies are having this yet, but the leaders will kind of create followers, is a positive-sum mentality, where some companies are realizing, you know, AI lets me fundamentally rethink my business.

[00:40:02] Serena Huang: Mm. 

[00:40:03] Val Bercovici: Lets me rethink the tasks that my staff performs. 

[00:40:06] Serena Huang: Yes. 

[00:40:06] Val Bercovici: And the jobs are still there. The jobs are still, you know, get this function, this business function done, get this particular unit of value sold to a customer- Yeah ... get this supply, you know, processed efficiently, but the tasks change and, and the, the- Yes

most progressive companies are actually increasing their revenues, increasing their productivity, not laying off- 

[00:40:27] Serena Huang: Right ... 

[00:40:28] Val Bercovici: but realizing their tasks have changed, not the jobs. 

[00:40:31] Serena Huang: Yeah. Yeah. Indeed, and, um, I, I've been running small focus groups with AI users, and one of the questions I really love asking them is, what are you able to do with AI that you absolutely could not before?

Not how much faster are you doing something, but what could you not have done before at all? And, and the answers are fascinating, and I, I see the creativity and innovation when, when people talk to me about just really recreating a job that they now potentially like a lot better, um, if they can change their workflows, and it's not just about doing the same things I was doing much faster- Yeah

using AI. 

[00:41:12] Val Bercovici: And my favorite example of that, you know, is dashboards. Like, there's- Yeah ... always information we want aggregated and synthesized and summarized- 

[00:41:20] Serena Huang: Yes ... 

[00:41:20] Val Bercovici: that was in that giant in-between zone of useful to me, but not worth hiring a developer for- Right ... explaining to them what I really want done, iterating over weeks or months- 

[00:41:31] Serena Huang: Yes

[00:41:31] Val Bercovici: and then getting, like, months later an expensive dashboard that works. Yeah. This can be done, like, in an hour now. Right. And, uh, so that, that when you co- compound that productivity benefit, that example i- is a big one. There's another kind of dark example to this as well, though, right? Which is what happened with Mythos a month ago- 

[00:41:49] Serena Huang: Mm-hmm

[00:41:49] Val Bercovici: when Anthropic released Project Glasswing, and we saw that this one particular class, this new class of super powerful model is able to be creative and chain together a bunch of attack sequences that were individually defensible- Yeah ... before, but collectively break not all software, but a lot of software.

Yeah. Thousands of new vulnerabilities in the past month alone, and some very mature, hardened software have been discovered. So that's just another example of, yes, you've got to reimagine both the positive/negative, the tool use and the weapon use of this technology because, um, it is possible, and it will be exploited for benefit or, or for, or for loss.

Yes. 

[00:42:30] Serena Huang: For good or for evil, right? 

[00:42:31] Val Bercovici: Exactly. Exactly. 

[00:42:33] Serena Huang: Yeah. And, and I felt that dashboard impact you described personally. I remember, and this is a while ago, but I wanted to try creating a dashboard from scratch, and this used to take my whole analytics team multiple sprints, right? And I need front end desi- you know, uh, developers, I need designers, I need QA testers, the, the whole, the whole team.

And now I was able to do it within, well, under an hour for sure. Yeah. And then, Val, I was very skeptical. I was like, "This is not going to work. It's not going to, y- you know, function. The filters will break for sure. It can't export anything." No. N- nothing. It was one prompt, and everything actually worked, um, and was quite beautiful, too.

So I, um, I, I think those moments of, um, genuine awe, like Sid- Sikh bathing no chair awe, is what, what a lot of innovation will come from. Um, and at the same time, a little bit scary for people who might be worried about, what does this mean for my job? 

[00:43:34] Val Bercovici: Yeah. 

[00:43:34] Serena Huang: Um, I, I'm curious if you have any encouraging words that you might say to someone who might see the power of AI but also worry about their future?

[00:43:45] Val Bercovici: Yeah. You know, I, I don't wanna be too one-sided here, but it's been quite positive, you know, on, on the WEKA front, where as chief AI officer, you know, I had these two roles, inbound and outbound. From like- Yes ... you know, uh, an outbound perspective, it was all about evangelizing the benefits of AI, and certainly working on our product strategy and so forth.

From an inbound perspective a year ago, there's a lot of convincing people to just try these AI tools. 

[00:44:10] Serena Huang: Yes. 

[00:44:10] Val Bercovici: Uh, and then, you know, once developers finally came around, knowledge workers and so forth. What we're seeing today is, you know, we have these amazing designers that are not out of work. They're now helping everyone at the company be great designers, and they're busier than ever.

Literally, we need to hire more because people are seeing the benefits, the productivity benefits, particularly of dashboards. And here's, like, the group benefit we're seeing is, as you know, um, uh, as we talked about earlier on, memory inventories, actual, you know, DRAM, dynamic RAM memory, and high bandwidth memory inventories are in very short supply.

Storage, you know, flash drives, NVMe devices are in very short supply. Many of the dashboards we created for our own, like, you know, procurement people and our own sort of service management people helped us make better decisions, and we are now in this weird enviable position of not being s- you know, short on supply, and actually having inventory of one of the hottest commodities in the industry today, which is, you know, storage and storage that works as memory.

So there have been individual and, and company-wide benefits towards, you know, reali- realizing the tasks improve, right? And we're able to do a lot more and things we couldn't imagine, and, and the c- the compounding benefits, we're seeing them in, in three to six months in our case. 

[00:45:30] Serena Huang: Yeah. Um, and, and I like to focus on thinking about AI in, in different levels, the individual, the team organization, um, and then also your constant focus on let's think about outcomes.

Let's talk about outcomes, measure those instead of tokens and usage. Um, and I think that's going to be really key for leaders who are considering the next chapter. I think the leaders who are going to win are, are not the ones with necessarily the cheapest AI or the most tokens somehow, but they are the ones who actually know what they want out of AI, and they're measuring the outcomes associated with that rather than tokens, and then be surprised by a price.

So yeah, definitely if you're building your strategy around people, uh, make sure your AI cost assumptions are correct and, and not outdated. We talked so much about that today, um, because that will get repriced, as Val said, when the subsidies come to an end. It's time to revisit, not time to panic. Um, yeah.

Um, Val, thank you so much for today's conversation. Any last words that you want to leave with us? 

[00:46:50] Val Bercovici: Maybe just a, a, a fun sort of hidden narrative that's rising to the fore right now, which is what I call the revenge of middle management. Ooh. So, you know, middle, middle management is always like, you know, the, like the, the whipping boy, if you will.

It's always something that is first to get cut, and leaders will proudly say, "We eliminated waste by that inefficient middle management layer." But then you take a look at where the productivity is coming from in AI, it's from agents. What do the best agents do? They organize in swarms. They actually literally form an org chart.

[00:47:19] Serena Huang: Yeah. 

[00:47:20] Val Bercovici: They take ambitious goals. Hmm. They understand how to actually split and divide the goals, and delegate subtasks to accomplish a goal, how to monitor and measure the subtasks for quality, for efficiency, how to synthesize that all back up into the final business outcome or requirement. That is a classic MBA middle management skill, and skills are like the hottest things to sort of- Yes

train AI agents with. So I'm not just saying there's gonna be this pendulum swing of everyone hiring a whole, a whole, you know, army of middle managers again. But middle managers have to realize that if you're on the job market all of a sudden, your skill is, is the hottest skill necessary for AI to work in business.

Yeah. It's being able to program an agent, if you will. It's being able to set policies, goals, understand how to measure, how to delegate, and h- how to summarize. And, uh, so that's, that's kind of one of my many hot takes is that, uh, there will be a revenge of the middle managers because it's middle management skills that make agents work, which makes AI give the ROI businesses expect.

[00:48:24] Serena Huang: I love that. It's so encouraging and such needed message today as people are feeling the anxiety everywhere. Thank you, Val, so much for joining us today. It's always a pleasure, and I look forward to our next conversations. 

[00:48:40] Val Bercovici: Can't wait. Always enjoy these. 

[00:48:42] Serena Huang: Thanks for listening to Deep Geeks. A huge thank you to my guest today, Val Bercovici.

If today's episode made you think differently about how AI gets built or powered, share it with someone who needs to hear it. Find Deep Geeks on Spotify, YouTube, or wherever you get your podcasts. Until next time.