Can AI innovation outpace its own energy demands? Dr. Serena Huang explores the real cost of scaling AI with Daria Mukhortova, Head of Sustainability at Nebius, and Val Bercovici, Chief AI Officer at WEKA.
As AI scales, energy demands are growing 100x and efficiency is no longer optional. Dr. Serena Huang sits down with Daria Mukhortova, Head of Sustainability at Nebius, and Val Bercovici, Chief AI Officer at WEKA, to examine whether innovation can outpace its own energy appetite. They explore how efficiency at the infrastructure and software layer is becoming a business imperative and what organizations should demand of their AI providers.
Timestamps:
(1:16) Meet Daria Mukhortova & Val Bercovici
(4:14) Inside Nebius's Helsinki data center & heat recovery
(6:36) What energy efficiency unlocks for AI customers
(8:59) The metrics that actually matter
(12:30) Goodput vs. throughput: are you measuring the right thing?
(14:37) Software's hidden role in AI sustainability
(19:23) How much energy does AI actually consume?
(21:37) Common misconceptions about AI's environmental impact
(26:31) How to choose an AI infrastructure provider
Links:
[00:00:00] Daria Mukhortova: The energy is then translated into AI outputs, but it doesn't tell you the full story, how much energy per workload has been used or how much energy per token has been used.
[00:00:12] Val Bercovici: These resource demands are so intense right now that every ounce of inefficiency is not just bad for the environment, it's bad for business, and it restricts the agility of these companies
[00:00:29] Serena Huang: Welcome to Deep Geeks. I'm Dr. Serena Huang. Today, we're going deep on a question that keeps a lot of people in our industry up at night, and it's not about models or data or compute in a way you might expect. Can AI innovation outpace its own energy demands, or are we building towards a power crisis? I have two incredible guests.
First, Dasha Mukhortova, Head of Sustainability at Nebius, and I'm also joined by Val Bercovici, Chief AI Officer at WEKA.
[00:01:04] Daria Mukhortova: Thank you, Serena. I'm very happy to be here. Thank you for having me.
[00:01:08] Val Bercovici: Pleasure, Serena. Looking forward to this.
[00:01:09] Serena Huang: Dasha, let's start with you. Tell us a little bit about your work at Nebius and why sustainability sits at the center of an AI infrastructure company.
[00:01:19] Daria Mukhortova: My key task fro-from the very beginning was to ensure that we don't just try figure out what Nebius is and what it builds, but also how. This actually translates into a few principles that myself and the team we agreed on early in the day. First is that we don't treat sustainability as something that comes after, but we treat it as a principle.
For us, sustainability is basically a synonym to efficiency, a synonym of reliability, which makes it pretty much understandable of how to treat it as an engineering approach. So that's, uh, pretty much summarizes, I guess, like, the core things that I focus on internally, uh, and, uh, looking forward to the discussion.
[00:02:04] Serena Huang: Likewise, can't wait to dive in, such an interesting perspective. And Val, you are approaching this from a different angle. As chief AI officer at WEKA, and given your background building the foundation layers of cloud and storage infrastructure, what does energy efficiency mean to you when you are thinking about how AI systems are actually built?
[00:02:26] Val Bercovici: Storage for accelerated compute for AI, as Jensen calls it, uh, is very different than storage for traditional computing. It's this very deep geek technical challenge of performance efficiency, not just capacity efficiency. And performance efficiency is very sophisticated engineering for that high performance.
You can say that general performance equals revenue, performance efficiency equals profit, and that's why there's, you know, as Dasha was saying, really great aligned incentives here. It's not just that, you know, we're doing this for altruistic sense. It really aligns with us as a business trying to help our customers extract maximum profits out of these, uh, you know, unprecedented capital and, and operational expenditures they have.
[00:03:12] Serena Huang: I am blown away by how these two perspectives are now actually converging, the sustainability lens and the infrastructure lens, because for a long time, I remember them as very separate conversations, and now we see how energy efficiency isn't just a check the box exercise for sustainability, it's real engineering constraint.
So let's dive into how companies like Nebius are actually solving for it Dasha, let me start with something concrete here. Nebius built the data center infrastructure in Helsinki, and that wasn't an accident. Could you please walk us through the thinking behind that? And specifically, I'm curious about the heat recovery piece, because the idea of recycling waste heat into usable energy feels like it flips this whole conversation.
[00:04:07] Daria Mukhortova: Yes, indeed. Uh, our data center in Finland, it's, uh, like 30 minutes away from Helsinki to get a bit to the south. It's, uh, what we internally call as our playground for all the specific innovations that then, then we try to roll out across other sites. The first one, it's first data center that we, we built.
Um, and, uh, what's interesting about that specific site is that, uh, is how it's engineered to be not just a consumer of energy, but also a contributor, uh, to the energy system. And, uh, the specific example that you mentioned about heat recovery, it's inbuilt in the cooling system cycle. For us, this is actually win-win across many fields, and I can walk you through the, actually the benefits of having the system.
So first of all, it's a great economic contribution to the local community because by using, by reusing server heat that we capture, which is basically free, you actually are able to reduce the cost of producing heat for the municipality And in, uh, last year, in 2025, uh, the households, they actually paid 10% less, they spent 10% less on heating because they were able the, the, to leverage the free server heats as a, as a resource.
And if we talk about numbers here, and if we treat heat as a byproduct of electricity consumption, so that would be, uh, recovering around 20, 30% on an annual basis and given back to the local energy system. And that's, I think, is key that when we think about infrastructure as an interplay of solutions that can be efficient on their own, but also connected with the local energy system, and also find ways to power that system back.
[00:05:57] Serena Huang: I am going to keep that story in my back pocket though, because I purposely spend a lot of time with people who are AI skeptics, I would call. And one of the reasons people say, "I don't want to touch AI," is because of energy consumption. And you just illustrated a very different way of approaching AI that can actually translate into greater good for the whole community, so thank you for sharing that.
And Val, from your vantage point, when customers come to WEKA, what does this energy efficiency Dasa just described actually unlock for them? What becomes possible when infrastructure is built this way?
[00:06:40] Val Bercovici: If you're striving for, again, these two forms of efficiency in the storage world, capacity efficiency and performance efficiency, it's a direct impact on your bottom line and your CapEx.
Particularly since late, uh, 2025, beginning of 2026, we are now seeing the rise of agents. And agents are, again, another, you know, if we go from chat to reasoning, that's an order of magnitude more energy consumption and efficient performance requirements for efficiency. As we move from reasoning to agents, it's another 10x, meaning 100x more, you know, intensive workloads, compute workloads than for chats.
And so the performance efficiency benefits I'm talking about get amplified 100x between WEKA and Nebius right now. It really puts a, a very big scope, a very big lens on how efficient you're being, because these resource demands are so intense right now that every ounce of inefficiency is not just bad for the environment, it's bad for business, uh, and it, it restricts the agility of these companies.
You're seeing some of the, the leading minds, the leading voices, the CEOs of OpenAI and Anthropic publicly talk about- Balancing cash flows to survive as these very high pace startups, these unprecedented sort of scale startups, uh, and performance efficiency and overall efficiency is a big part of that.
[00:08:01] Serena Huang: So now we're seeing that efficiency at the infrastructure layer is not just about saving money or electricity bills, it's about what you can do with AI. And I see that organizations that crack this, they're going to be able to run more experiments, train more models, serve more inference, which brings me to the next question.
So Val, help me out here. What should we be thinking about? What should we be measuring? And I've heard terms like microwatts per token, PUE, but help us understand the benchmarks that actually matter.
[00:08:37] Val Bercovici: I'll start with the transparent one. Now, maybe one of the things to remind people is this is a very transparent industry, particularly on the inference side.
So pricing, token pricing is a very public metric, very, very competitive.
[00:08:49] Daria Mukhortova: Yes.
[00:08:50] Val Bercovici: And, you know, Nebius and us do a lot of work on optimizing token pricing right now. Uh, but I think with regards to this discussion, I'll let Dasha focus on PUE, because I think that's one of her specialties. We can focus on tokens per watt.
You know, very often when it comes to performance, we try and measure tokens per second, sometimes tokens per GPU, and the GPU now, the LPUs, and other kinds of accelerator types are diversifying, but tokens per watt is a really, really key metric. And inefficient inference, which is really where we're at today, you know, think of, of Amazon.
Amazon isn't legendary for their factories. Amazon is legendary for their warehouses and their delivery logistics. The way inference works right now, being so nascent, is that there are no warehouses in the inferencing reality of delivering tokens to users. There's just factories with lots of inefficiencies.
So introducing a concept of token warehouses, uh, lets you not waste factory output, lets you optimize and scale the delivery of tokens to your users. It really shows up very much in the actual tokens per watt. It shows up so much so that inefficient implementations end up consuming as much as any individual household would use in a single day per chat session or per agent session.
[00:10:06] Daria Mukhortova: Wow.
[00:10:06] Val Bercovici: Whereas, yeah, exactly. Whereas efficient, you know, uh, warehousing of tokens, if you will, and inefficient memory for tokens, not just storage, ends up reducing that by 80% or more, you know, is an active project we have with Nebius right now.
[00:10:21] Serena Huang: Wow, that's a huge difference. Dasha, do you want to chime in there on some of the
[00:10:26] Daria Mukhortova: metrics?
Yeah, so I, I guess I can definitely cover the PUE part. PUE basically is a metric that men- that, that measures how efficient energy is delivered to IT equipment, right? Um, however, I would make this note that power usage effectiveness is great on its own to see how you design the physical infrastructure that's built around the specific GPUs.
However, it doesn't tell you how efficiently the energy is then translated into AI outputs. So from my perspective- Right ... it's very important to, uh, have a broader look at, uh, how we can measure efficiency, and what Val was talking about is should be definitely part of the picture. And what I see also in the industry, there is a certain, uh, there is a good understanding what PUE is, and hence people tend to also compare different sites and different providers based on PUE, but it doesn't tell you the full story.
So from my perspective, uh, it's important to also, um, introduce the measurement of, let's say, uh, how much energy per workload has been used, or how much energy per token has been used. And another, like, a level of depth here would be how much of useful outputs we also differentiate between, I guess, goodput and throughput here has been produced per megawatts or per, per watts of energy that entered the system.
And this brings me to, uh, one thought about how we can treat the infrastructure. And while the discussion is largely focused on hardware, like what the chips, how the chips are performing, how the servers are performing, what the cooling systems are, how they are contributing to that, the big question is about the software, because software plays an important role in, um, orchestrating all of that, all of that, uh, hardware and infrastructure, and allocating the workloads in a way that leverages all the available piece of fabric leads to servers not, uh, staying idle.
Uh, idling is actually a killer to efficiency. I would set this discussion as something that needs to take into account all the pieces and all the layers that infrastructure is composed of, starting from how you build a chip, how you build the enclosure of the chip, onto what your software is doing, how it's performing, how it's orchestrating the workload.
[00:12:57] Val Bercovici: And so goodput literally shows that you can have a very busy system, right? We, we talk about GPU efficiency a lot because they're such expensive assets and they're so expensive to operate. You can have a very busy system that's not really producing a lot of useful output or producing at a very, very slow rate, and that's if you measure throughput or just general utilization.
But if you really focus on goodput, it's what is the actual useful output, what is the actual, you know, performance level, efficiency level that you want, and you want that utilization to focus on that versus inefficiently being busy, you know, and, uh, and focusing on and delivering inefficient output that's not adequate.
So the actual Final tokens per second is something that's more important, more relevant to measure than just general utilization of the infrastructure to generate that output. And it differentiates and, and elevates the conversation from just throughput, for example.
[00:13:53] Serena Huang: I think we keep hearing about hardware.
Frankly, it's all we hear about most of the time, but we forget about the software. Dasha, is there anything else you want to add on software optimization when it comes to sustainability decisions? Yeah, sure.
[00:14:10] Daria Mukhortova: And from my perspective, I can also tell that this is something that Nebius, uh, decided to invest in from the very beginning, like to have the software layer on top of the, of the, the hardware stack, because we see efficiency gains at that level, and I can provide a couple of examples of how this can work.
Well, software layer is actually responsible for... can be responsible for tracking certain failures in your workloads in the nodes. Your in- it's in your interest, uh, to actually fix this seamlessly, so the workload doesn't start to retrain, for example. Because retraining, it means that actually drawing, uh, twice as much resources.
Autoscaling is something that is very necessary to ensure that the bursty consumption that the AI workloads are, like, normally characterized with, right? This bursty consumption would be met with the right-sized cluster. That would mean to, uh, ensure that the servers or, like, GPUs are not overprovisioned, meaning that they're not staying idle.
And idling can actually draw quite a lot of power, especially when we talk about GPUs. Uh, so it's, uh, always the question about, uh, leveraging as much of the, of the available capacity to meet the workload versus just provisioning it, overprovisioning it, like, let's say, to one specific client, and then locking it for any useful outputs that could be there.
So that's what autoscaling, uh, at the software layer is also solving. And there are many examples like that, and I would actually invite also Val to contribute with his knowledge of the storage systems because what I also see that storage, it has to be AI, um, like tailored. There are several types of wo- o-o-of data that needs to be stored.
It could be the question of having active storage or cold storage, and, uh, it depending on what your workload actually needs. If it's a training data, it has to be going to the active storage, meaning that you need to have access to it faster, meaning fewer bottlenecks is actually less energy being drawn.
And then if you are talking about some historical logs, it's, it's best to allocate this to another type of storage that actually consumes zero, uh, close to zero power when, when idling, right? So, um, this type of questions about how we actually manage and orchestrate the workload when already running can bring us to significant efficiency gains on top of what infrastructure can provide in terms of cooling savings and drawing less power per, like, per run.
[00:16:49] Val Bercovici: So tiering memory is probably the hottest thing in AI right now for inference to respond to the agent demand, and the ability to provide storage capacities with memory performance to, you know, leading offerings like the Nebius Token Factory is probably one of the most exciting things that we're working on together.
So one of the reasons why memory tiering is so important is that today it's a general best practice that if you wanna scale inference, the, the memory that traditionally comes with GPUs is very tightly coupled to the GPUs.
[00:17:23] Daria Mukhortova: Mm-hmm.
[00:17:23] Val Bercovici: So as you need more and more memory for inference, which is really inference is a memory problem, you have to overprovision GPUs just for the memory they bring along for the ride.
[00:17:33] Serena Huang: Mm.
[00:17:33] Val Bercovici: And during inference, those GPUs are largely idle, so it's really a waste of the capital resources, and it's inefficient energy-wise as well. If you decouple for the first time in the AI era, the compute from the memory, the GPUs from the GPU memory tiers, you can rebalance the system and only provision the memory needed for inference without overprovisioning those idle, wasteful GPUs at inference time.
And that's the magic here is the software combined with the hardware to balance a system, weed out all the inefficiencies, and yield that good put that we're all chasing.
[00:18:09] Serena Huang: For most organizations, even very sophisticated ones, what happens between plugging in power and getting tokens out, that process is completely opaque, and that's actually very expensive.
So we're going to open this up Let's contextualize, uh, the scale here because I think the numbers are shocking to people who haven't looked at this closely. Val, what is the intelligence on how much energy AI actually consumes from training to inference to just keeping the lights on at an AI data center?
[00:18:49] Val Bercovici: We're now in the era of GPUs, which instead of thousands of cores per rack, have millions of cores per rack. So it's a fundamentally different level of parallelization. And, uh, not just having to understand the engineering for how to make something operate a million times in parallel versus 1,000 times in parallel, that's already a very steep engineering challenge from a compute networking, you know, storage software perspective.
But the energy consumption for a rack of GPUs is hundreds of kilowatts, and we're now in the era of the latest generation processors. I think by the time we'll be, we'll be airing this, the latest generation of GPUs will be consuming up to a megawatt per rack.
[00:19:36] Serena Huang: Wow.
[00:19:36] Val Bercovici: Which is kind of, you know, in, in historical context, almost an insane amount of energy consumption.
And I think there's a real reason why many of the big announcements between the frontier labs that get a lot of headlines and the major GPU suppliers like Nvidia and AMD and Google and others now, the announcements aren't in, you know, aren't in like a performance metric or even a dollar value. Mm-hmm.
A lot of these deals are announced at how many gigawatts of capacity they've agreed to provide each other. And, uh, and so that's just a really interesting, you know, evolution of how we discuss an industry is instead of going from a compute metric or just an economic metric like dollars, we're already talking about gigawatts now, multiple gigawatts of energy for these deployments, largely for inference now.
[00:20:24] Serena Huang: And it has evolved so quickly. This is really innovation and scale. We couldn't imagine this probably even three years ago, right?
[00:20:33] Val Bercovici: A gigawatt is typically what one nuclear power plant generates .
[00:20:36] Serena Huang: Right .
[00:20:37] Val Bercovici: So we're talking about multiple- Multiple ... nuclear power plants now dedicated just to one or more large scale AI data centers, AI factories, and, and with efficiencies, AI warehouses, so can warehouses as well.
[00:20:49] Serena Huang: Yeah. Incredible. Well, Dasha, in your work with organizations on sustainability, I'm curious what you've heard as the most common misconceptions that you run into.
[00:21:02] Daria Mukhortova: Well, first of all, uh, now apart from the discussion around energy consumption, there is also this big focus on what happens with water resources.
And I guess the first, um, like, the intuitive reaction of many people would be assuming that data centers and this infrastructure consumes a lot of water, and it can be the case. However, the question is: How is the cooling system designed? So this is the question I think that the industry should be asking first before making jumping into assumption that, uh, one side or another consumes a lot of water.
And providing an example of Nebius here, we actually do not rely on water intake, even though we're now introducing the liquid cooling system, and it's because the system that, uh, we are building is closed loop. It doesn't include any evaporative components, and we actually use the outside air, uh, through the dry coolers to get the temperature of the fluid that circulates within the same loop, uh, like, actually thousands, millions of cycles to get the temperature down.
Also, another question is why was this design possible? Like, if going deeper into understanding the technology and linking back to, like, seeing this as a system as a whole, because in our case, the reason why we were designing the system this way is also because our servers operate at higher temperatures, so they're specifically designed to be resilient and to be high performance under temperatures of up to forty-five degrees.
So it means that we don't need as much power and as much of cooling. So that's one thing about this water consumption. When, uh, thinking about the energy footprint of the AI workload, and we discussed it, uh, uh, actually quite a few times already with you, that the first, uh, thought is about the hardware and the physical infrastructure.
However, what gets overlooked is also how the model itself is being designed because the design of the model, the setup of the model actually defines, to a big extent, how efficiently it will operate on a given hardware. Actually, at Nebius, we were also thinking about that, uh, and the Token Factory, uh, actually recently introduced the post-training optimization tool, which is, uh, linking us back to the discussion about the importance of software.
And what this tool does, it basically, like in a very, in a nutshell, it tunes and tweaks the setup, the post-training setup of the model so it can be more efficiently running on the hardware that, uh, hosts it, meaning that the, uh, post-training will be completed faster without failures, and it means that it will draw less power.
So this, um, I would say overlooking the, the model set up when talking about the footprint, it also another misconception that, uh, I face a lot in the discussions And probably the third one that I would mention, um, is, uh, indeed that there is a big, uh, discussion now, and it gets a lot of traction in the news, the adding energy built, like capacity.
So there's a lot of investment in the area, like, uh, literally building some facility energy that would be linked to your data center supplying power. And the question that I would ask is, how much of this new build will actually be translated into useful computes? Or is it adding and then losing on overheads?
So this is the question that I think will, uh, become even more important than, uh, the discussion about how much capacity company A or facility A needs to add to meet the demands.
[00:24:55] Serena Huang: We all hear about how expensive, how, uh, energy-consuming data centers are, but there's a lot happening inside the black box than most people realize.
[00:25:07] Val Bercovici: And the other point of emphasis, we just can't repeat it enough in this conversation, is that systems efficiency of hardware and software Because one of the transparency metrics hopefully that will become more prominent throughout 2026 is are you using, you know, proper memory technologies, augmented memory technologies?
Are you warehousing tokens or dropping them on the floor as you're pumping them out from your AI factories? Because in doing that, it's proven now that you get anywhere from 75 to 90% efficiencies by balancing systems with the appropriate level of memory, uh, versus not doing it in the nascent era of inference.
[00:25:46] Daria Mukhortova: What advice we could give to clients that are the customers that are choosing the provider, AI infrastructure provider, I would say that first question that needs to be... Like, first every infrastructure provider has to be challenged in how they're building their stack. I understand that, of course, customers would largely be driven by such things as cost efficiency, like the price of the offering, as well as, of course, reliability.
But also, uh, another question is: How is that achieved? And not every, uh, every company would be the same here. So like in our case, in case of Nebius, the price and the affordability is managed through the efficiencies achieved throughout the stack, which gives us this freedom of actually setting the price that would be more favorable for the clients, not because we are just doing it, but, but also because we can optimize our cost.
Since our servers, they actually draw 20% less power than any off-the-shelf solutions because we design this in-house. This also results in a very tangible cost saving for the team, right? Then it translates into savings at the cooling system level, at the PUE level that we discussed, right? And it's adds on top of that, and I can even provide you an example of Finland since I have these numbers already.
So last year, uh, because of hardware efficiency, meaning server efficiency and cooling system efficiency, we were able to avoid 50 gigawatt hour of, of electricity use, and using that power, we can actually run our Paris site for, I guess, five to seven months, 24/7. So, and like this is a lot about cost saving that then translates into cost efficiency for clients.
And I think actually this will become a growing interest from, from many clients from enterprises trying to understand how the workload that they run with you, what kind of footprint, energy footprint, con- car- carbon footprint it actually provides. As a provider, you should be able to do that for the clients, and if you're not able to do that, that should be also something to probably discuss and consider.
[00:28:01] Serena Huang: Very helpful and very practical. As we close out today, I'm curious if you can leave us with some parting thoughts to our listeners. What is one thing you want them to remember? What is one mindset shift? What is one key message?
[00:28:18] Val Bercovici: Yeah. I think your message is probably even more for providers and for consumers is that this is such a competitive industry, and it's literally at the forefront of science and technology and engineering, that the market forces will drive efficiency as part of overall competitive offerings.
You know, you really won't be able to compete, particularly on token pricing, which is, uh, which is almost everything in the inference industry, if you don't have an efficient system, hardware, software, et cetera, environmentals. So I'm actually v- generally optimistic, right? There's a easy... It's easy to be a doomer in this industry, uh, but when you take a look at just the, the realities of how you build systems to compete, efficiency is not optional.
It was optional in the past for IT systems. It was nice to have. It's just on a critical path. It's fundamental and essential towards having a competitive offering in AI training, and particularly inference markets.
[00:29:10] Daria Mukhortova: From my perspective, I think we all agree that AI brings benefits. It can accelerate like drug discovery, research on diseases, and many, many other like, uh, critical fields can benefit from AI.
So it means the use of energy actually wor- it, it makes AI worth use of energy Uh, and what will define the future of AI is how well we engineer it. And if we think about this as a ecosystem of different decisions made at different levels with different players, from chip, uh, from chip providers to infrastructure providers and how they design the system onto software and how different tools are being built, this is...
And if we all work together, this is how I think we can achieve a sustainable AI that it produces maximum value, but also is mindful and conscious about the resources that it's using.
[00:30:15] Serena Huang: I feel so inspired and just encouraged by that. Thanks for listening to Deep Geeks. A huge thank you to our guests today, Dasha and Val, for bringing such depth and honesty to a conversation the whole industry really needs right now.
If our episode today make you think differently about how AI gets built or powered, please share it with someone who needs to hear it. Find Deep Geeks on Spotify, YouTube, or wherever you get your podcasts. Until next time