The AI boom is often framed as a race to build more data centers, with headlines focusing on megawatts deployed, GPUs installed, and square footage under construction. That framing, however, misses the more important shift. The real transformation surrounding AI is happening inside the systems that power it.
AI workloads are fundamentally different from traditional cloud workloads, and as they scale, they are forcing a redesign of the infrastructure stack. That shift is changing how these systems are built and what it takes to build them. To understand why, it’s important to look at where traditional data center architectures are beginning to break.
Why AI Is Breaking the Data Center Stack
Modern data centers were designed for distributed, loosely coupled workloads. These workloads could be spread across machines, scaled incrementally, and tolerated latency between components. Data centers were built for flexibility and distribution, while AI workloads demand coordination and intensity. This mismatch is where the system begins to break.
Power delivery, cooling, compute architecture, interconnect, and networking are no longer independent layers that can be optimized in isolation. A constraint in one immediately impacts the others. For example, increasing compute density may not be limited by chip design, but by the ability to deliver power or remove heat.
Recent advances in model efficiency further reinforce this dynamic. Techniques like TurboQuant significantly reduce the compute and memory requirements needed to run large models, improving performance per watt at the chip level. But rather than removing system constraints, these gains shift where pressure shows up. As inference becomes cheaper and more widely deployed, overall demand increases, driving higher utilization across infrastructure. This amplifies bottlenecks in power delivery, cooling, and data movement, reinforcing the need to design across the full system rather than optimizing individual components in isolation.
As compute becomes more concentrated and runs continuously at high utilization, the system strains. Power delivery becomes more difficult, and the heat generated by that density forces cooling to become a primary constraint on how systems operate.
These pressures reshape infrastructure design. Power and cooling are no longer individual considerations. They directly influence how compute can be deployed, how racks are structured, and how tightly systems can be integrated.
As systems become more tightly packed, data movement becomes a second constraint. At smaller scales, this was manageable. At AI scale, bandwidth and latency begin to define overall performance and push existing interconnect technologies to their limits.
At the largest scale, coordination even becomes a bottleneck. Training large models requires thousands of GPUs to function as a single system, and traditional networking architectures struggle to handle that level of synchronization efficiently.
To understand where these constraints emerge, it helps to look at the system as a whole.

Pressure builds across each of these layers as AI systems scale, and what begins as a shift in workload quickly becomes a system-wide constraint. Pressure in one part of the stack impacts the others and turns previously manageable limitations into defining bottlenecks.
Founders Solving the Hard Problems
The most interesting solutions are coming from founders working directly on these pain points. As AI systems continue scaling, teams building across power infrastructure, cooling, interconnect, and networking are encountering the limits of current architectures in real time.
We asked some of our portfolio company founders for their perspectives on how these systems are breaking and how these systems need to evolve as demand increases.
Frore Systems
Seshu Madhavapeddy (Founder)
Frore Systems’ advanced cooling technologies enable higher-density compute by addressing thermal limits at the chip and system level.
Question: How does cooling influence the way AI chips and systems get designed in the first place?
Answer: “Cooling fundamentally defines the boundaries of AI system design. Every AI GPU is thermally limited, which means GPU performance and rack density are shaped not just by compute and networking capability, but by how efficiently heat can be extracted. As power densities increase, cooling has become a primary design enabler that influences chip layout, compute tray architecture, and rack-level design from the outset.
This is why we talk about the Thermal Stack as a foundational system layer of AI factories. Liquid cooling innovations from Frore Systems, like LiquidJet coldplate that delivers 75% higher heat removal efficiency and LiquidJet Nexus that simplifies liquid cooling from a collection of discrete components in the compute tray to an integrated architecture, enable higher compute density, more efficient system design, and higher performance that are not possible with traditional thermal solutions.”
Lumilens
Ankur Singla (Founder)
Lumilens enables high-speed data movement across AI systems through optical technologies built for next-generation bandwidth and latency requirements.
Question: What is changing in AI system architecture that makes optical interconnect essential?
Answer: “AI system architectures are rapidly shifting toward massively scaled, distributed GPU clusters, driven by large model training and an ever-increasing volume of real-time inference that requires low latency and high bandwidth between nodes. As AI clusters grow to tens or hundreds of thousands of GPUs, and connectivity speeds continue increasing from 800G to 1.6T and beyond, traditional electrical interconnects hit fundamental limits in bandwidth density, power consumption, reach, and signal integrity.
Simply put, electrical networking cannot efficiently or cost-effectively support the throughput, density, or energy efficiency required for massive AI compute clusters, especially those with disaggregated and/or scale-up architectures. Photonic (advanced optical) interconnects overcome these legacy constraints by enabling much higher bandwidth over longer distances with lower power per bit, making them essential for building next-generation AI infrastructure stacks.”
Upscale AI
Barun Kar (Founder & CEO)
Upscale AI designs networking architectures optimized for large-scale AI clusters.
Question: What breaks first in traditional data center networks when running large AI clusters?
Answer: “What breaks first in traditional data center networks is the assumption that traffic is loosely coupled and can be distributed efficiently across the fabric.
Large AI clusters generate tightly synchronized east-west traffic and collective communication patterns that traditional networks were not designed to handle efficiently at scale. The first visible symptom is congestion-driven tail latency: incast and hot spots slow collective operations, and in synchronization-bound workloads, the slowest path holds back the job. The result is that the network becomes the bottleneck, limiting cluster performance and the economics of scaling.”
Across each of these perspectives, the pattern remains consistent. AI exposes the limits of the system as a whole and requires deep technical expertise to address. Building in this environment requires a different kind of founder and a different kind of support.
Backing the Builders of the AI Infrastructure Stack
The challenges across the AI infrastructure stack are deeply technical, highly interdependent, and often unforgiving of partial solutions. That changes what it takes to build in this category. Execution requires precision across all levels of the system, and mistakes compound quickly.
As a result, the profile of founders building in this space looks different. They aren’t approaching these problems from first principles alone. They bring direct experience that changes how they build. Solving these problems requires coordination across hardware, software, and systems engineering. Teams must be able to navigate long development cycles, complex supply chains, and high-stakes customer environments. The most successful organizations are those that have worked together before and can operate with a high degree of integration from the start.
The broader ecosystem reinforces this dynamic. Customers are making infrastructure decisions that define their competitive position. In these environments, reliability and trust matter just as much as innovation. Access to critical partners is earned through credibility and a proven ability to execute, making this the point where alignment between founders and investors becomes critical.
At MVP Ventures, we consistently back teams building at this level of complexity. Our team invests in early-stage companies at the intersection of AI, hardware, and software. MVP’s role is to help founders operate across the full system they are building, from product development to go-to-market strategy, talent, leadership hiring, capital markets, government relations, and the broader set of relationships required to scale.
The AI Infrastructure Cycle Ahead
AI is creating a new class of infrastructure problems that are harder, more capital-intensive, and less forgiving than what venture has traditionally supported. Progress is driven by the ability to execute across long timelines, complex systems, and tightly coupled constraints.
Founders developing technologies in this space face an environment where every aspect - from technical decisions to supply chains and customer relationships - is tightly linked. They need both deep technical expertise and the right support.
At MVP, we partner with founders building in these conditions and help them execute across the full system required to scale. The next phase of AI will be defined by those who can do that consistently.
Want to learn more about how we back founders early and deliver after the check clears?
Explore what makes partnering with MVP a no-brainer: mvp-vc.com | LinkedIn | X