Architecture & constraints
Memory Wars
How memory movement shapes the economics of inference, and why the right architecture depends on the workload.
BEP research coverage
From silicon architecture to the economics of intelligence.
BEP Research connects engineering, market data, and investment analysis across the infrastructure stack. Our work combines Ben Pouladian’s operating and investment perspective with specialist research contributions. Start with the frameworks, inspect an observation, and follow how the work develops.
Start here
Architecture & constraints
How memory movement shapes the economics of inference, and why the right architecture depends on the workload.
Compute & capital
A framework for following capital through AI infrastructure, useful model output, and the cash flows that finance the next buildout.
Collaborative memory research
A worked example of connecting memory engineering to business economics: which AI data earns a place on NAND flash.
BEP research frameworks
Architecture & constraints
Ben Pouladian
Memory Wars examines how moving and storing data shapes AI infrastructure economics. A faster accelerator does not by itself resolve the capacity, bandwidth, or latency constraints of a particular workload. The research follows how those constraints influence memory tiers, system design, and supplier demand.
The business question is where spending moves when the bottleneck changes. That requires looking at the full system and its workload, rather than treating a chip specification as a financial result.
This is a framework for investigation; its implications depend on the system and workload.
Original December 2025 essay on Substack ↗Compute & capital
Ben Pouladian
The Token Dollar connects the physical buildout of AI to the financial flows supporting it. Capital funds chips, datacenters, and power; infrastructure produces model output; customer revenue must then support operating costs and the next investment cycle.
The framework asks where economic value is captured along that chain, and how financing and dollar-denominated markets may shape it. The thesis is something to test against contracts, pricing, utilization, and cash generation as the market develops.
Collaboration in practice
Ben Pouladian & Rui “Rick” Xie
Which AI bytes actually belong on NAND? The answer depends on access patterns. Read-heavy model weights and frequently updated working state can impose very different demands on a storage or memory tier.
Rick’s engineering and standards analysis and Ben’s investment framing examine those distinctions together. The report considers writes, reuse, endurance, and movement between tiers before drawing conclusions about demand for memory suppliers. It also corrects an earlier HBF timing expectation using a later company roadmap.
A public observation from The Stack
LLMflation compares a weighted basket of model output prices with GPT-4’s March 2023 launch price. Switch between index points and dollars to inspect the same observation.
This is a price comparison with different model compositions. It does not measure equivalent quality or the cost of completing a task successfully.
Explore The Stack ↗Full dashboard · subscriber sign-in requiredLLMflation
As of 2026-09-12
A price-basket comparison, not a measure of equivalent model quality or task success. Weights: OpenAI 30%, Anthropic 25%, Google 20%, DeepSeek 15%, open-weight models 10%. The launch reference is $60 per million output tokens.
How we work
Start with the architecture: what must move, where capacity binds, and which engineering tradeoffs matter.
Examine source boundaries, dates, and assumptions. Distinguish vendor specifications, market observations, and modeled estimates.
Connect technical constraints to utilization, costs, supplier exposure, and the business question being tested.
Use new evidence and specialist review to refine the explanation. Make material changes visible in the published work.
In the memory series, Rui “Rick” Xie contributes engineering and standards analysis; Ben Pouladian develops the investment framing. See how the research changed ↗
The research, over time
Selected milestones from the memory series, including a changed draft premise and a published roadmap correction. Each entry links to the dated original so you can inspect the reasoning.
Initial thesis
The essay centered memory bandwidth in the inference architecture debate. It established a research question to revisit as systems and workloads change.
Read the dated publication ↗Architecture refined
The analysis broadened into memory tiers and workload-specific placement, distinguishing capacity, latency, and bandwidth requirements.
Read the dated publication ↗Draft corrected
Rick’s review of standards and vendor disclosures changed the draft’s premise: the published data did not support a worsening efficiency curve inside the HBM stack. The analysis shifted to costs around the stack.
Read the dated publication ↗Roadmap revised
The report corrected the earlier 2027 HBF-device expectation, citing a later company roadmap. It also separated read-heavy model weights from write-intensive KV cache when assessing flash demand.
Read the dated publication ↗Research areas
How do the layers work together?
Compute performance depends on the interaction of silicon, memory, networking, and software. We examine how those design choices affect system costs and the economics of suppliers such as NVIDIA.
Public interactive tool
Explore the datacenter supply chain ↗Which bytes belong where?
High-bandwidth memory, DRAM, and NAND serve different roles in AI systems. Our collaborative memory research connects architecture, bandwidth, capacity, and standards to the businesses supplying those layers.
Public research overview
Explore the collaborative memory research ↗What changes when data must travel farther?
Moving data between accelerators, racks, and facilities creates networking constraints. We study optical interconnects alongside bandwidth, power consumption, and the requirements of AI infrastructure.
Public interactive tool
Explore networking in the supply chain ↗What can the system actually connect?
Packaging brings logic and memory together. We follow technologies such as CoWoS and consider how integration choices and capacity constraints affect AI semiconductor supply and system economics.
Public interactive tool
Explore the chip-to-datacenter stack ↗What does usable compute cost?
GPU rental prices are one part of a provider decision. Our research considers hardware, workload requirements, capacity, utilization assumptions, and cluster total cost of ownership, alongside token pricing and inference margins.
Subscriber dashboard · sign-in required
Explore cluster TCO in The Stack ↗How do watts become useful output?
Power delivery, cooling, and the cost of operating hardware shape AI economics. We connect those physical constraints to infrastructure spending and the unit economics of serving model requests.
Subscriber dashboard · sign-in required
Explore the inference cost waterfall ↗For investment and strategy teams
Team research access, private briefings, and scoped projects build on these coverage areas. We agree on the question, output, and scope before an engagement begins.