BEP RESEARCH

BEP research coverage

AI infrastructure research.

From silicon architecture to the economics of intelligence.

BEP Research connects engineering, market data, and investment analysis across the infrastructure stack. Our work combines Ben Pouladian’s operating and investment perspective with specialist research contributions. Start with the frameworks, inspect an observation, and follow how the work develops.

Start here

Three ways into the research.

Explore the frameworks ↗

Architecture & constraints

Memory Wars

How memory movement shapes the economics of inference, and why the right architecture depends on the workload.

Compute & capital

The Token Dollar

A framework for following capital through AI infrastructure, useful model output, and the cash flows that finance the next buildout.

Collaborative memory research

The Write Problem

A worked example of connecting memory engineering to business economics: which AI data earns a place on NAND flash.

BEP research frameworks

A question. A mechanism. A way to test it.

Architecture & constraints

Memory Wars

Ben Pouladian

Memory Wars examines how moving and storing data shapes AI infrastructure economics. A faster accelerator does not by itself resolve the capacity, bandwidth, or latency constraints of a particular workload. The research follows how those constraints influence memory tiers, system design, and supplier demand.

The business question is where spending moves when the bottleneck changes. That requires looking at the full system and its workload, rather than treating a chip specification as a financial result.

What we examine

  • Where model weights and working data sit, and how often they move.
  • How capacity, bandwidth, and latency requirements differ by workload.
  • Whether an architecture change shifts system cost or supplier economics.

This is a framework for investigation; its implications depend on the system and workload.

Original December 2025 essay on Substack ↗

Compute & capital

The Token Dollar

Ben Pouladian

The Token Dollar connects the physical buildout of AI to the financial flows supporting it. Capital funds chips, datacenters, and power; infrastructure produces model output; customer revenue must then support operating costs and the next investment cycle.

The framework asks where economic value is captured along that chain, and how financing and dollar-denominated markets may shape it. The thesis is something to test against contracts, pricing, utilization, and cash generation as the market develops.

What we examine

  • Who finances capacity and who bears utilization risk.
  • How useful output becomes revenue after serving costs.
  • Whether pricing, financing terms, and reinvestment support the capital cycle.
Original April 2026 essay on Substack ↗

Collaboration in practice

The Write Problem

Ben Pouladian & Rui “Rick” Xie

Which AI bytes actually belong on NAND? The answer depends on access patterns. Read-heavy model weights and frequently updated working state can impose very different demands on a storage or memory tier.

Rick’s engineering and standards analysis and Ben’s investment framing examine those distinctions together. The report considers writes, reuse, endurance, and movement between tiers before drawing conclusions about demand for memory suppliers. It also corrects an earlier HBF timing expectation using a later company roadmap.

What the distinction changes

  • Capacity demand alone does not establish which memory tier wins.
  • Write behavior and reuse matter when evaluating flash deployment.
  • Product-roadmap dates need to remain dated and revisable.
Read the full public report on Substack ↗

A public observation from The Stack

From a price claim
to a measurable comparison.

LLMflation compares a weighted basket of model output prices with GPT-4’s March 2023 launch price. Switch between index points and dollars to inspect the same observation.

This is a price comparison with different model compositions. It does not measure equivalent quality or the cost of completing a task successfully.

Explore The Stack ↗Full dashboard · subscriber sign-in required

LLMflation

As of 2026-09-12

21.4GPT-4 launch = 100
How to read this comparison

A price-basket comparison, not a measure of equivalent model quality or task success. Weights: OpenAI 30%, Anthropic 25%, Google 20%, DeepSeek 15%, open-weight models 10%. The launch reference is $60 per million output tokens.

How we work

From technical detail to a decision.

01

Understand the system

Start with the architecture: what must move, where capacity binds, and which engineering tradeoffs matter.

02

Check the measurement

Examine source boundaries, dates, and assumptions. Distinguish vendor specifications, market observations, and modeled estimates.

03

Interpret the economics

Connect technical constraints to utilization, costs, supplier exposure, and the business question being tested.

04

Revisit the thesis

Use new evidence and specialist review to refine the explanation. Make material changes visible in the published work.

In the memory series, Rui “Rick” Xie contributes engineering and standards analysis; Ben Pouladian develops the investment framing. See how the research changed ↗

The research, over time

What changed, and why.

Selected milestones from the memory series, including a changed draft premise and a published roadmap correction. Each entry links to the dated original so you can inspect the reasoning.

  1. Initial thesis

    Memory Wars

    The essay centered memory bandwidth in the inference architecture debate. It established a research question to revisit as systems and workloads change.

    Read the dated publication ↗
  2. Architecture refined

    The Hierarchy Rewrites

    The analysis broadened into memory tiers and workload-specific placement, distinguishing capacity, latency, and bandwidth requirements.

    Read the dated publication ↗
  3. Draft corrected

    The Bandwidth Tax

    Rick’s review of standards and vendor disclosures changed the draft’s premise: the published data did not support a worsening efficiency curve inside the HBM stack. The analysis shifted to costs around the stack.

    Read the dated publication ↗
  4. Roadmap revised

    The Write Problem

    The report corrected the earlier 2027 HBF-device expectation, citing a later company roadmap. It also separated read-heavy model weights from write-intensive KV cache when assessing flash demand.

    Read the dated publication ↗

Research areas

The full stack, connected.

For investment and strategy teams

Bring a specific question.

Team research access, private briefings, and scoped projects build on these coverage areas. We agree on the question, output, and scope before an engagement begins.