Positron raised $875m on Thursday at a $5bn post-money valuation. In February it raised $230m at about $1.06bn. Its valuation has more than quadrupled in seven months, and its first serious product does not go into production until the second half of 2027.

The interesting thing is not the number. It is what the chip is built around.

The bet, stated plainly

Positron's Asimov design is organised around memory — between 288GB and 2,304GB per chip — and it gets there using LPDDR5X, the commodity low-power memory that goes into phones and laptops, rather than the high-bandwidth memory stacked beside a conventional AI accelerator.

That is a claim about what inference actually costs.

Training a model is compute-bound in a fairly straightforward way: you are doing an enormous number of matrix operations and the limiting factor is how many you can do per second per watt. Serving one is different. Every token generated requires reading the model's weights, and for a large model with a long context the machine spends most of its time moving data rather than multiplying it. On that view the binding constraint is how much memory you can attach and how fast you can read it, and compute is comparatively cheap.

If that is right — and it is a real position held by serious people, not a marketing line — then the winning inference architecture is the one that buys memory most efficiently.

Why the supply chain point may matter more than the architecture

HBM is manufactured by a very small number of firms, in constrained volume, and it has been a real limiter on how many accelerators reach the market. This desk wrote that the constraint on the buildout was described as electricity for two years while the nearer one was a procurement queue for hardware, and the queue exists substantially because of memory.

LPDDR5X is made in enormous quantity for consumer devices. A design that uses it is drawing on a supply chain sized for the phone market rather than one sized for the accelerator market, which is a different order of magnitude and a different set of suppliers.

Whether the resulting bandwidth is adequate is the whole engineering question, and nobody outside the company can answer it, because the chip has not taped out. Commodity memory is cheaper and more plentiful and slower per pin; the architecture has to make up the difference in how it is organised. That is exactly the sort of claim that looks either obvious or naive in retrospect and cannot be adjudicated in advance.

It fits the money, which is the part that is checkable

Inference-focused companies took about 67 percent of disclosed AI chip capital between late January and mid-August, across twelve rounds totalling roughly $5.37bn.

This single round is $875m — about a sixth of that entire eight-month total, in one company, three weeks later. The direction of travel that piece described has not just continued, it has accelerated, and it has done so at a valuation that quadrupled in seven months for a product still more than a year from production.

That is either conviction about the shape of demand in 2028 or it is the late stage of a capital cycle, and the honest position is that both are consistent with the evidence available today.

The three ways this goes wrong

Worth setting out, because the enthusiasm is unanimous and unanimity is a poor signal.

The tape-out is at the end of this year and production is targeted for H2 2027. Silicon schedules slip, and eighteen months is long enough for the workload to change underneath the design — smaller models running closer to the work have already shifted once.

The incumbent is not standing still, and the memory thesis is not secret. Nothing prevents a conventional accelerator design from attaching more commodity memory if that turns out to be where the constraint is.

And a chip is not a product. It needs the compilers, kernels and framework integration that make a model actually run on it, which is an enormous amount of unglamorous software — and the layer that supplies it has just been bought by a competitor.

What to watch

Not the tape-out, which will happen and will be announced.

Watch for a published inference benchmark on a real model — tokens per second at a stated context length and a stated cost per million tokens — measured by somebody other than the company. Memory-first architectures live or die on long-context performance, which is precisely where the claim is strongest and where the arithmetic is least forgiving.

Positron AI's raise of $875m at a $5bn post-money valuation announced on 10 September 2026, structured as a $375m Series C at a $3.5bn pre-money valuation co-led by NEA, Atreides Management, Valor Equity Partners, Andra Capital and SemiAnalysis Capital with a follow-on Series C-1 of up to $500m anchored by NEA and Jim Clark; the February 2026 raise of $230m at a valuation of about $1.06bn; the Asimov chip's design around memory capacity and bandwidth using between 288GB and 2,304GB of commodity LPDDR5X per chip; the scheduled tape-out on TSMC's N3P process at the end of 2026 with production targeted for the second half of 2027; the Titan system combining four to eight Asimov chips per node for models beyond 16 trillion parameters and context windows beyond 10 million tokens; and the planned 2MW engineering data centre are as reported by Reuters, SiliconANGLE, Converge Digest and the company's own announcement on 10 September 2026. No independent benchmark of the architecture exists, since the chip has not taped out. The account of HBM supply constraints is general industry knowledge and the inference drawn is our own.

Topics aisemiconductorsinferenceventure capital

Technology Correspondent

Alison Acosta

Alison Acosta reports on artificial intelligence, enterprise software and the infrastructure behind the modern internet, with a focus on how technical decisions become business decisions.