Silicon · Early access

HTX301

The first reference chip on the HyperThought™ platform to run ultra-large models on a single PCIe card. 700B-parameter LLM inference, on-prem, at about 240W — no GPU cluster required.

HTX301 evaluation board
700Bparameters, on a single card
~240Wtotal board power
384GBon-card memory

The idea

A different architecture for LLMs

Ultra-large models don't need a superscalar GPU cluster. HTX301 disaggregates prefill and decode workloads and prioritizes a decode-first silicon design, using the LISA™ instruction set to scale from 4B to 700B parameters on the same architecture.

The result eliminates massive GPU clusters, NVLink/NVSwitch interconnects, and complex liquid cooling — delivering data sovereignty, predictable cost, and deterministic performance for agentic AI.

"The era of needing superscalar GPU clusters for ultra-large LLMs is over." — William Wei, CMO

No cluster tax

Skip NVLink/NVSwitch, multi-node orchestration, and the cooling that comes with them.

Sovereign by default

Inference never leaves your premises — data, weights, and prompts stay in-house.

Scale without waste

One architecture from 4B to 700B; provision for the model you run, not the worst case.

Specifications

HTX301 at a glance

Configuration6× HTX301 LPU chips on a single PCIe card
Memory384 GB (standard LPDDR-class)
Model reach4B → 700B parameters, no over-provisioning
Power≈ 240 W total
ProcessT28nm mature node
ArchitectureLISA™ ISA · prefill/decode disaggregation · decode-first
Form factorPCIe card — on-prem deployment
TargetEnterprise & agentic AI inference
HTX301 EVB front
Evaluation board — front
HTX301 EVB back
Evaluation board — back

Figures reflect the HTX301 reference platform and are subject to change ahead of general availability.

Availability

Get HTX301

HTX301 is entering early access. Register your interest and our team will follow up with specifications, evaluation details, and next steps.