HTX301
The first reference chip on the HyperThought™ platform to run ultra-large models on a single PCIe card. 700B-parameter LLM inference, on-prem, at about 240W — no GPU cluster required.
The idea
A different architecture for LLMs
Ultra-large models don't need a superscalar GPU cluster. HTX301 disaggregates prefill and decode workloads and prioritizes a decode-first silicon design, using the LISA™ instruction set to scale from 4B to 700B parameters on the same architecture.
The result eliminates massive GPU clusters, NVLink/NVSwitch interconnects, and complex liquid cooling — delivering data sovereignty, predictable cost, and deterministic performance for agentic AI.
"The era of needing superscalar GPU clusters for ultra-large LLMs is over." — William Wei, CMO
No cluster tax
Skip NVLink/NVSwitch, multi-node orchestration, and the cooling that comes with them.
Sovereign by default
Inference never leaves your premises — data, weights, and prompts stay in-house.
Scale without waste
One architecture from 4B to 700B; provision for the model you run, not the worst case.
Specifications
HTX301 at a glance


Figures reflect the HTX301 reference platform and are subject to change ahead of general availability.
Availability
Get HTX301
HTX301 is entering early access. Register your interest and our team will follow up with specifications, evaluation details, and next steps.