← All announcements

Skymizer Announces HTX301 — Reinventing On-Prem AI Inference

HTX301 enables 700B-parameter LLM inference locally at just ~240W on a single PCIe card — no GPU cluster required.


Skymizer today announced HTX301, a reference chip built on the HyperThought™ LPU platform that brings ultra-large language model inference on-premises. A single PCIe card carrying six HTX301 chips and 384GB of memory runs 700B-parameter LLM inference at approximately 240W.

A different architecture for LLMs

HTX301 employs a novel architecture that disaggregates prefill and decode workloads and prioritizes a decode-first silicon design, using LISA™ (Language Instruction Set Architecture) for unified scaling from 4B to 700B-parameter models — without over-provisioning.

The result: no massive GPU clusters, no NVLink/NVSwitch interconnects, and no complex liquid cooling. Enterprises get data sovereignty, predictable costs, and deterministic performance for agentic AI workflows.

“The era of needing superscalar GPU clusters for ultra-large LLMs is over.” — William Wei, CMO

“Purpose-built decode hardware paired with intelligent software orchestration enables prefill/decode disaggregation at scale.” — Luba Tang, CTO

Availability

HTX301 is entering early access. Register your interest to receive specifications, evaluation details, and channel availability.

Explore HTX301 More announcements