Skymizer · Accelerate AI
Run ultra-large LLMs
on your own hardware.
Purpose-built LPU IP and silicon for on-prem and on-device inference. 700B-parameter models locally, at ~240W — no GPU cluster, no cloud, no compromise on data sovereignty.
Product line
From compiler IP to silicon
One architecture — LISA and a decade of compiler expertise — scaled from edge devices to on-prem racks.
HTX301
Reference chip on the HyperThought platform. Six chips, 384GB, 700B-parameter LLMs on a single PCIe card at ~240W.
Learn more →
LPU IPHyperThought™
Octa-core LPU IP on LISA v3 — multi-chip scalable, purpose-built for multimodal, agentic AI at the edge and on-prem.
Learn more →
LPU IPEdgeThought™
Compiler-centric single-core LPU IP for on-device LLM inference on resource-constrained edge devices.
Learn more →
Why Skymizer
Efficiency is an architecture, not a node
Decode-first silicon
HTX301 disaggregates prefill and decode and prioritizes decode-first design, scaling 4B → 700B via LISA without over-provisioning.
Data sovereignty
Inference stays on your hardware. Predictable cost, deterministic performance, nothing leaving the building.
Mature-node economics
Built on T28nm with standard LPDDR4/5. Compression and architecture do the work — not cutting-edge fabrication.
Compiler heritage
A decade of AI compiler and runtime work underpins the whole stack — models mapped to silicon, efficiently and reliably.
Developers
Built to be deployed, not just licensed
Quickstarts, an SDK, a supported-models list, and a community where operators help each other bring HTX301 online — from unboxing to first token.
Latest
Announcements
Skymizer Announces HTX301 — Reinventing On-Prem AI Inference
HTX301 enables 700B-parameter LLM inference locally at just ~240W on a single PCIe card — no GPU cluster required.
AwardHyperThought™ Wins "Best IP/Processor of the Year" and "Most Promising Product" at EE Awards Asia 2025
Skymizer's HyperThought LLM Accelerator IP receives dual recognition at EE Awards Asia 2025.
PartnershipSkymizer and JFE Shoji Electronics Join Forces to Advance AI Acceleration in Japan
JFE Shoji Electronics becomes an official distributor of Skymizer's AI acceleration solutions in Japan.
FAQ
Common questions
What is Skymizer?
Skymizer is a software-hardware co-design IP company founded in 2013 in Taiwan. It builds LPU (Language Processing Unit) IP and silicon for on-prem and on-device large language model inference, drawing on a decade of AI compiler and runtime expertise.
What is an LPU (Language Processing Unit)?
An LPU is a processor purpose-built to run large language models efficiently. Skymizer’s LPU IP is based on its LISA (Language Instruction Set Architecture) and maps LLMs onto efficient silicon on mature process nodes, rather than relying on cutting-edge fabrication.
What is the HTX301?
HTX301 is Skymizer’s reference chip on the HyperThought platform. Six HTX301 chips with 384GB of memory on a single PCIe card run 700B-parameter LLM inference on-prem at approximately 240W — without a GPU cluster, NVLink/NVSwitch, or complex cooling.
How large a model can HTX301 run, and at what power?
A single HTX301 PCIe card scales from 4B to 700B parameters and runs 700B-parameter inference at roughly 240 watts, using 384GB of on-card memory.
What is the difference between HyperThought and EdgeThought?
HyperThought is octa-core, multi-chip-scalable LPU IP (LISA v3) for LLMs up to 600B parameters, on-prem and at the edge. EdgeThought is a single-core, compiler-centric LPU IP for on-device inference on resource-constrained edge devices. HTX301 is the reference chip that productizes HyperThought.
Does Skymizer sell chips or license IP?
Skymizer’s core business is licensing LPU IP and integrated software to semiconductor companies, device manufacturers, and system integrators. HTX301 is a reference chip built on that IP.
Why does Skymizer use a mature 28nm process instead of a leading-edge node?
Skymizer achieves efficiency through architecture and compression — weight compression and KV-cache compression plus a decode-first, prefill/decode-disaggregated design — rather than leading-edge silicon. This keeps cost and power low while running very large models.