← All announcements

Skymizer Launches EdgeThought LLM Accelerator IP for On-Device LLM Inferencing

EdgeThought is a compiler-centric LPU IP engineered to accelerate large language models on resource-constrained edge devices — the game-changer for the on-device GenAI era.


Skymizer today launched EdgeThought™, a compiler-centric Language Processing Unit (LPU) IP designed for on-device LLM inference on resource-constrained edge devices.

Compiler-driven co-design

Leveraging a decade of compiler expertise, EdgeThought uses a single-core, resource-efficient design with a dynamic decompression engine that decompresses model weights on the fly — dramatically reducing storage cost and memory bandwidth while maintaining inference precision. It performs reliably on mature process nodes, no cutting-edge silicon required.

Ecosystem from day one

EdgeThought (LISA v2 & v3) integrates with the tools developers already use:

  • Serving: HuggingFace Transformers, NVIDIA Triton Inference Server, OpenAI API, LangChain API
  • Fine-tuning: HuggingFace PEFT, QLoRA
  • RAG: LlamaIndex, LangChain

Target deployments include low-power IoT devices, AI PCs, edge servers, and multi-user, multi-batch inference scenarios.

Explore HTX301 More announcements