LPU IP

EdgeThought™

Compiler-centric LPU IP for on-device LLM inference on resource-constrained edge devices — the game-changer for the on-device GenAI era.

EdgeThought platform
EdgeThought core processing engine

Architecture

The core engine for next-gen edge AI

EdgeThought pairs a single-core, resource-efficient design with a dedicated accelerator that optimizes memory-bandwidth utilization, and a dynamic decompression engine that expands model weights on the fly. Built on LISA v2 & v3, it fits low-power devices without sacrificing inference quality.

  • Low-power IoT devices
  • AI PCs and edge servers
  • Multi-user, multi-batch inference

Why EdgeThought

Small footprint, full capability

Single-core efficiency

A resource-efficient design that minimizes hardware demands for constrained edge devices.

Dynamic decompression

On-the-fly model-weight decompression cuts storage cost and memory bandwidth while preserving precision.

Mature-node ready

Performs reliably on mature process nodes — no cutting-edge silicon required.

Compiler-driven co-design

A decade of compiler expertise maps LLMs onto efficient hardware, end to end.

Ecosystem

Works with the tools you already use

Serving

  • HuggingFace Transformers
  • NVIDIA Triton Inference Server
  • OpenAI API
  • LangChain API

Fine-tuning

  • HuggingFace PEFT
  • QLoRA

RAG

  • LlamaIndex
  • LangChain