EdgeThought™
Compiler-centric LPU IP for on-device LLM inference on resource-constrained edge devices — the game-changer for the on-device GenAI era.


Architecture
The core engine for next-gen edge AI
EdgeThought pairs a single-core, resource-efficient design with a dedicated accelerator that optimizes memory-bandwidth utilization, and a dynamic decompression engine that expands model weights on the fly. Built on LISA v2 & v3, it fits low-power devices without sacrificing inference quality.
- Low-power IoT devices
- AI PCs and edge servers
- Multi-user, multi-batch inference
Why EdgeThought
Small footprint, full capability
Single-core efficiency
A resource-efficient design that minimizes hardware demands for constrained edge devices.
Dynamic decompression
On-the-fly model-weight decompression cuts storage cost and memory bandwidth while preserving precision.
Mature-node ready
Performs reliably on mature process nodes — no cutting-edge silicon required.
Compiler-driven co-design
A decade of compiler expertise maps LLMs onto efficient hardware, end to end.
Ecosystem
Works with the tools you already use
Serving
- HuggingFace Transformers
- NVIDIA Triton Inference Server
- OpenAI API
- LangChain API
Fine-tuning
- HuggingFace PEFT
- QLoRA
RAG
- LlamaIndex
- LangChain