Cactus Needle
8-29MB foundation models that run locally on tiny devices at 4K tok/s
Visit website →About
Cactus Needle 3 is a family of automation foundation models designed to run on devices with severe constraints — wearables, AR glasses, smart home hubs, robots, microcontrollers, and game consoles. The models are quantized to 2-bit precision and range from 8 to 29 megabytes, yet Cactus claims they match DeepSeek V4 Flash on standard benchmarks while decoding at up to 4,000 tokens per second.
The models run entirely on-device with no cloud dependency. Pebble's founder Eric Migicovsky publicly endorsed Needle, saying the company runs it locally in the Pebble Index Ring app instead of relying on cloud inference — a device with no screen that needs to trigger actions reliably without internet.
YC-backed and open-source on Hugging Face and GitHub. The Cactus Platform provides a dashboard for deploying and managing models across device fleets.
AI-assisted draft, human-reviewed before publishing — see how we choose & review tools.
Why we picked it
The edge-AI space is full of demos that compress a model and show it running on a Raspberry Pi. Cactus Needle is different: a YC-backed team shipping production models that match a 70B-class model at 8-29MB, with a real customer (Pebble) running it on a screenless ring. If on-device AI is going to matter, this is the kind of infrastructure that makes it possible.