Cactus Needle

8-29MB foundation models that run locally on tiny devices at 4K tok/s

Visit website →

About

Cactus Needle 3 is a family of automation foundation models designed to run on devices with severe constraints — wearables, AR glasses, smart home hubs, robots, microcontrollers, and game consoles. The models are quantized to 2-bit precision and range from 8 to 29 megabytes, yet Cactus claims they match DeepSeek V4 Flash on standard benchmarks while decoding at up to 4,000 tokens per second.

The models run entirely on-device with no cloud dependency. Pebble's founder Eric Migicovsky publicly endorsed Needle, saying the company runs it locally in the Pebble Index Ring app instead of relying on cloud inference — a device with no screen that needs to trigger actions reliably without internet.

YC-backed and open-source on Hugging Face and GitHub. The Cactus Platform provides a dashboard for deploying and managing models across device fleets.

AI-assisted draft, human-reviewed before publishing — see how we choose & review tools.

Why we picked it

The edge-AI space is full of demos that compress a model and show it running on a Raspberry Pi. Cactus Needle is different: a YC-backed team shipping production models that match a 70B-class model at 8-29MB, with a real customer (Pebble) running it on a screenless ring. If on-device AI is going to matter, this is the kind of infrastructure that makes it possible.