Needle 3 at a glance
INPUTS
Text prompts, plus tool definitions or an extraction schema
OUTPUTS
Tool calls or structured extraction, always well-formed
MODEL
29-121M Laddered Simple Attention Networks, CQ2 quantisation
TRAINING
360B tokens of proprietary structured dataset
SPEED
Up to 4k tokens/s decode and 10k tokens/s prefill on a Raspberry Pi 5
CUSTOMISATION
Fine-tuning lifts every subnetwork 18 to 36 points on DroidCall;
from 4 layers (29M) up the tuned subnetwork passes DeepSeek V4 Flash