Needle 3 at a glance INPUTS Text prompts, plus tool definitions or an extraction schema OUTPUTS Tool calls or structured extraction, always well-formed MODEL 29-121M Laddered Simple Attention Networks, CQ2 quantisation TRAINING 360B tokens of proprietary structured dataset SPEED Up to 4k tokens/s decode and 10k tokens/s prefill on a Raspberry Pi 5 CUSTOMISATION Fine-tuning lifts every subnetwork 18 to 36 points on DroidCall; from 4 layers (29M) up the tuned subnetwork passes DeepSeek V4 Flash