Video from July, before the Tiiny arrived (MacBook inference, the pauses are the point): https://www.youtube.com/watch?v=2VZUEtDk78Y
My daughter is autistic. What she needs from a companion is not cleverness — it is the same answer to the same question, every single time, and a memory of what she told it last week. And because it is a child's diary and moods we are talking about, nothing may ever leave the house. No cloud account, no API bill, no telemetry.
So I built one. It is called Robi. It speaks German, runs on a Reachy Mini robot (or as an app on her iPad when the robot is not around), and since mid-August the Tiiny is its brain.
Open conversation with long-term memory of what she tells it
The evening diary ritual — she talks as long as she wants, it listens
Social role-play: greeting someone, a phone call, ordering food, asking for help — scenes I write in a parent dashboard in two minutes
Routines with step-by-step check-ins, and it waits until she says she is done
Calm-down breathing, paced by the robot's head motion or the face on the iPad
Special-interest mode: it remembers what she loves and weaves it into everyday answers
Explains idioms literally when she meets them (matters a lot for her)
Bedtime stories it invents, ten minutes long, stored so "nochmal" replays the exact same one
Pictures she asks for, printed on the family printer when she says so
"Vergiss das" deletes the last diary entry — her data, her call
Safety: real distress always reaches us, and she knows that rule exists
"Was kannst du?" gets the identical answer every time. Predictability everywhere.
Three tiers on the home LAN, nothing else:
Robot / iPad — mic, speaker, face, motion. Deliberately thin, no logic.
Backend on a NAS — orchestrator, skills, safety, scheduler, parent dashboard, and the encrypted memory (SQLCipher + sqlite-vec). The brain that outlives the robot.
Inference host — STT, TTS, LLM and embeddings behind OpenAI-compatible HTTP. This is where the Tiiny lives.
The inference tier was a MacBook until August. It worked, but the MacBook had to stay on, and a good reply took 11s on mains and 22s on battery. You can feel that pause in the video. A child's patience is about 2s.
Job | Model | Status |
|---|---|---|
Chat | Qwen/Qwen3.6-35B-A3B | live since 2026-08-16 |
Embeddings (memory search) | Qwen/Qwen3-Embedding-0.6B | works, switching over |
Speech recognition | Qwen/Qwen3-ASR-1.7B | loads, waiting for the transcript fix in 1.0.0 |
Pictures | Tongyi-MAI/Z-Image-Turbo (+ the Children-Drawing finetune) | waiting for the image engine fix |
STT, TTS and pictures still run on the Mac for now. The device is the brain and the memory index; ears, voice and eyes come next.
End of speech to first spoken reply: 3.5–4.1s across a four-turn session. About 2.4s of that is Whisper on the Mac, which is why ASR on the device is the one I am waiting for most.
A full bedtime story, 7.5 minutes of spoken audio, generated in 92s — decode outruns playback, so she never waits mid-story.
Warm prefill 630–765 tok/s, decode 25–32 tok/s. Four models loaded at the same time (chat, embeddings, ASR, image) at ~80% NPU RAM, no fighting.
Embedding calls do not evict the chat prompt cache. On my previous setup they did, on every turn. This alone bought me seconds.
Thinking is on by default, and reasoning_effort is ignored. Qwen will silently spend your whole token budget thinking and hand you an empty reply. The switch that works:
"chat_template_kwargs": {"enable_thinking": false}
One system message only. The chat template rejects a second one (HTTP 400, "System message must be at the beginning"). If your framework injects context as extra system messages, fold it into the user turn.
The prompt cache is a single slot, keyed on the byte prefix. Keep your system prompt byte-identical across turns, put retrieved context into the user message, store history exactly as sent — and the whole conversation stays cached. Any other prompt (a story, a picture request, a second app) evicts it. I warm the persona prefix at startup.
Two smaller ones: the device serves the same API directly on its LAN IP on port 80, so your backend does not need the Mac app in the path. And the tiiny CLI (ls, download, load, unload) is the way to script model loading — models do not come back on their own after a reconnect.
ASR and pictures onto the device once the fixes land, then the Mac goes to sleep for good.
The code is not public yet — it is very much shaped around one child. But if you are building something for a kid who needs things to be the same every time, I am happy to go into any detail. Ask away.
Really thoughtful build. What stayed with me is that you’re optimizing for consistency and trust, rather than simply trying to make the model “smarter.” The “same answer every time” requirement, replayable bedtime stories, and the “Vergiss das” command giving her control over her diary all feel like small details that could make a huge difference in everyday life. The performance notes were also incredibly useful, especially the prompt-cache behavior and the two-second patience window. I’m curious how you handle memory mistakes or situations where Robi isn’t sure what she means. Wishing you both the best as you move ASR and image generation onto the Tiiny—I’d love to hear how that goes.
Thanks — those two are exactly what I still lose sleep over. Memory mistakes: the safety net is her, not the model. Retrieval is deliberately simple (the five closest diary entries, no scoring), and when Robi brings up something wrong or unwanted, "Vergiss das" deletes it from the store, not just the chat. Old entries expire on their own. Uncertainty: Robi answers with what it has and doesn't ask back. For her, a clarifying question is a second wait and a second thing to process — a short, slightly wrong reply she can correct is less stressful. Routines and role-play are scripted, so the LLM only fills in the wording there. ASR and image numbers will follow once they're on the device.
I will be sure to check this out!
Make the logo a little more appealing