ㅤ
Hey guys -So we’re all excited to start playing with our tiiny devices. As more and more of us get our devices online, I wanted to share a few tips I’ve learned so far. My background: I work in a tech-adjacent field (engineering) but my only experience with coding is a few classes back in high school, so I am writing t
Video from July, before the Tiiny arrived (MacBook inference, the pauses are the point): https://www.youtube.com/watch?v=2VZUEtDk78Y
My daughter is autistic. What she needs from a companion is not cleverness — it is the same answer to the same question, every single time, and a memory of what she told it last week. And
I've had a Tiiny on my bench and I wanted to see what it does with a real job rather than one prompt at a time. So I built Daybreak.
It pulls world news from 14 free feeds every five minutes and hands every single article to the Tiiny. Not a selection, all of them. Roughly 300 a day.
For each one the Tiiny does five thin
If you're setting up a Tiiny this week, here's the first thing that's going to confuse you.
The Tiiny thinks about one thing at a time. Ask it for a second thing while the first is still going and you get this:
{"code": 150004, "message": "The operation failed to complete."}Inside one program that's easy to deal with. Yo
This is the small version of my Titanium Bot, cut down to run on your Mac or Windows box next to the Tiiny with the models on the device doing the thinking. You get a chat page, memory that fills itself after a conversation and recalls what matters, seeded skills, routines on a schedule, and voice through the device's
I wanted background removal on the Tiiny and went looking in the Model Store. There are four models sitting right there for it. Turns out the image endpoint has no image input at all, so none of them can be handed a picture.
So Image Studio does what the box can actually do. You upload a png, a vision model reads it, a
If you have poked at the Vault and were not sure what it actually does, this is the plain version.
It is two things under one name. The first is a document store you can search by meaning: you drop files in, the device chunks them, an embedding model on the NPU turns them into vectors, and they land in a local database.
A benchmark for the Tiiny that shows what your box does under real work: prefill against prompt length, throughput over a long generation, what happens when several callers arrive at once, and what a reasoning model costs in wall time for the tokens nobody reads. Results are plain JSON and the report is one HTML file y
The short version is that almost none of the model runs on the CPU.
This was read off a unit rather than out of a doc. A chat model is served by a llama.cpp server, but the only weights sitting on the CPU side are the token embedding table. Every transformer layer and the output head run from a PowerInfer bundle compile
The farm is where Tiiny apps live. Eight of them today, each one a card that says what it does, which models it wants loaded on your box, and the one command that installs it.
The launcher is a real desktop app for Mac, Windows and Linux, and it does the installing and starting for you if you would rather not open a ter
Point it at a folder of markdown and it reads every note, works out which names keep turning up, and draws the ones that turn up together. Click a name and its neighbourhood pulls forward. Ask anything and the answer cites the notes it came from, or says plainly that nothing it found covers the question instead of maki
Ask for a story and half a minute later a painted page shows up in an open book and a warm voice reads it. Characters come back across nights looking and acting the same, which is the part a five-year-old notices. Everything runs on the device: the words, the picture and the voice. Nothing leaves the house.
farm install
A Tiiny runs one inference at a time, and the second caller gets device error 150004. OneLane is a single Python file you copy next to your own code. It holds a lock across processes so callers wait their turn instead of failing, rides out busy retries, and can track a model budget. Every app on the farm uses it. There
A spec sheet does not tell you what a box does when you give it real work, so this is a set of numbers measured on one instead.
What is on the page: five models, 23 runs, measured between 15 August and 15 September. Every result carries the build, the host and the NPU cost it was measured with, because a figure taken of
The store lists 50 models now, and the thing nobody says up front is that the number beside each one is a budget rather than a spec.
Your box has 100 NPU units. Everything you keep resident spends some: a 35B chat model is around 50, an image model around 32, a text-to-speech voice around 7, a small embedding model 1. T
A library that lives on the Tiiny. You ask in plain words, it answers from real books, and every answer names the book and the page it came from. If the shelf has nothing on it, it says so instead of guessing. No internet anywhere in the loop, so it works with the power out and a phone on the box's own network.
It's on
Daybreak pulls world news from 14 public feeds every five minutes and hands every article to your Tiiny. The device writes the summary, picks the category, works out where on earth it happened, scores how serious it is, and builds the embedding that groups related stories into threads. Then it draws the whole thing as
AINode Pocket runs on your computer next to your Tiiny and gives you one OpenAI-compatible endpoint in front of all of them. /v1/models is the union across your devices, and chat requests route by model id to a box that has it loaded. Each device gets a lock so several callers queue instead of hitting error 150004. The