Experiment
AI agents programmed a Google chip without a manual
In about 22 hours, AI agents wrote a driver and a compiler for Google's undocumented Coral chip, and the $60 stick now runs all six layers of a small language model.
The experiment
The Google Coral USB Accelerator is an AI accelerator the size of a USB stick. It costs about $60. Google ships a compiler that translates finished models for the chip. It cannot compile transformers, the architecture behind today's language models. Google has never published how the chip is programmed internally.
On 1 October 2026 I plugged the stick into my MacBook and gave AI agents a short brief: bring tinygrad to this chip, with your own compiler instead of Google's. tinygrad is a lean open-source framework for neural networks, small enough for an agent to understand. There was no specification. One piece of advice came with the brief: fast feedback loops.
What the agents built
A coordinating agent broke the work down and handed subtasks to 23 other agents. Four of them decoded the chip's four instruction families in parallel. The result:
- a USB driver in Python that needs none of Google's runtime libraries,
- the decoded instruction set, documented instruction by instruction,
- a code generator for the chip,
- a backend through which tinygrad runs models directly on the stick.
The test was strict. Wherever Google's compiler produces an equivalent, every generated program was compared with it byte for byte. Results on the chip were checked bit for bit against a numerical model on the laptop.
After about 22 hours, all six transformer layers of the TinyStories-15M language model ran on the chip, with a single chip call per token. That gives 176.7 tokens per second for a single text. Google's own compiler cannot compile this model. The agents worked through much of the night without me.
My part
I set the goal and made the call whenever a direction had to be chosen. For example: a real tinygrad backend that takes ordinary models, instead of a special solution for a single model. And: all layers of the language model on the chip, not just the matrix multiplications. I drew the lines: no changes to the tinygrad project itself, nothing published without my OK. The video had a budget of two dollars for external AI services. It used 22 cents.
There was also manual work. Whenever the stick hung, I unplugged it and plugged it back in, about half a dozen times. Then the agents taught the driver to recover a stalled chip on its own: clear pending USB transfers, re-initialise the chip, run a control program that has to come back bit-exact. The next four stalls needed no help from me.
What was hard
Early on, the stick kept hanging. The agents found the decisive hint in a comment in Google's driver source: no new data may go out while the chip still wants to send results. Otherwise the USB link locks up.
On top of that came undocumented rules. Parameters must be aligned to 256 bytes. The chip rounds halves to even, unlike Google's reference in TFLite. After an L2 normalisation, a helper row in memory becomes unusable, and a sigmoid right after it stalls the chip. The agents found such rules by narrowing down crashes and deviations step by step.
Some bugs were silent. An image model got only 11 percent of the test images right on the chip, against 90 percent in the simulator. A call-by-call comparison found two causes. One of them: programs overwrote other layers' weights with their working buffers.
What this means for your company
First: agents can take on demanding engineering work end to end. Here that meant reverse engineering, a driver, a compiler, tests and documentation, and in the end even editing and voicing the video. What mattered was not long instructions but an objective test, fast feedback and clear guard rails. Only one process could touch the hardware at a time, work stopped at the first failure, and anything irreversible stayed with me.
Second: small local hardware can run models itself. The data never leaves the building. tinygrad's digit-recognition example (MNIST) processed almost 23,000 images per second on the stick. For narrow tasks, such as a pass/fail check on a production line, small models are often enough.
Third: hardware without proper software from its vendor is no longer necessarily a dead end. Given a reference and a test setup, agents can close the gap.
The limits
- TinyStories-15M is a very small model with 15 million parameters. It writes simple children's stories in English. Larger models do not fit into the chip's 8 MiB of weight memory.
- The chip computes with 8-bit integers. That costs measurable accuracy (perplexity 2.13 instead of 1.88), but the stories stay coherent.
- My laptop's GPU is faster: 316 tokens per second with the same model. The stick is meant for weak hosts.
- The laptop still does part of the work, mainly the output layer over 32,000 tokens, which does not fit next to the six layers.
- The whole-model program still bypasses tinygrad's scheduler.
- The agents did not start from zero. George Hotz had published a first map of the chip with edgetpuxray. Google's open-source driver provided register names and firmware, and Google's compiler served as the reference.
- Everything was developed and measured on one Mac with a single stick.
Code, video and thread are linked below. The tinygrad project shared the post on X; by 5 October it had been viewed just over 25,000 times.
Links
Which process in your company costs too much time?
In a process check we look at up to three processes and tell you honestly where an agent pays off.
Request a process check