Drag to move · right-drag to pan · wheel to zoom · double-click to reset
One training step
Hide a real training picture under noise. Ask the model what noise was added. Score how close it got. That is the whole loop — running here, live, on the real model.
Pick a training picture
The step—
—
The score
answer implies
Training would now nudge the weights to shrink that error, and repeat 30,000 times. This tab cannot — these weights are already finished.
Thirty thousand steps
Ten prompts, drawn by the real model at 27 points in its own run. Not a simulation — this is what it could actually do at each of them.
Drag the slider to move through the run.
Why the curve lies
The loss, with how much the pictures actually changed drawn over it. They agree closely — which is the trap. Both fall 84% of the way by step 200, while the output is still coloured smears, and both look flat by step 1,000 with a factor of seven still to go. That last 7× is the difference between a blob and a monster.
Can it draw what it never saw?
A second model was trained from scratch with some combinations removed — every green thing and every ghost, but never a green ghost. Then it was asked for one. Because the data is generated rather than collected, the right answer still exists for a picture the model has never met, so its attempt can be scored exactly like everything else.
What it saw, what it never saw, what it drew
What this does and does not show
Drawing a green ghost is not invention — it is recombination, and the purest possible demonstration of it. What it rules out is narrower and worth stating plainly: that the model can only return pictures it was shown. This one returns a picture nobody ever gave it. Whether it could invent a feature outside its six words is a different question, and nothing here can test it, because those six words are all it can be asked for.
Starting…