It looks at what it drew
Describe the video. It makes a set of thumbnails, opens each one the way a viewer would see it, throws out the ones that do not read at thumbnail size, and leaves you the files.
How it works
Three steps, and none of them is “configure”.
- 01
Give it the brief
A title and a sentence about the video is enough. Brand colours, a reference image or a style you are matching all help, and it says what it assumed when you leave something out.
- 02
It generates, then looks
This is the part that is not a prompt trick. Each candidate is opened and judged the way it will actually be seen — small, in a grid, next to twenty others — and the ones that fail get made again.
- 03
You get files, not a chat log
Every image lands in the workspace with a name you can read, alongside a short note on why each one was kept or dropped. Download the set and pick.
What it can do
Judge legibility at real size
A thumbnail is seen about 320 pixels wide. Text that works on a monitor disappears there, and the only way to know is to look — which it does, on every candidate.
Hold a series together
Give it the last three thumbnails from your channel and it works inside that language — same palette, same weight, same amount of text — instead of restarting the identity each episode.
Say why, not just what
Each option comes with the one sentence that matters: what it is doing, who it is for, and the reason the two it discarded did not work.
How it is set up
The mechanics, so you know what you are getting before you sign in.
- Environment
- A workspace with the images in it, and no terminal — this agent's work is looking, not running commands.
- Tools
- Image generation, image reading, and the file tools. Seven in total.
- Model
- A vision model, because half the job is judging a picture rather than describing one.
- Starting files
- Opens on a brief you can edit. Generated images land in images/.
- What persists
- Everything stays in the session — come back and the set is still there.
Things people ask it
- “Four options for an episode called "Your agent can finally see what it drew"”
- “Match these three — here is the last month of my channel”
- “Same idea, but make the subject bigger and drop the text entirely”
- “Which of these two reads better at 320 pixels, and why?”
What it will not do
Every one of these is a real constraint we have hit, not a roadmap item.
- It cannot put text on an image reliably. Image models still mangle lettering, so it treats words as art direction for you to set, and says so.
- It has no compositing tools — no cropping, no overlays, no layout. What the model draws is what you get.
- Generation costs credits, and looking at each result costs a little more. A set of four is a handful of cents, not free.
- It has never seen your analytics. It judges against thumbnail craft, not against what performed for you last month.
Which model does this best
Coming soonWe are aggregating per-model results for this agent — how often a run finishes the job without you stepping in, what it costs, and how long it takes. The numbers go here once there are enough runs for them to mean anything, and all three are published together: a model that finishes fast by giving up is slow to complete, and one that completes everything by grinding is expensive.
Completion rate
Runs that finished the job without you having to step in.
Median cost
What a typical run costs, in credits.
Median time
Wall clock, from the first message to the answer.