Agents

It looks at what it drew

Describe the video. It makes a set of thumbnails, opens each one the way a viewer would see it, throws out the ones that do not read at thumbnail size, and leaves you the files.

Make some thumbnailsFree to try. Nothing to install.

How it works

Three steps, and none of them is “configure”.

  1. 01

    Give it the brief

    A title and a sentence about the video is enough. Brand colours, a reference image or a style you are matching all help, and it says what it assumed when you leave something out.

  2. 02

    It generates, then looks

    This is the part that is not a prompt trick. Each candidate is opened and judged the way it will actually be seen — small, in a grid, next to twenty others — and the ones that fail get made again.

  3. 03

    You get files, not a chat log

    Every image lands in the workspace with a name you can read, alongside a short note on why each one was kept or dropped. Download the set and pick.

What it can do

  • Judge legibility at real size

    A thumbnail is seen about 320 pixels wide. Text that works on a monitor disappears there, and the only way to know is to look — which it does, on every candidate.

  • Hold a series together

    Give it the last three thumbnails from your channel and it works inside that language — same palette, same weight, same amount of text — instead of restarting the identity each episode.

  • Say why, not just what

    Each option comes with the one sentence that matters: what it is doing, who it is for, and the reason the two it discarded did not work.

How it is set up

The mechanics, so you know what you are getting before you sign in.

Environment
A workspace with the images in it, and no terminal — this agent's work is looking, not running commands.
Tools
Image generation, image reading, and the file tools. Seven in total.
Model
A vision model, because half the job is judging a picture rather than describing one.
Starting files
Opens on a brief you can edit. Generated images land in images/.
What persists
Everything stays in the session — come back and the set is still there.

Things people ask it

  • “Four options for an episode called "Your agent can finally see what it drew"”
  • “Match these three — here is the last month of my channel”
  • “Same idea, but make the subject bigger and drop the text entirely”
  • “Which of these two reads better at 320 pixels, and why?”

What it will not do

Every one of these is a real constraint we have hit, not a roadmap item.

  • It cannot put text on an image reliably. Image models still mangle lettering, so it treats words as art direction for you to set, and says so.
  • It has no compositing tools — no cropping, no overlays, no layout. What the model draws is what you get.
  • Generation costs credits, and looking at each result costs a little more. A set of four is a handful of cents, not free.
  • It has never seen your analytics. It judges against thumbnail craft, not against what performed for you last month.

Which model does this best

Measured over real runs of this agent, per model. A run counts as finished when it produced the answer on its own — nothing failed, nobody was asked to approve anything, and it did not run out of steps. Read all three columns together: a model that finishes fast by giving up scores badly on the first, and one that completes everything by grinding scores badly on the second.

ModelCompletion rateMedian costMedian timeRuns
gemini-3.5-flash-litegoogle100%$0.01794s4
Completion rate
Runs that finished the job without you having to step in.
Median cost
What a typical run costs, in credits.
Median time
Wall clock, from the first message to the answer.

Medians over the runs behind each row. A model appears once it has 3 runs on this agent, and the run count is shown so you can judge how much a figure rests on.

It looks at what it drew

It is already set up. Open it and ask it something.

Make some thumbnails

Browse every agent