Agents

It looks at what it drew

Describe the video. It makes a set of thumbnails, opens each one the way a viewer would see it, throws out the ones that do not read at thumbnail size, and leaves you the files.

Make some thumbnailsFree to try. Nothing to install.

How it works

Three steps, and none of them is “configure”.

  1. 01

    Give it the brief

    A title and a sentence about the video is enough. Brand colours, a reference image or a style you are matching all help, and it says what it assumed when you leave something out.

  2. 02

    It generates, then looks

    This is the part that is not a prompt trick. Each candidate is opened and judged the way it will actually be seen — small, in a grid, next to twenty others — and the ones that fail get made again.

  3. 03

    You get files, not a chat log

    Every image lands in the workspace with a name you can read, alongside a short note on why each one was kept or dropped. Download the set and pick.

What it can do

  • Judge legibility at real size

    A thumbnail is seen about 320 pixels wide. Text that works on a monitor disappears there, and the only way to know is to look — which it does, on every candidate.

  • Hold a series together

    Give it the last three thumbnails from your channel and it works inside that language — same palette, same weight, same amount of text — instead of restarting the identity each episode.

  • Say why, not just what

    Each option comes with the one sentence that matters: what it is doing, who it is for, and the reason the two it discarded did not work.

How it is set up

The mechanics, so you know what you are getting before you sign in.

Environment
A workspace with the images in it, and no terminal — this agent's work is looking, not running commands.
Tools
Image generation, image reading, and the file tools. Seven in total.
Model
A vision model, because half the job is judging a picture rather than describing one.
Starting files
Opens on a brief you can edit. Generated images land in images/.
What persists
Everything stays in the session — come back and the set is still there.

Things people ask it

  • Four options for an episode called "Your agent can finally see what it drew"
  • Match these three — here is the last month of my channel
  • Same idea, but make the subject bigger and drop the text entirely
  • Which of these two reads better at 320 pixels, and why?

What it will not do

Every one of these is a real constraint we have hit, not a roadmap item.

  • It cannot put text on an image reliably. Image models still mangle lettering, so it treats words as art direction for you to set, and says so.
  • It has no compositing tools — no cropping, no overlays, no layout. What the model draws is what you get.
  • Generation costs credits, and looking at each result costs a little more. A set of four is a handful of cents, not free.
  • It has never seen your analytics. It judges against thumbnail craft, not against what performed for you last month.

Which model does this best

Coming soon

We are aggregating per-model results for this agent — how often a run finishes the job without you stepping in, what it costs, and how long it takes. The numbers go here once there are enough runs for them to mean anything, and all three are published together: a model that finishes fast by giving up is slow to complete, and one that completes everything by grinding is expensive.

Completion rate

Runs that finished the job without you having to step in.

Median cost

What a typical run costs, in credits.

Median time

Wall clock, from the first message to the answer.

It looks at what it drew

It is already set up. Open it and ask it something.

Make some thumbnails

Browse every agent