Frames Desk

A short file of notes on evaluating generative image tools.

Ten minutes to evaluate an image model

Most evaluations of an AI image tool go the same way. Somebody types "a cat astronaut", gets something impressive, and forms an opinion. Three weeks later the tool turns out to be wrong for the actual work, and nobody can point to the moment the decision was made.

The problem is that the impressive first image tests the model, and you are choosing a tool. Those are different purchases. Here is a ten-minute routine that tests the tool.

Minute one to three: run your own worst case

Do not run a showcase prompt. Run the thing you already know is hard about your work.

If you make product shots, run a product shot with legible packaging text. If you make editorial illustration, run the composition with three figures and a specific spatial relationship. If you work in a house style, run something that has to match it.

The showcase prompt tells you what the model does well, which you can learn from the marketing page. Your worst case tells you where the ceiling is for you specifically, and that is the number that decides.

Minute four and five: make one edit

Generate something, then ask for one change. Something narrow: change the jacket colour, remove the object on the left, move the light.

Two things to watch. Did the requested change happen? And — more importantly — what else moved? Faces shifting, colours drifting warm, text re-rendering into different text. A tool that grants the edit and quietly repaints everything else is not usable for revision work, and revision work is most of the job.

Then make a second edit on top of the first. If quality falls off between the first and the second, the tool is feeding its own output back in, and you have found the ceiling on how many rounds you can do.

Minute six and seven: read the receipt

Look at what the interface tells you about the run it just did.

You want, at minimum, the model identifier, because vendors ship several and route between them. Ideally also the resolution actually delivered — which is not always the resolution requested — and the time it took. If none of this is visible, you cannot reproduce a good result later, and every future comparison you run on this tool will be against a moving target.

This is the fastest quality signal in the whole routine, because it takes five seconds to check and it correlates with everything else.

Minute eight and nine: find the limits page

Go looking for what the tool says it cannot do.

Every tool has gaps. Transparency, text at small sizes, exact colour matching, consistent characters across sessions. A vendor that documents those has thought about them. A vendor whose site contains only capabilities has left you to discover the gaps during a project.

While you are there, check the failure billing policy and what the free tier actually licenses you to do. Both are usually one page away and both are the sort of thing that matters exactly once, expensively.

Minute ten: decide what you are buying

At this point you know the ceiling on your own hard case, whether edits are usable, whether runs are reproducible, and what the tool admits it cannot do. That is enough to choose, and it is four things more than the cat astronaut told you.

None of it requires an account, on any tool worth evaluating. If you cannot get through minute five without a card, that is itself a result.

For the one I build, what the free tier covers is written out rather than discovered — what the free image covers, where commercial rights start, what the provenance metadata is — and the model that answered is printed under the result alongside the tier, the delivered size and the elapsed time. Those are the two pages I would check first on somebody else's tool, so they are the two I try to keep honest on mine.