← Blog

The product designer's AI workflow: what to delegate and what never to

Aug 25, 2026 · Flamel

The product designer's AI workflow: what to delegate and what never to

Ask ten product designers about their AI workflow and you get ten stacks. One names four tools, one names a plugin and a prompt library, one has a folder of screenshots and a lot of opinions. All of it will be out of date by the spring, which is the first clue that the stack was never the workflow.

A workflow is not a list of tools. It is a division of responsibility: a standing answer to what the machine produces, what you decide, and where the handover between the two sits. Get that division right and the tools become interchangeable, which is exactly what you want from something that changes every quarter. Get it wrong and no tool rescues it. So this is an ai workflow for designers written as a split rather than a stack, and the split has a rule underneath it.

The rule that decides what gets delegated

Here is the test, and it is one sentence. If the quality of a piece of work can be judged from the artefact alone, it can be delegated. If judging it requires knowing why the work was asked for, it cannot.

Run a few things through it. Whether a table has a sensible empty state can be judged by looking at the table. Whether the table should exist at all cannot: that depends on what the screen is for, who it serves, and what the team agreed to optimise this quarter. Whether a headline reads well can be judged from the headline. Whether it is making the right promise cannot.

The test works because it separates execution from context. A model has the artefact and everything ever written about artefacts like it. What it does not have is the meeting, the constraint the client would not say out loud, the last three things that failed, and the number the business needs to move. That is not a temporary gap in training data. It is the difference between producing something and being answerable for it, which is the same line that runs under everything the models still cannot do.

The rule has a corollary that is worth as much as the rule. Anything you delegate has to be described well enough that a stranger could execute it, which means the act of delegating forces you to write down what you actually want. Most designers discover, the first time they try, that they had not decided. That discomfort is not a sign the workflow is wrong; it is the workflow working, several days earlier than it used to.

What to delegate, generously

Production is the obvious one and it should go entirely. Screens, states, variants, the twelve versions of an empty list, placeholder copy that is not embarrassing, the tidy-up pass on spacing. If you are still producing these by hand as a matter of pride, the pride is expensive and nobody is paying for it.

Variation is the more interesting one. The old constraint on exploring directions was that each one cost days, so you explored two and called the second one a comparison. That constraint is gone. Ask for eight takes on the same problem, deliberately far apart, and use them the way a photographer uses a contact sheet: not to find the finished thing, but to see the range of what the problem allows.

First drafts belong on the same side of the line. A first draft of a flow, of an information architecture, of a survey, of the questions for a research call. Editing something wrong is faster than starting from nothing, and it is also more honest, because a bad draft makes the disagreement visible immediately.

And so does the grind that surrounds design work without being it. Reformatting research notes. Summarising forty support tickets into recurring complaints. Turning a rough set of decisions into something readable by someone who was not in the room. That work is real, it is not judgement, and it eats weeks.

One caution on that last category, because it is where the line gets blurred in practice. Summarising forty tickets is delegable. Deciding which of the resulting complaints is a symptom and which is the cause is not, and a model asked for both in one go will hand you a tidy ranking that looks like analysis. Take the summary, keep the diagnosis, and be suspicious of any output that arrives already sorted by importance.

A bench of identical cast parts beside one hand-cut master
Production is the cheap half now. The master it is cast from is the part that still has to be made by someone who can answer for it.

What never to delegate, strictly

Three things, and they are not chosen for difficulty. They are the three places where the answer depends on something outside the artefact.

The first is the frame: deciding what the problem actually is. Every brief arrives as a proposed solution, and the work is to hear the constraint underneath it. A model will happily accept the proposed solution as the problem and execute it beautifully, which is the most expensive failure available to you, because the output is correct and the direction is wrong.

The second is the criterion: naming what this work is meant to move, before it exists. Activation or retention. Fewer support tickets or faster checkouts. Those choices are not visible in any screen and they set which of the eight variants is the good one. Hand this over and you get work that optimises for whatever the model assumed, which is usually whatever is most common in its training data.

The third is the final call, and it is the one people give away without noticing. Not because they asked the machine to decide, but because they generated six options, felt no strong preference, and picked the one that looked most finished. That is a decision made by rendering quality. When the number moves the wrong way, "it was the best-looking option" is not an answer anyone can use.

The handover is a brief, and a brief is a prompt

Between the two halves there is one artefact, and almost every broken AI workflow is broken here. What you hand over is a brief. It states the problem as you have framed it, the criterion the work will be judged against, the constraints that are real, and the ones that only look real. Written well, it is also, exactly and without translation, a prompt.

This is why people who write good briefs got sharply more effective this year and people who write vague ones got noisier. A vague brief used to cost you a week of a junior's time and produced one wrong thing you could correct in conversation. Now it costs ninety seconds and produces forty wrong things, all plausible, all polished, and correcting them in conversation does not scale. The answer is not a longer prompt but a better-shaped one: the five fields that turn a prompt into a brief are the same five a junior would have come back and asked you for.

The skill has a name and it comes apart into parts you can practise: saying the thing you actually want in terms that cannot be misread is the same discipline whether the reader is a contractor, a colleague or a model. And the fastest way to get better at it is repetition against real constraints rather than tutorials, which is what briefs written the way clients actually write them are for.

Evaluation is the part that quietly collapses

There is a fourth thing that is technically delegable and should not be, because delegating it takes the whole structure down with it. Evaluation.

It is tempting. The model produced the work, the model can critique the work, and its critique reads well. The problem is that its critique is drawn from the same distribution as the output, so it finds the things that are commonly wrong rather than the things that are wrong here. It will tell you the contrast is low. It will not tell you the flow assumes a user who already trusts you, when your whole quarter is about people who do not.

If you cannot look at a generated screen and say what is wrong with it in terms of the criterion you set, you do not have a workflow. You have a slot machine with a design degree standing next to it. Critique is the skill the whole split rests on, which is why it is worth building deliberately rather than assuming it comes free with taste.

Building it is less mysterious than it sounds, and it is mostly a matter of sequence. Write the criterion before you look at anything, then predict how the work will fail against it, then look. Doing it in that order is what stops the polish from voting, because a rendered screen is persuasive in a way a written criterion is not, and whichever of the two you meet first tends to win.

What the week looks like when the split holds

Monday, you spend on the frame, and it feels unproductive because nothing is being made. You read the tickets, talk to the person who actually owns the number, and write down the problem and the criterion in a form somebody could disagree with. That page is the week's real output.

The middle of the week runs fast, and this is where the tools live. Briefs go out, variants come back in volume, you cut hard against the criterion, and you push the two survivors further than you would have had time to before. The ratio inverts from the way it used to be: most of the hours go to judging rather than to making.

The end of the week is the defence. Not a deck of what you produced, but the argument for why this is the problem, why this criterion, why this option, and what you expect the number to do. That page is what makes the work reviewable by someone senior and, incidentally, what makes it yours rather than the tool's. It is also why designers who worried this year about being replaced tend to arrive at the same answer once the split is explicit: the machine took the half that was already commoditised.

The two halves of the workflow with the brief as the single handover between them
Everything on the left is delegable because it can be judged from the artefact. Everything on the right needs the reason the work was asked for. The brief is the only bridge, which is why it is where workflows break.

Where this leaves the tools

Nowhere important, which is the point. Any tool that turns a brief into candidates fits the left half, and they are converging fast on being much the same thing. What does not converge is the quality of your frame, your criterion and your critique, and none of those live in software.

So the useful version of how designers use ai is not a stack to copy. It is a question to answer before the next piece of work starts: what am I handing over, what am I keeping, and is the thing I hand over specific enough that forty polished wrong answers are not the likely result.

That question is answered in practice rather than in theory, and the cheapest place to practise it is on a brief that behaves like a real one — vague in the places clients are vague, immovable where they are immovable. Take one and direct the machine with it, then judge what comes back against a criterion you wrote down first. The gap between what you get and what you wanted is the whole skill, and it is measurable from the first attempt.