How to Ask an AI for Design
I am a product designer. This site is mine from design through code, but I don’t type the code. I write the requests instead.
Work like that long enough and a strange pattern shows up. Same AI, same screen, and some days it lands on the first try while others take five rounds back. I used to call that luck. Not anymore. The quality of the output rode far more on the precision of the request than on the model.
What follows are calls I actually made over the last few days on this site, and the places where I was wrong. These aren’t principles I set out with. They’re what I learned after being wrong.

1. Looking at it was not checking it
I was replacing the gallery entry card on the case pages with a 3D depth stack. Five images recede inward, dimming as they go back.
I built it and looked at the screen. The back layers were hazy, so I judged the blur was working. I nearly moved on.
Out of habit I printed the computed values instead:
d0: filter: none
d1: filter: none
d2: filter: none
d3: filter: none
d4: filter: none
Not one layer had blur on it. What I read as haze was the brightness drop I’d applied to the back layers. The eye can’t separate the causes of blur, so it settles on whichever one sounds right.
One CSS line did it. In max(0px, var(--d) - 1), 0px is a length and var(--d) - 1 is a number. Mix units and the browser throws out the whole expression, which kills the entire filter property. No error, no warning. Nothing happens, quietly.
Here’s the part that stings. The same bug was in the prototype file. I looked at that prototype and decided “the blur version is better.” The screen I was looking at had no blur in it at all. I praised something that wasn’t there.
So I changed how I ask. Not “add blur” and done, but “add blur, then print the computed filter per layer.” Values, not eyes.

2. Adjectives can’t be passed through as-is
Next was deciding what the same card does on mobile. A screen driven by a finger has no hover, so it has to move on its own. Here’s what I asked for:
Could it loop on mobile? Not too fast, though.
That sentence settles nothing. “Not too fast” could be 4.2 seconds or 9, and those are completely different impressions.
So instead of asking for a value, I asked for options. One card every 4.2 seconds, one at 6.5, and a third that leaves the images alone and only breathes the spacing between layers. Three of them, actually running, side by side.
Seeing them made the answer obvious. The card doesn’t sit on screen for long, so at 6.5 seconds not even one image turns over as you scroll past. The breathing one is quiet but never shows you a new image. I went with 4.2.
Translating an adjective into a value isn’t something an AI can do for you. That part is taste. But turning an adjective into comparable options is exactly what it can do.
3. The constraint was a search range, not an answer
The card started as three square images overlapped at five degrees. I wanted it gone, so I asked:
Not overlapping. Something different, nothing like it.
Three came back. Vertical slices each sliding at a different speed, a film strip cut off at both edges, an image lying under ink that clears wherever the cursor passes. None of them used overlap.
And all three fell short somewhere. The third one especially: at rest it read as an empty card. Nothing but a grey field until you hover. On a device with no mouse, empty forever.
That’s when I saw the constraint was the problem. “No overlapping” was a means to make it look unlike the old one, and somewhere along the way it had become the goal. I lifted it.
Overlapping is fine. Give me more range.
Six more came back. Photos dropping along the cursor’s trail, a depth stack receding inward, a corner curling to reveal the sheet beneath, a hand of cards fanning out, a contact sheet laying everything out small, a deck dealing one card at a time. What I chose in the end used overlap.
A constraint narrows the search. If narrowing produces nothing, there is no shortage of answers. The range is wrong.
4. State the technical concern, but don’t hand over the decision
Building the depth stack, I dropped the blur. I had a reason. Blurring all five layers is expensive to composite on mobile, and the brightness step already reads as depth.
Then I saw the prototype and reversed myself. The blurred version was clearly better.
At that point there are two easy moves: push the performance concern through, or ignore it for the sake of the look. I took a third. The front two layers stay sharp; blur only starts at the third. Full depth, less than half the blur work.
That is the design engineer’s seat, I think. Performance is a negotiable condition, not an absolute one, and so is taste. Only someone holding both can build the compromise.
5. Rules get written into code, not into conversation
The footer has a band of words scrolling past. It used to read:
Designer · Typography · Editorial · Grid systems · Identity · Brand
A graphic designer’s vocabulary. It no longer matched the work, so I changed it:
Design engineer · Prototyping · Interaction · Design systems · Motion · Interface
Then I looked again and this was wrong too. Motion, Interface, Interaction: that’s a list of surfaces I can handle. Only the job title had changed; the character was identical. Third pass:
Design engineer · Problem framing · First principles · Prototyping · Judgment · End to end
All three times, a test broke. A test was holding on to the old list. It’s a nuisance and it’s the point: the code remembers what the promise was.
So on the last pass I added a ban list. Stack names like React and TypeScript, and alongside them Motion, Interface, Typography, Grid systems. It blocks the drift back to a skills list. The next time I break my own rule absentmindedly, the test speaks first.
Writing gets held the same way. Korean written by AI has tells: a comma in 61% of sentences where a person uses 26%, phrases like “by means of” and “in conclusion” on repeat, sentence endings collapsing into one form. Rather than catching that by eye each time, I built a checker. This piece went up after its comma ratio and stock phrases were measured.
6. When a request keeps failing, the level is wrong
Back to the word band. I asked twice and failed twice. The first pass swapped the vocabulary, the second swapped the job title. The result was still a list of things I can do.
The third time, what I said was this:
Emphasise fundamental problem solving rather than things like Motion. Something more first-principles than design skills.
That was the first time the level of the request changed. The first two asked for different words inside the list. The third asked for a different kind of list.
When the same failure repeats, don’t look for a more precise word. Ask again one level up. Not what to fix, but what you’re asking to be fixed.
In short
The six come down to one thing. An AI doesn’t make the call for you. It lets you test the call quickly.
Nine interaction prototypes in a day. Three loop speeds running side by side. Blur values printed per layer. Any of those used to require borrowing an engineer’s time, so most of it got decided by imagination. Now it gets decided by building it.
Choosing what to make, measuring whether it’s right, admitting when it isn’t: still mine. And that share hasn’t shrunk. If anything, more chances to test means more calls to make.
A request is a spec. A blurry spec gives you a blurry result.