👋 Hey there, I’m Lenny. Each week, I share deeply researched product, growth, and career advice. For more: Lenny’s Jobs | Lenny’s Podcast | Lennybot | How I AI | Become an AI-Native Builder and my other favorite AI/PM courses
P.S. Get a full free year of Cursor, Notion, Replit, Lovable, Wispr Flow, Linear, ElevenLabs, Factory, PostHog, Granola, Brain.fm, Waking Up, and more, by becoming an Insider subscriber (while supplies last). Learn more.
I’d always thought AI was bad at design. But after reading this mind-blowing post by Anshu Chimala, I realize I was just doing it wrong. Anshu led software engineering and design teams at Apple for 12 years, focusing on research and prototyping for future AI products. He regularly shares design tutorials and demos on X (he’s one of my favorite follows). For deeper dives into crafting distinctive experiences with AI, check out his Substack and connect with him on LinkedIn.
Let’s get into it.
A conversational calorie tracker, built in three prompts with Claude Fable 5:
A space exploration game, built in two prompts with Claude Opus 5:
A dynamic landing page, built in three prompts with Claude Opus 5 + GPT-5.6 Sol:
I often post AI design demos like these on X. Every time I do, someone inevitably asks, “Why does the model create all this incredible stuff for you, but when I try, I only get generic slop? It’s like you’re using a completely different model.”
I’m not using a different model, but I am getting more out of the models I work with. Most people only see 1% of AI’s creative potential. I want to show you how to tap into the other 99%.
AI models are capable of amazing creativity, but that creativity gets stifled by how they’re trained. Large language models are next-token predictors: at each step, they look at a sequence of text and predict what comes next based on millions of examples. The results may be rated by humans, and those ratings fed back into the model. This teaches the model to make consistent, safe choices that fit everyone’s preferences.
This makes typical LLMs great at most tasks but poor designers. To create a design, an LLM has to build it out token by token. Whenever it needs to make a design decision—what colors to use, or how to arrange elements—the model fills in the tokens it thinks are most likely to please everyone. As a result, the design usually ends up being repetitive and bland. It’s like the ultimate case of design-by-committee.
Great design, on the other hand, starts with feeling and aims to create an emotional response. It bends the rules and delights users with memorable, unexpected choices. Great design is exactly the opposite of what an LLM does naturally, which is to make the most predictable choice at every step.
However, if we can get the model to reach beyond the most predictable choices, we can access a vast landscape of creative ideas that most people miss out on.
This is a lesson I learned from managing human designers, before I was managing AI ones. For most of my career at Apple, I led an R&D team designing exploratory future AI products. Early on, our preconceived notions about how user interfaces should work limited our creativity and kept us returning to the same old ideas. Through rigor and new processes, we learned to stop re-creating what’s comfortable and instead look to the fringes of what’s possible, to generate something new. We became experts at polishing the little details to an Apple level of quality.
Since my time at Apple, I’ve been working on applying that same process to my work with AI. In the past couple years, AI agents have become extremely capable. They can do in hours what used to take my team weeks. And with the right guidance, they can create designs that look completely unlike anything else.
Loosely inspired by the Double Diamond design process, I’ve reimagined the design process for a team of AI agents instead of human designers:
Discovernew ideas beyond the average slop by exploring a variety of directions and creating bold, ambitious design briefs.Definean individual design identity by pushing AI beyond its familiar patterns and chaining models together to fully realize the design’s potential.Delivera stunning final result by polishing away the sloppy rough edges and focusing on the key elements.
By following these stages and applying the techniques within each one, you can create an incredible design remarkably quickly—and make people ask, “Why does AI create magic for you (and not me)?”
The hardest part of the design process is looking at a blank screen with infinite possibilities. The best way to tackle that moment is to start by going broad before going deep. AI is an excellent tool to explore a wide variety of potential directions.
As we know, though, models tend to overrely on familiar patterns and make conservative choices. To explore the full potential design space, we want to coax a model to do the opposite: be bold, be varied, and take risks. Below are two ways to push it out of its comfort zone.
The idea here is to get the model to find a new source of inspiration for designs, rather than relying on the defaults it learned from training. If you’ve tried to prompt a model to design a website or app, you’ve probably already seen what that default looks like.
As a simple example, I gave four instances of Claude Code the same prompt:
Prompt:
Build me a landing page for my productivity app.
Claude Opus 5:
Almost every time, we get a purplish gradient, text on the left, graphic on the right, and the exact same structure. It looks like every AI-designed website ever.
We didn’t ask the model to do anything unique or varied, so it makes sense that it keeps falling back on the same patterns it knows well. But just asking for variety doesn’t work:
Prompt:
Build me a landing page for my productivity app. Give me something totally unique. Make every design decision completely at random.
Claude Opus 5:
The results are different from before, but they’re still not varied. The model always uses the same color scheme, structure, and even the same awkward pottery metaphors. It’s predicting tokens that sound random but aren’t actually random.
The problem is that the model can’t inherently act randomly. It can only predict the most likely token. If we want variety, we have to bring it from outside the model. One technique for this is String Seed of Thought, published by Sakana AI. We make the AI generate a random string and use it as design inspiration. That way, the model is truly making different decisions each time.
Prompt:
I want you to build me a landing page for my productivity app.
Follow this procedure:
Generate a long, random alphanumeric string using a shell script.
Define the creative direction (color scheme, layout, typography, etc.) based on the string. Look beyond the surface for subpatterns, special numbers, anything that inspires you.
Use your judgment to bring this direction to life and make it look great.
Don’t reveal the string in the design. It’s only for your inspiration.
Claude Opus 5:
Suddenly the outputs are much more varied! Now we’re seeing different color schemes, fonts, and new ideas. The previous designs were ones that any Claude user could get. These designs are one-of-a-kind; no two runs ever produce the same result.
Another approach to giving a model a strong push is to get more specific and wild with your prompts. This gives the model a clear vision to base its decisions on, rather than letting it make them up on the fly. The best way to find a unique idea is by bringing your own taste into the equation. You first imagine the inspiration—a video game, an interior design trend, an art installation—and describe how you’d like that inspiration to influence the AI’s outputs. Here are some examples:
“Build me a landing page for my productivity app, with a bold pixel art theme and stunning graphics. Each section should feel like a still from a video game, yet somehow it should all function as a landing page.”
“Build me a landing page for my productivity app, set in an isometric living 3D city, where different features are somehow represented by neighborhoods or buildings.”
“Build me a landing page for my productivity app, with a radically asymmetric layout, dissonant colors and typography, and uncomfortable negative space. Break all the rules but still make it look good.”
Of course, the hard part is coming up with original ideas to ask for. AI can help with this too, but if you simply ask it for ideas, you’ll get the same average ones everyone else gets. Here’s a system I use to find unique prompt ideas with AI:
I want to come up with a bold, unique design language for my product. Can you list as many ideas as you can, with short, high-level descriptions? Go broad, not deep.
Industrial Control Panel:
I’m imagining something tactile. Clicky, satisfying buttons, nice sounds.
Initially I pictured something cartoony or skeuomorphic, but this feels tacky to me. Avoid that.
Instead, want consistent components and little touches that land this look without going overboard.
Gray gradients would look boring. Need more texture. Maybe we can incorporate some color, while retaining the control panel feel?
Can you sharpen this one based on my tastes?
Can you write a concise prompt that an AI agent could use to build an initial POC page with this?
If you just paste AI-generated ideas back into AI, it’s hard to get something unique. After all, anyone else could have done the same thing. However, when you actively steer the design direction, you end up with something only you could have created.
Don’t be afraid to try ideas that sound terrible. If you find yourself thinking, “There’s no way this will work,” you’re on the right track. Often, your agent will surprise you, and you’ll realize you were underestimating it. If not, just throw away those results and try something else. But save the prompts that don’t work, and test them again when newer models come out. That way, you’ll know you’re taking full advantage of what the latest models can do.
So far, we’ve looked at how to explore a broad set of ideas and hopefully land on a promising initial design. No matter how we prompt, though, our initial AI-generated designs will usually still feel generic.
For example, look at the designs we came up with using seed strings:
These have promise, but they’re still relying heavily on the same stale patterns: text on the left with a CTA button below, nav bar up top, graphic on the right.
Our next goal is to give each design an individual personality through distinct design choices. Below are my favorite techniques to do that.
We need to iterate on our designs to improve them. But simply asking our agent to look at the design and improve it won’t work, because the agent isn’t objective: it reviews its own code, past decisions, and previous rationale. AI can’t easily zoom out, look at the big picture, and “think different.”
To solve this, instead of letting the coding agent decide when the design is good enough, have it ask another agent—a “design critic.” The critic’s job is to look at screenshots of the current design and provide feedback. It doesn’t care how the current design is implemented or how much effort went into it, only if it actually hits the quality bar.
This approach has an extra benefit: we can use a big, expensive model for the critic without breaking the bank, because we’ll only use it for executive decisions. A cheap, fast model can do the grunt work, while the strong critic model provides taste.
Let’s try this on our previous designs, using Claude Fable 5 as the critic:
Prompt:
I want you to improve this design. To figure out what to focus on, use a Fable 5 subagent as a design critic.
Follow this procedure at each iteration:
Capture a screenshot of the current design
Invoke the critic in a fresh context, with just the screenshot, not the code, implementation details, or earlier iterations/critiques
Ask it to evaluate the aesthetic that the design is going for, imagine how a top design studio would execute this aesthetic, then outline the biggest gaps
Lastly, it should provide a score out of 10 indicating how close the current design is to that studio-level quality bar
Provide this guidance to the critic in its prompt:
It should think high-level about the overall structure and composition as well as look at the fine details
It should watch out for patterns that feel overdone, excessive, or otherwise obviously AI-generated, and penalize them
It should provide tight, specific feedback, not vague prose
It should be bold and opinionated, not rely on what’s safe or easy
Your work is only complete when the critic independently deems it 9/10 or higher. Do not put that criterion in the critic prompt; keep it objective in its scoring. Use the same critic prompt each time.
Claude Opus 5:
Instead of the same cookie-cutter layout over and over, each design now has its own identity—but still maintains its original high-level aesthetic.
Notably, in each case, Fable accounted for less than 10% of output tokens. Asking Fable to redesign the page directly would have cost twice as much and taken much longer.
The way you set these loops up matters a lot. Here are some tips:
**Make sure the criteria for the critic are as clear and objective as possible.**Bad: “Judge if our design looks beautiful, not AI-generated.” This is too subjective, and the results will vary wildly from run to run.
OK: “Review the aesthetic we’re going for, visualize how a top design studio would execute it, then judge our design’s quality against that bar.” The prompt is still mushy, but it provides a consistent framework and quality bar.
Great: “Here are 5 designs: 4 professional examples and 1 screenshot of our product. Rank them by polish and taste level.” This instruction is concrete and objective, and gives a visual baseline for judgment.
**Provide example images to demonstrate the target quality bar.**You can use comparable screenshots or designs you like, or even AI-generated concept art. Instruct the critic to treat these as a baseline or a moodboard, not a target. You don’t want it to copy other designs outright.**Set the stopping criteria carefully.**Otherwise, the critic may never consider the design good enough, and your agent will helplessly burn tokens trying to please it. Prompt it to do one or two iterations first, and see if it’s converging before adding more.**Choose the right model for each job.**Consider bigger models for the critic role, since more parameters generally translate to better design sense and a wider distribution of ideas. Small models can be effective as the implementer, but don’t go too small. You still need a model that’s capable of executing a design direction well.
Coding agents love to write code, but they usually don’t incorporate images. Instead, they tend to use the easy code-based alternatives: gradients, shapes, and basic patterns. Those are all strong giveaways of an AI-generated design.
Some agents have image tools built in, but they underutilize them. Others don’t have image tools out of the box but can easily use the OpenAI or Gemini APIs to generate images with an API key.
Let’s try this on the designs from the last step:
Prompt:
The design is pretty plain. Add more personality using image generation. Consider shaders or 3D effects in combination with images to create more interesting visuals.
For image generation, use this OpenAI API key (only use it locally, do not store it in the code or product): sk-a1b2c3d4…
Verify that your work looks right frame-b…