Building Better Frontends with an Image-to-Code approach使用图像到代码的方法构建更好的前端

AI 工作流243 次阅读约 37 分钟

图像

Most AI-generated websites still look average or sloppy.大多数 AI 生成的网站仍然看起来平平无奇或杂乱无章。

Not because the models are useless, but because the workflow is usually wrong. People ask a coding agent to “make it modern,” “make it clean,” or “make it premium” and then expect great taste to appear right away.不是因为模型没有用,而是因为工作流程通常不对。人们让一个编程代理“让它变得现代”、“让它变得干净”或“让它变得高端”,然后期待立刻出现出类拔萃的品味。

Sometimes that works. Most of the time, you get the same AI landing page again: centered hero, gradient blob, random cards, weak spacing.有时候这样可行。大多数情况下,你得到的还是同样的 AI 着陆页:居中的英雄图、渐变光团、随机卡片、间距不足。

The better approach: separate the visual design step from the coding step.更好的方法:将视觉设计步骤与编码步骤分开。

-> Instead of asking an agent to invent and code the entire website at once, generate high-quality website images first. -> Then turn those images into code section by section.-> 不需要让代理一次性发明并编写整个网站,先生成高质量的网站图像。 -> 然后将这些图像逐段转换为代码。

That is the image-to-code frontend workflow.这就是图像到代码的前端工作流程。

What image-to-code means什么是图像到代码

Image-to-code means you start with visual references instead of code.将图像转换为代码意味着你从视觉参考开始,而不是从代码开始。

The workflow looks something like this:工作流程大致如下:

  1. Generate images of the website or app screens you want生成你想要的照片网站或应用程序屏幕
  2. Pick the best visual direction选择最佳视觉方向
  3. Extract or recreate the assets from those images从这些图像中提取或重新创建资源
  4. Give the images and assets to a coding agent将图像和资源交给编码代理
  5. Build the frontend section by section逐段构建前端
  6. Refine with screenshots until it matches the design通过截图进行优化,直至与设计一致

图像

Instead of asking the model to invent taste, layout, assets, responsiveness, and implementation at once, you give it a visual target.与其让模型一次性发明品味、布局、资源、响应式和实现,你给它一个视觉目标

Why this works better than prompting code directly为什么这种方法比直接提示代码更有效

Frontend quality is visual before it is technical.前端质量在技术之前是视觉化的。

A good website depends on spacing, hierarchy, composition, typography, color, assets, and small decisions.一个优秀的网站依赖于间距、层次结构、构图、排版、颜色、资源和小的决策。

图像

Now Image generation is different. You can describe the look in plain words and the model can focus almost entirely on visual output. In my experience, it is much easier to tune an image skill for taste than a coding skill.现在图像生成是不同的。你可以用平实的语言描述外观,而模型几乎可以完全专注于视觉输出。根据我的经验,调整图像技能以适应品味比调整编码技能容易得多。

A coding skill needs to care about implementation, structure, edge cases, performance, accessibility, responsiveness, naming, and visual polish. An image skill can be much more direct:编码技能需要关注实现、结构、边缘情况、性能、可访问性、响应性、命名和视觉润色。图像技能可以更加直接:

This is the kind of layout. This is the type of composition. This is the level of minimalism. This is the brand feeling. Do not make it generic.这是布局的类型。这是组合的方式。这是极简主义的程度。这是品牌的感觉。不要让它变得泛泛而谈。

Why Images 2.0 is currently the best choice为什么图像 2.0 目前是最好的选择

Right now, my preferred model is ChatGPT Images 2.0. It is especially strong for website-images because it understands layout, hierarchy, interface structure, typography and visual systems better than most image models I have tried. (nb, grok imagine,...)目前我最喜欢的模型是 ChatGPT Images 2.0。它在处理网站图片方面特别强大,因为它比我所尝试过的其他大多数图片模型更能理解布局、层级、界面结构、排版和视觉系统。(注:grok imagine,...)

And you can use reference images to get the exact style you need.你可以使用参考图片来获取你需要的精确风格。

Some tips: A good frontend reference image needs:一些小贴士: 一个良好的前端参考图像需要:

  • clear section structure清晰的区域结构
  • readable hierarchy可读的层级结构
  • usable interface patterns可用的界面模式
  • realistic spacing逼真的间距
  • strong art direction
  • assets that make sense and match the theme 有意义的且符合主题的素材
  • enough details for an ai agent to copy足够的细节供 AI 代理复制

You can use it inside ChatGPT. If your coding environment supports image generation, you can also generate it there, but I personally prefer the ChatGPT interface for exploration because it is simple, fast to iterate and easier to control.你可以在 ChatGPT 内部使用它。如果你的编码环境支持图像生成,你也可以在那里生成,但我个人更喜欢使用 ChatGPT 界面进行探索,因为它简单、快速迭代且更容易控制。

Image skills图像技能

A skill is basically a reusable Markdown instruction file that tells the model what kind of output you want, what to avoid and how to think about the design.一项技能基本上是一个可重用的 Markdown 指令文件,它告诉模型你想要什么样的输出、要避免什么以及如何思考设计。

For this workflow, I built/wrote two image generation skills that you can use here:对于这个工作流程,我构建/编写了两个图像生成技能,你可以在这里使用它们:

The web skill is for landing pages, websites, dashboards and multi-section frontend concepts.网页技能用于落地页、网站、仪表板和多部分前端概念。

The mobile skill is for app screens, mobile-first products and onboardings.移动技能适用于应用界面、移动优先的产品和入职流程。

You can copy the Markdown skill file into ChatGPT, Codex, or any environment where the image model is available. Then you prompt the model to use that skill when generating designs.你可以将 Markdown 技能文件复制到 ChatGPT、Codex 或任何可以访问图像模型的环境中。然后提示模型在生成设计时使用该技能。

The nice part is that image skills are easy to customize. You can tune them toward your own taste好的部分在于图像技能很容易定制。你可以根据你自己的喜好进行调整

图像

Step 1: generate the website section by section第一步:逐个生成网站部分

Do not generate one huge full-page screenshot.不要生成一个巨大的全页截图。

That usually gives you less detail, weaker layout decisions and a design that becomes hard to recreate accurately.这通常会让你得到较少的细节、较弱的布局决策,以及一个难以准确重制的设计。

Instead, generate one image per section.相反,为每个部分生成一张图片。

For example:例如:

图像

This gives every section enough visual space and detail.这为每个部分提供了足够的视觉空间和细节。

A simple prompt can look like this(with included skill):一个简单的提示可以像这样(包含技能):

Use the frontend web image skill above. Generate a section-by-section website concept for a premium local flower business called Seven Flowers. Create one separate image for each section: 1. Hero 2. About 3. Product / bouquet showcase 4. Services 5. Testimonials 6. CTA 7. Footer Style direction: minimal, editorial, calm, high-end, warm, natural. Avoid generic SaaS visuals, fake dashboard cards, gradient blobs, and overdesigned AI aesthetics.使用上述前端网页图像技能。 为名为“七花”的高端本地花店生成逐部分网站概念。 为每个部分创建一个单独的图像: 1. 英雄 2. 关于 3. 产品/花束展示 4. 服务 5. 客户评价 6. 行动号召 7. 页脚 风格方向:简约、编辑式、平静、高端、温暖、自然。避免通用 SaaS 视觉效果、虚假仪表板卡片、渐变块和过度设计的 AI 美学。

If you are just brainstorming, you can stay broader:如果你只是在进行头脑风暴,你可以保持更广泛的范围:

Use the frontend web image skill above. Generate 6 separate website section concepts for a clean SaaS landing page. Keep it minimal, creative, and implementation-friendly. One image per section.使用上述前端网页图像技能。 为简洁的 SaaS 着陆页生成 6 个独立的网站区域概念。保持简约、创意且易于实现。每个区域配一张图片。

But do not be too vague. Give at least a basic product, industry, style, or mood. So another example:但不要过于模糊。至少给出一个基本的产品、行业、风格或氛围。 再举一个例子:

Bad: make a cool website差:做一个酷炫的网站

Better: Design a minimal but creative landing page for an AI note-taking app for students. Make it calm, structured, slightly playful and not like a generic startup website. Generate one image per section.更好:为面向学生的 AI 笔记应用设计一个简约但富有创意的登录页面。使其平静、结构化、略带趣味性,不要像通用的创业公司网站。每个部分生成一张图片。

Step 2: use style words that actually mean something步骤 2:使用有实际意义的风格词汇

Do not say “modern”. “Modern” means almost nothing now.不要说“现代”。“现代”现在几乎没什么意义。

Be more precise. Use style directions that affect the layout and art direction.更加精确。使用影响布局和艺术方向的样式指示。

Examples:示例:

图像

You can also describe composition:你也可以描述组合:

图像

In image generation words like “very creative” can actually help though. In coding prompts, “be creative” often gives weird or messy results. In image generation, it often pushes the model toward more interesting layout ideas based on my tests.在图像生成中,像“非常有创意”这样的词语实际上可能会有帮助。在编码提示中,“要创意”往往会产生奇怪或混乱的结果。根据我的测试,在图像生成中,它通常会使模型更倾向于基于有趣布局想法。

Use it, but combine it with constraints.使用它,但要结合约束。

Example: Be creative, but keep the layout realistic, premium, clean, and easy to implement in a real frontend.示例:发挥创意,但保持布局真实、高端、简洁,并易于在真实前端中实现。

Step 3: generate multiple directions and pick the best parts步骤 3:生成多个方向并挑选最佳部分

Run the same prompt multiple times. Open different chats. Compare the outputs.多次运行相同的提示。打开不同的聊天。比较输出结果。

Usually, one generation has the best hero, another has the better feature section, and another has a stronger CTA. You can mix the best ideas.通常,一代产品有最好的英雄单元,另一代有更优的功能区域,还有一代有更强的行动号召。你可以融合最好的想法。

Once you find something you like, refine it directly.一旦你找到喜欢的东西,就直接优化它。

Examples:示例:

图像

Step 4: extract the assets步骤 4:提取资源

Good websites need good assets.好的网站需要好的资源。

The generated website images will often contain strong visual elements: product photos, abstract 3D shapes, illustrations, background textures, device mockups, icons, or decorative objects.生成的网站图像通常包含强烈的视觉元素:产品照片、抽象的 3D 形状、插图、背景纹理、设备模型、图标或装饰性物体。

The easiest method to get these assets is to use image generation again.获取这些资产最简单的方法是再次使用图像生成。

Open a new chat, upload the website section image and ask the model to generate and extract the asset.打开一个新的聊天,上传网站部分图像,并要求模型生成并提取资产。

Important: say “generate and extract,” not only “extract” because otherwise ChatGPT will just crop the images out.重要提示:要说“生成并提取”,而不仅仅是“提取”,否则 ChatGPT 只会裁剪图像。

Example:示例:

Generate and extract the abstract 3D object from the top-right of this website hero image. Keep it isolated, clean, high resolution从该网站英雄图片右上角生成并提取抽象 3D 物体。保持其独立、干净、高分辨率

-> Then repeat this for every asset you need.然后对您需要的每个资源重复此操作。

If the asset needs transparency, remove the background with a background remover. I usually use Adobe Express Remove Background for this.

After that, convert the final images to WebP before using them on the website. For that, I use cxnvert. (built that myself, opensource and fully free, no signup or anything)

图像

Step 5: turn the images into code步骤 5:将图像转换为代码

Now you bring everything into your coding agent.现在你将所有内容都带到你的编码代理中。

This can be Cursor, Codex, Claude Code or any agentic coding setup that can read images and modify your project.这可以是 Cursor、Codex、Claude Code 或任何能够读取图像并修改你项目的智能编码设置。

The key is patience.关键在于耐心。

Do not dump the whole website into the agent and say:不要把整个网站扔给代理,然后说:

copy all of this exactly复制所有这些内容

That usually produces weak results.那样通常会产生很弱的结果。

Build section by section.逐段构建。

Start with the hero.从英雄部分开始。

Give the agent:给代理:

  • the hero section image英雄区域图片
  • the extracted assets for that section该部分的提取资源
  • the names of the asset files资源文件名
  • clear instructions to recreate the layout清晰的重新创建布局的说明
  • the tech stack and project structure技术栈和项目结构

Review the result.检查结果。

The first attempt probably will not be perfect. That is normal.第一次尝试可能不会完美。这是正常的。

Refine it.进行优化。

Step 6: refine with screenshots步骤 6:使用截图进行优化

After the agent builds a section, take a screenshot of the result. Compare it to the reference image.在代理构建完一个部分后,对结果进行截图。将其与参考图像进行比较。

Pay attention to:注意:

图像

Then give the screenshot back to the agent with specific feedback.然后给代理提供具体的反馈,并将截图还给他们。

Even better: draw on the screenshot. Use arrows, circles, or notes.更好:在截图上绘制。使用箭头、圆圈或注释。

Example:示例:

This is close, but the hero image is too small and too low. Move it higher, increase its size by around 20%, and align the headline baseline more closely with the reference.这个接近了,但英雄图片太小且位置太低。把它移高,将其尺寸增加约 20%,并将标题基线与参考更紧密地对齐。

Repeat this for every section.对每个部分重复此操作。

That is the workflow:这就是工作流程:

  1. Reference image参考图像
  2. Code attempt代码尝试
  3. Screenshot截图
  4. Visual feedback视觉反馈
  5. Refinement优化
  6. Next section下一节

What to pay attention to需要特别注意:

Pay special attention to:特别关注:

图像

Alignment Bad alignment instantly makes a site feel cheap. Make sure text, cards, images, and grids sit on a consistent system.

Spacing Spacing decides whether a website feels premium or messy. Avoid cramped sections, random gaps, and overlapping elements.

Responsiveness Do not only build for desktop. Check tablet and mobile early. Some image-generated layouts look great in one aspect ratio but need adaptation for real screens.

Brand consistency Every section should feel like the same website. Colors, typography, image style, buttons, and visual details need to connect.

Assets Assets carry a lot of the quality. If the images, icons, or 3D objects are weak, the frontend will feel weak too.

Smoothness The layout should feel intentional. Not random. Not stitched together. Smooth transitions between sections matter.

Why this unlocks more creative frontends

The main advantage of this workflow is that it gives you a visual design phase right away

So you can explore many directions quickly:

图像

You still need taste. You still need to judge what is good. You still need to refine.

But the workflow gives you a better starting point than asking a coding agent to invent everything from scratch.

Image first. Code second.

That is the whole idea.

Useful links

Final thought

Image-to-code is not that crazy, just takes time. It kind of changes the workflow in a useful way.

For me, this is currently one of the best ways to get AI-built frontends that actually look good.

Generate the design first. Extract the assets. Code section by section. Refine with screenshots.

That is the approach.