YouMind
Войти

Open Sourcing a Skill: Completely Solving the Illustration Problem for Xiaohongshu and WeChat

@op7418
УПРОЩЁННЫЙ КИТАЙСКИЙ28 мая 2026 г.
464K
1.5K
255
98
2.7K

Суть

Guizang's new open-source tool, guizang-social-card-skill, offers a professional alternative to generic AI layouts by using magazine-style aesthetics and smart image-text placement for high-quality social media content.

歸藏(guizang.ai) - inline image

Some time ago, I open-sourced guizang-ppt-skill, and later I discovered something while using it to create content myself.

The web pages generated by it, when screenshotted and posted to image-text platforms, received much better feedback and data than my manual layouts.

歸藏(guizang.ai) - inline image

I believe you have found some prompts or Skills before that generate these 3:4 card images.

They almost all have the same flavor: Tailwind + large color blocks + emoji stacking + mediocre font hierarchy.

After looking at them, I can roughly understand why AI-generated image-text cards are so easy to spot at a glance—they are making web pages, not magazines.

Image-text cards are a completely different beast compared to PPTs: vertical screens, a 1-second decision in the information flow whether to stop or not, and relying on pictures rather than words.

Different layouts, different rhythms, different readers.

So I split it out from the PPT Skill and made it into a standalone guizang-social-card-skill (https://github.com/op7418/guizang-social-card-skill).

Below, I'll talk about why it's good and why I'm willing to spend so much time on it.

2. Why exactly is it good?

3:4 vertical images are the main battlefield for image-text cards. Most of the design energy of this Skill was spent on 3:4—font hierarchy, layout proportions, and line-breaking rules.

Everything has been calibrated according to the real scenario of being swiped through in a mobile information flow. 21:9 and 1:1 WeChat Official Account header images are also supported.

歸藏(guizang.ai) - inline image

Let's start with what image-text creators care about most.

2.1 It distinguishes what you are writing and matches it with the right style

Content on image-text platforms is categorized. A movie review and a product review require completely different visual languages;

A travel diary and workplace tips should use different layouts.

But most AI tools don't care about this; they use the same template for whatever you write.

The result is that everyone's cards look like they came off a WeChat Official Account cover assembly line.

This Skill has built-in adaptation rules for 11 common image-text categories:

  • Travel / Lifestyle: Magazine style, warm color palette, full-screen images, large serif titles;
  • Workplace / Tips / Business Insights: Grid style, dark backgrounds, data-heavy poster layouts;
  • Film / Culture: Cool-toned magazine style, movie poster layouts, character close-ups prioritized;
  • Product Reviews / Digital: Grid style, comparison matrices, device-framed screenshots;
  • Reading / Notes: Magazine style, serif fonts, centered quote layouts, maximum white space;
  • Food / Shop Visits: High-saturation magazine style, top-down shots prioritized, text moved to corners;
歸藏(guizang.ai) - inline image

I even specifically made a map component for travel bloggers. You can mark shop locations and travel routes on it, and the AI will automatically generate annotations for you.

歸藏(guizang.ai) - inline image

Feed it the same text: tell it it's a movie review, and it gives you a movie poster-style card; tell it it's a product review, and it gives you a comparison chart with device frames.

歸藏(guizang.ai) - inline image

More importantly, it has clear "no-go" zones:

  • Fan-oriented content: the required visual language is completely different;
  • Pure promotional ads: contradicts its design philosophy of emphasizing content;
  • Long tutorials exceeding 12 screens: image-text format is not the optimal carrier for long tutorials.

In these scenarios, the Skill will tell you at the beginning, "You might want to use another tool."

I left this intentionally. Boundaries define a product more than the capabilities themselves; a Skill that tries to do everything usually ends up doing nothing well.

2.2 How to overlay text on images

Overlaying text on images is the hardest part of image-text cards and the place where "AI feel" is most easily exposed.

If not done well, three types of failures occur:

  1. Text covers faces or the center of the product.
  2. White text on light backgrounds or black text on dark backgrounds is unreadable.
  3. Text spans the entire image, ruining an otherwise beautiful composition.
歸藏(guizang.ai) - inline image

The Skill handles this in three steps:

  1. Identify subjects in the image: faces, products, text-dense areas, and automatically avoid them in the layout;
  2. Calculate the color and brightness of the landing area: decide text color, whether to add a mask, and how deep the shadow should be;
  3. Adaptive font size and line breaking: dynamically adjust font size and line break positions based on the size of the landing area, rather than hardcoding a font size that overflows.
歸藏(guizang.ai) - inline image

With these rules, the "premium feel" of the card is established. Readers can't tell the difference between "text pressed on" and "text that was originally there."

2.3 Where do images come from: This is the biggest difference from other AI card tools

Most AI tools for generating image-text cards either ask you to upload your own images, use emojis instead, or generate illustrations that look like AI at first glance.

The result is that manual image hunting is exhausting, or stacking emojis looks fake.

This Skill is connected to three free commercial-use image libraries by default:

  • Pexels: supports Chinese search, good for general scenarios;
  • Unsplash: strongest photographic texture, the first choice for people, life, and space content;
  • Wallhaven: games, photography, and wallpapers are all here, though copyright is messy.
歸藏(guizang.ai) - inline image

It automatically dispatches search terms based on the semantics of the text paragraphs, retrieves images, crops them to fit the layout, and avoids cutting off faces or subjects.

You get a card with real photography, not a color-block card.

And it's not rigid about finding images that are absolutely free of copyright issues. It will tell you which images it found, and you decide whether to use images with unclear copyright.

Also, platforms are currently very strict about AI watermarks. Most AI-generated images you use now have watermarks, which get flagged by platforms and can lead to shadowbanning. This is a major pain point.

2.4 Screenshots are also images: The Four-Piece Beautification

Much of our content can't use photography; it needs to be software screenshots, chat records, or product interfaces.

The Skill has a built-in screenshot beautifier:

Adds macOS/iOS style device frames (browser chrome or phone borders), uses different background textures to hold the screenshot—grid paper, dot matrix, warm white, or dark—so the screenshot no longer floats with a white background on a white background;

At the same time, it automatically matches shadow layers and corner radius parameters based on the visual style. Two styles each have a screenshot recipe, ensuring consistency without manual adjustment.

歸藏(guizang.ai) - inline image

Simply put, a random screenshot you take will look like an official product promotional image after passing through it.

2.5 AI Image Generation: Use with Restraint

The Skill only calls for AI image generation when all previous image sources fail to find suitable material.

When generating images, it forces style constraint words to avoid the mediocre visuals of "obvious AI illustrations."

I'd rather it use AI less than have it use AI in a way that makes all image-text cards look like sisters. This also avoids content exposure being affected by using AI images.

歸藏(guizang.ai) - inline image

2.6 Visual System: Two Styles + 28 Layout Skeletons

Those familiar with my previous PPT Skill will find this familiar. These two visual systems and layout skeletons were adapted and recalibrated from the PPT Skill.

I won't repeat the details, but here's how they look in the image-text card context.

Two visual systems:

  • Magazine Style: The kind of typography you see on the covers of The New Yorker or Shanghai Translation Publishing House. Large white space, large serif titles, asymmetrical layouts, and text with room to breathe.
  • Grid Style: In the vein of Massimo Vignelli and Helmut Schmid's Swiss graphic design. Strong grids, sans-serif fonts, geometric feel, restrained but precise use of color.
歸藏(guizang.ai) - inline image

28 layout skeletons, which I selected from magazines, posters, album covers, and movie posters I've seen over the past decade—those that stand up to scrutiny.

AI is still mediocre at "free layout design." By giving it a proven skeleton, its task is downgraded from "designing" to "filling," and the stability of the finished product immediately improves.

10 sets of theme color palettes, fixed font pairings, and limited icon libraries—I won't list these details one by one.

歸藏(guizang.ai) - inline image

The logic is the same: constraints are not obstacles; they are the baseline.

Give a content creator infinite color choices, and they are more likely to make something ugly; give them 10 sets of proven palettes, and the probability of them making something decent approaches 100%.

3. Why do this?

3.1 Design Perspective: Magazine feel is very effective

Why go with magazine and grid styles instead of more "modern" card designs?

The essence of an image-text card is the same as a printed poster, pictorial, or album cover. Use a static image to convince a stranger to stop in 1 second. Magazines and posters have studied this thoroughly over the past hundred years.

Web design language is made for scrollable, interactive scenarios. Moving it to a static image makes it look forced and the information flat.

歸藏(guizang.ai) - inline image

So, all the "whys" of this Skill's visual decisions:

  • Why large white space? White space is the magazine's way of telling you "the focus is here."
  • Why prioritize serif fonts? Serif fonts have the weight of printed matter at large sizes.
  • Why asymmetrical layouts? Asymmetry creates visual rhythm, letting the eye know where to look first.
  • Why restrained colors? In social media feeds, a restrained palette is actually more eye-catching than high saturation; it stands out from all the "loudly shouting" cards around it.

These decisions sound "abstract," but they are all specific constants in the code. Font size ratios, white space ratios, grid column counts, contrast thresholds, line-breaking rules. These constants are the true moat of this Skill.

3.2 Product Perspective: It is a product, not just a Prompt

After making so many Skills, I've formed a judgment on "what a Skill actually is":

A Skill is essentially a small product.

歸藏(guizang.ai) - inline image

Applied to this project:

I wrote a PRODUCT.md for it, clearly explaining what problem it solves, who it's for, and what it doesn't do. It's to force myself to think clearly about "what I am actually doing." If I can't explain it, the Skill shouldn't be released.

I give it version numbers (v0.5 / v0.9 / v0.10 / v0.12), and every version has a CHANGELOG. I can tell you why v0.10 was a failed attempt and how v0.12 fixed it.

I wrote a HANDOVER.md for it, explaining what the deliverables look like, where the boundaries are, and when to use other tools. I hope anyone taking it over can have a complete understanding of it within 30 minutes.

I list what it's not good at in advance to save users from trial and error.

Why go to all this trouble?

Because the biggest problem with the Skill ecosystem is that most Skills are satisfied with "I can make one," and few people pursue "doing it to the extreme."

A Skill should be a small product that can stand on its own. A Prompt can be copied by a competitor in ten minutes; a product cannot.

The flip side is, if I can't even explain the boundaries of my own Skill, I have no right to ask others to hand over their workflow to it.

Final Words

This Skill helped me understand what my PPT Skill actually got right.

What it got right was that it was treated as a product from the very beginning. Many templates, detailed rules, and beautiful colors are all by-products of this.

If anyone asks me what a Skill is in the future, I will answer with two sentences:

A Skill is a product. To judge if a Skill is good, see if it has been favored by its author.

歸藏(guizang.ai) - inline image

If you are also creating image-text content, I hope it helps you save those good topics ruined by bad layout.

If you are also making Skills, I hope it makes you rethink whether the thing you made is worth having a PRODUCT.md.

GitHub: https://github.com/op7418/guizang-social-card-skill

Tell your Codex, Xiaolongxia, ClaudeCode, or Workbuddy: Help me install this Skill: https://github.com/op7418/guizang-social-card-skill

Сохранение в один клик

Используйте YouMind для глубокого чтения вирусных статей с помощью ИИ

Сохраняйте источники, задавайте точные вопросы, обобщайте аргументы и превращайте вирусные статьи в полезные заметки в одном рабочем пространстве ИИ.

Исследовать YouMind
Для авторов

Превратите ваш Markdown в аккуратную статью для 𝕏

Когда вы публикуете длинные тексты, изображения, таблицы и блоки кода, форматирование в 𝕏 становится мучением. YouMind превращает полный черновик в Markdown в чистую статью, готовую к публикации в 𝕏.

Попробовать Markdown для 𝕏

Другие паттерны для анализа

Недавние виральные статьи

Смотреть другие виральные статьи