From Now On, You Can Also Be an AI Photographer

@MANISH1027512
SIMPLIFIED CHINESESep 03, 2026
144K
533
84
26
1.0K

TL;DR

This article provides a detailed framework for AI photography, focusing on bounded randomness through variable pools for expressions, scenes, and camera techniques to achieve realistic results.

As I said before, if the retweets exceeded 100, I would release this Skill.

The goal has been met; everyone is truly supportive.

https://x.com/MANISH1027512/status/2094786658914930945

Alright, I won't just release the Skill; I'll also explain the logic behind it. After reading this, you can become an AI photographer and even expand your own style of skills.

古一 - inline image
古一 - inline image
古一 - inline image
古一 - inline image

First, the Core Prompt Framework

A Korean Instagram influencer, fair skin, exaggerated figure.

[Expression] + [Clothing Style] + [Scene] + [Moment] + [Photography Technique]

Is that it?

Yes, the skeleton is basically that simple.

I have been repeatedly emphasizing to everyone: playing with GPT-Image 2 is not like other tools—don't over-control it.

If you strictly define what the person wears, how their hands are placed, which way their head tilts, what angle the light comes from, and exactly what is in the background, GPT tends to struggle with conflicting demands. It will try to satisfy everything, but this can cause rendering conflicts (for example, that's how noise is generated).

So, control the key anchor points and appropriately delegate the rest to the model.

But delegation doesn't just mean saying "everything is random."

What's truly important is:

Prepare a thick enough set of enumerated values for the key anchor points.

For example, if you tell the model "random expression," it easily cycles through a few of the most common ones.

But if you give it:

Indifferent and spacing out, slightly squinting, looking down and distracted, unable to help laughing, suddenly looking back, a brief moment of loss, a natural expression stung by the sun, the laziness of just waking up...

The boundaries of randomness are defined by you.

This is also what I understand to be a major difference between a Skill and a regular Prompt.

What is a Skill Doing?

It's actually not that mysterious.

It's about continuously thickening these key variables into variable pools / material banks.

For example:

Expression Pool

Clothing Style Pool

Scene Pool

Moment Pool

Shot Type Pool

Focal Length Pool

Camera Position Pool

Composition Pool

Foreground Pool

Lighting Pool

When generating, you recombine from these pools.

You can understand it as a kind of bounded randomness:

It's not letting the model guess blindly, nor is it pre-defining every pixel.

This is a very common underlying logic for many image generation Skills now.

Skeleton + Variable Pools.

Here is a Version to Run Directly in ChatGPT

Copy the following, and it's basically ready to use:

n=10, 9:16 real-life candid photography portrait. The subject is fixed as: a Korean Instagram influencer with great camera presence, fair skin, and an exaggerated figure. Random expressions: indifferent and spacing out, slightly squinting, natural expression stung by sunlight, a half-smile, looking down and distracted, eyes closed feeling the wind, suddenly looking back, looking up into the distance, unable to help laughing, the laziness of just waking up, a natural pause when captured, a slight frown, thoughtful, a brief moment of loss. Random clothing styles: BM style, pure desire style, stepmother style, mature style, Chanel style, home style, Korean light mature style, French lazy style, minimalist cold style, Old Money style, vacation style, urban commuter style, sweet and cool style, Y2K style, Korean girl group private style, plain water high-end look, low-saturation high-end private clothes, relaxed private style. Random scenes: midsummer lotus pond, old wooden boat, lakeside wooden walkway, shallow country stream, bamboo forest waterside, white-walled courtyard, southern old house window, modern glass flower house, lakeside stone steps, mountain road, seaside white building, city rooftop, old bus, old-fashioned train carriage, minimalist B&B, convenience store freezer, washing fruit by the river, bench under tree shade, alley after rain, old pier, summer poolside, reed lakeside, country orchard, white pavilion, under lotus leaves. Random moments: just sat down, preparing to get up, looking down to tidy the skirt, raising a hand to adjust a hat, wind blowing hair, peeking from behind leaves, bending over to touch the water surface, barefoot stepping in water, walking slowly along the wooden walkway, dazing at the stern of the boat, turning to look into the distance, raising a hand to tidy hair after it was messed up by the wind, looking down to wash fruit, napping by the window, suddenly looking back, eyes closed sunbathing, walking from shadow into sunlight, hand on the railing, swinging legs while sitting on stone steps, using a huge lotus leaf to block the sun, squatting by the water, leaning back slightly, resting with hands on the ground, accidentally captured while passing the lens. Random shot types: extreme close-up of face, shoulder close-up, chest-up close-up, waist-up medium close-up, above-knee medium shot, half-body environmental portrait, full-body environmental portrait, small-scale figure in large environment composition. Random focal lengths: 24mm ultra-wide angle close range, 28mm wide angle, 35mm environmental portrait, 50mm natural perspective, 85mm compressed portrait, 105mm long-distance voyeuristic telephoto. Random camera positions: very low angle close to the water, low angle looking up from under a huge lotus leaf, ground-level low angle, low angle below the waist, eye-level candid, slight high angle, obvious high angle, shooting down from above stairs, side shot through plants, shooting from behind a door frame, shooting into the interior from outside a window, long-distance shot from the front row of a carriage, shooting from behind the person, over-the-shoulder perspective, Dutch Angle, person completely off-center, photographer hidden behind environmental objects in a voyeuristic perspective. Random compositions: extreme negative space, large area of sky, large area of water, person pressed into the bottom third of the frame, person placed at the extreme side, center composition but blocked by foreground, asymmetrical composition, diagonal composition, perspective leading lines, huge natural objects pressing against the person, small-scale environmental portrait, foreground occupying one-third of the frame, foreground occupying half the frame, peeking through plant gaps, door frame framing, window frame framing, lotus leaves forming a natural circular frame, architectural geometric cutting of the frame. Random foregrounds: completely out-of-focus lotus flowers, huge lotus leaves, tree leaves, bamboo leaves, white curtains, glass reflections, water reflections, petals, reeds, door frames, window frames, car glass, bus seats, transparent plastic curtains, railings, edges of white walls, wet branches. The foreground must naturally intrude into the frame, forming obvious obstruction; do not show all elements completely. Random lighting: strong midsummer noon hard light, dappled tree shade light, afternoon side-backlight, evening low-angle backlight, clear scattered light after rain, high-contrast natural light by the window, water surface reflection fill light, white wall reflection light, fine light and shadow formed through leaves, partial overexposed highlights, person at the boundary of light and dark, bright background with slightly dark person, person's face lit by only a small patch of sunlight. Random but restrained colors: blue sky + emerald green + white, dark green + milky white + skin tone, water blue + white + gray, cyan green + off-white + a few pink flowers, dark blue night + cool white light, gray-blue after rain + wet green. The frame retains at most 3-4 main color blocks, avoiding multicolored clutter. Random photography states: natural candid, non-posed, accidental capture by the photographer, person completely unaware of the lens, slight motion blur, partial out-of-focus, glass reflection ghosting, real lens flare, edge highlight overflow, slight overexposure, shallow depth of field, long-distance compression, close-range wide-angle distortion, natural grain, slight digital noise, the instant feel of a portable camera. Overall requirements: real-life photography, strong photographic texture, modern Oriental summer portrait, clean frame, clear large-scale relationships, distinct spatial layers, no studio feel, no commercial studio shooting, no standard influencer posing, not all subjects looking at the camera, no conventional centered portraits, no complex props, no cluttered backgrounds, no excessive skin smoothing, no plastic skin. Each image must randomly combine different scenes, clothing styles, moments, shot types, focal lengths, camera positions, compositions, foregrounds, and lighting, prioritizing unconventional camera positions and an accidental candid feel; avoid repetition between images in the same batch.

Skill Version

Open source repository:

https://github.com/vibeshotclub/vsc-skills/tree/main/vibeshot-candid-photography

There is no black magic in the Skill version itself.

The core is to continue increasing the thickness of the variable pools based on the above framework, allowing for more combinations between different variables.

I believe you, being smart, can already learn by analogy.

Finally, let me introduce the community I have been working on, VibeShotClub:

This is a gathering place for AIGC visual creators who are truly tinkering for the long term.

We will continue to organize and share Prompts, Skills, workflows, model techniques, and excellent cases produced by the community.

If you are also seriously playing with AI images, AI videos, or building your own creative workflow, you are welcome to come and visit.

Remix in YouMind

Turn one viral article into a full content workflow

Collect the source, decode the pattern, create assets, draft the story, and distribute from one AI workspace.

Explore YouMind
For creators

Turn your Markdown into a clean 𝕏 article

When you publish your own long-form writing, images, tables, and code blocks make 𝕏 formatting painful. YouMind turns a full Markdown draft into a clean, ready-to-post 𝕏 article.

Try Markdown to 𝕏

More patterns to decode

Recent viral articles

Explore more viral articles