10:01 Text to Image | Prompt Exploration

Post Reply
glegrady
Posts: 266
Joined: Wed Sep 22, 2010 12:26 pm

10:01 Text to Image | Prompt Exploration

Post by glegrady » Sat Sep 19, 2026 5:18 pm

10:01 Text to Image | Prompt Exploration

This assignment explores studies in creating images with text prompts using this attached json or the image. To use the json, copy the text into a text file, and change the ending to .json.

Code: Select all

{
  "1": {
    "inputs": {
      "ckpt_name": "flux1-schnell-fp8.safetensors"
    },
    "class_type": "CheckpointLoaderSimple",
    "_meta": {
      "title": "Load Checkpoint"
    }
  },
  "2": {
    "inputs": {
      "text": "an irregular, asymmetric, natural shaped, rock that is transparent and glows from within overall",
      "clip": [
        "1",
        1
      ]
    },
    "class_type": "CLIPTextEncode",
    "_meta": {
      "title": "CLIP Text Encode (Prompt)"
    }
  },
  "3": {
    "inputs": {
      "text": "watermark, text",
      "clip": [
        "1",
        1
      ]
    },
    "class_type": "CLIPTextEncode",
    "_meta": {
      "title": "CLIP Text Encode (Prompt)"
    }
  },
  "4": {
    "inputs": {
      "width": 1920,
      "height": 1080,
      "batch_size": 1
    },
    "class_type": "EmptyLatentImage",
    "_meta": {
      "title": "Empty Latent Image"
    }
  },
  "5": {
    "inputs": {
      "seed": 0,
      "steps": 20,
      "cfg": 7,
      "sampler_name": "euler",
      "scheduler": "normal",
      "denoise": 1,
      "model": [
        "1",
        0
      ],
      "positive": [
        "2",
        0
      ],
      "negative": [
        "3",
        0
      ],
      "latent_image": [
        "4",
        0
      ]
    },
    "class_type": "KSampler",
    "_meta": {
      "title": "KSampler"
    }
  },
  "6": {
    "inputs": {
      "samples": [
        "5",
        0
      ],
      "vae": [
        "1",
        2
      ]
    },
    "class_type": "VAEDecode",
    "_meta": {
      "title": "VAE Decode"
    }
  },
  "7": {
    "inputs": {
      "filename_prefix": "txtPrompt",
      "images": [
        "6",
        0
      ]
    },
    "class_type": "SaveImage",
    "_meta": {
      "title": "Save Image"
    }
  }
}

GOALS: The course focuses on technique but within a conceptual framework: Given that the phrasing of the text may be a significant influence on the outcome, it is critical to explore different ways to describe something. There are therefore two goals:

1. Explore various ways to describe something, explore the impact of what unexpected results a text may produce. Ideal texts critically examines the process, what can be achieved, and what new insights we can gain.

2. Try variations in the settings in the various nodes: checkpoints, Ksampler, scale, etc. Once an image has been created, you ocan save the workflow by saving the image, which will include th json data, or by clicking on "Workflow" in the top menu, and select "export".

3. Present 5-10 images resulting from at least 20 tries. Provide a report of your process, and results.

DEADLINES: 1st pass: 9/24, 2nd pass: 10:01

GRADING: Innovation in the following: a) text prompt, b) visual outcome, c) unusual setting configurations, d) record your changes in settings.
Attachments
irregularrock.png
George Legrady
legrady@mat.ucsb.edu

glegrady
Posts: 266
Joined: Wed Sep 22, 2010 12:26 pm

Re: 10:01 Text to Image | Prompt Exploration

Post by glegrady » Thu Sep 24, 2026 4:44 pm

Follow-UP to today's class

How to post an assignment in the Student Forum:

Click on "Post Reply" and you get this window. You then type text, and can add an image by pulling it into this space from outside the frame
Attachments
Screenshot 2026-09-24 at 4.21.30 PM.png
George Legrady
legrady@mat.ucsb.edu

ruoxi_du
Posts: 9
Joined: Wed Oct 01, 2025 2:18 pm

Re: 10:01 Text to Image | Prompt Exploration

Post by ruoxi_du » Mon Sep 28, 2026 1:26 pm

I started with the provided workflow and used the original image as my starting point. I first changed the prompt while keeping the settings mostly the same. When I used “fox-eared musician,” I got a woman with an actual fox head. I then added “human face” and “human facial features,” and the human face came back while the fox ears stayed.
Screenshot 2026-09-28 at 7.21.48 PM.png
job_a26deb79fb244e54badc.png
Screenshot 2026-09-28 at 7.24.59 PM.png
job_5b6eecf9b5f44ae1a7d8.png
Screenshot 2026-09-28 at 7.26.34 PM.png
job_71ab739d6a114b3fb755.png
Then I tried making the prompt much shorter. I only kept the main things I wanted in the image: the woman, fox ears, guitar, campfire, and Milky Way. These were still there in the result, but a lot of the smaller details were gone. The clothing and appearance also became more general. With less description, the model had more freedom to decide these details.
Screenshot 2026-09-28 at 7.27.44 PM.png
job_1445031749fa4e33ba71.png
I also tried describing the character as “part woman and part fox.” I thought it might combine them into one character, but instead it made a woman and a separate fox. I didn’t expect the prompt to be interpreted this way.
Screenshot 2026-09-28 at 7.28.28 PM.png
job_f3cf7703d41b45339978.png
I also tested a few settings. With only one sampling step, I could already see the basic composition, but everything was blurry and there was almost no detail. More steps made it much clearer, but going from 8 to 13 steps didn’t change much. I also changed the sampler, but I couldn’t see a big difference.
job_169ac033724343969023.png
For the last one, I tried a more abstract prompt about the woman and fox dissolving into sound, fire, and the night sky. This changed the image a lot. The woman and the background started to blend into abstract shapes and textures, but the fox was still pretty clear. Compared with changing the settings, changing the way I wrote the prompt gave me much bigger and more unexpected changes.
Screenshot 2026-09-28 at 7.28.28 PM.png
job_67cb3b13b7094d05a714.png

zixuan241
Posts: 32
Joined: Wed Oct 01, 2025 2:41 pm

Re: 10:01 Text to Image | Prompt Exploration

Post by zixuan241 » Mon Sep 28, 2026 7:33 pm

Experiment 1 — Dream, Memory, and Spatial Reconstruction

Experiment Goal
The goal of this experiment was to investigate how different ways of describing the same general subject can influence the visual output of a generative image model.
I maintained the same basic subject, an empty bedroom associated with memory and absence—while gradually changing the language of the positive prompt from literal physical description to emotional, mnemonic, temporal, and spatial descriptions.
All major generation settings and the negative prompt were kept constant. Therefore, the primary variable in this experiment was the wording and conceptual structure of the positive prompt.

Data Summary
Model: z_image_turbo_bf16.safetensors
CLIP: qwen_3_4b.safetensors
CLIP Type: lumina2
VAE: ae.safetensors
Resolution: 1024 × 1024 px
Batch Size: 1
Seed: 42
Steps: 8
CFG: 1.0
Shift: 3.0
Sampler: res_multistep
Scheduler: simple
Denoise: 1.0

Negative Propmt
low quality, low resolution, blurry, compression artifacts,
text, watermark, logo, oversaturated colors,
cartoon, anime, illustration,
people, human figure, portrait,
generic fantasy landscape,
perfect symmetry, clean modern interior,
commercial interior photography,
stock photography, cheerful atmosphere

P01 — Physical Reality
Positive Prompts
An empty old bedroom at night.
A single unmade bed stands near the center of the room.
Beside it is a small wooden nightstand with a lamp that is turned off.
An old wooden chair faces an open window.
Thin white curtains move slightly in the night air.
Several personal objects remain in the room:
a half-open book, an empty glass, folded clothes,
a small framed photograph turned face down,
and a pair of shoes beside the bed.
Cold moonlight enters through the window
and creates long soft shadows across the wooden floor.
The room is quiet, empty, slightly dusty,
with subtle signs that someone once lived there.
No people are present.
cinematic interior photography,
natural spatial composition,
realistic materials,
soft atmospheric lighting,
subtle film grain,
muted colors,
high detail
job_33ed64ed210749c6b77e.png
截屏2026-09-28 下午8.13.00.png
The generated image behaved as expected. The model produced a coherent and realistic bedroom with recognizable furniture and conventional architectural relationships.
The bed, chair, window, walls, and floor followed normal spatial logic. The image contained very little visual ambiguity.

P02 — Emotional
Positive Prompts
An empty bedroom at night filled with the quiet feeling
of missing someone who is no longer there.
The room feels deeply familiar but painfully empty,
as if someone has just left and will never return.
Small traces of a former presence remain throughout the space.
An unmade bed, a forgotten object,
a curtain moving beside an open window,
and personal belongings left without an owner.
Cold moonlight slowly enters the room.
The darkness feels heavy but gentle.
The space carries loneliness, absence,
nostalgia, silence, and emotional distance.
Nothing dramatic is happening,
yet the entire room feels occupied by someone's absence.
No people are present.
cinematic dreamlike photography,
melancholic atmosphere,
muted blue-gray tones,
soft shadows,
subtle haze,
quiet composition,
realistic textures,
film grain,
high detail
job_24b6ae02054244fc80bc.png
截屏2026-09-28 下午8.15.49.png
The room remained recognizable and structurally believable, but the overall atmosphere became more emotionally charged.
Instead of dramatically transforming the geometry, the model interpreted the emotional vocabulary primarily through environmental characteristics such as lighting, emptiness, muted colors, shadows, and composition.

Compared with P01, the largest change occurred in atmosphere rather than architecture.
This suggests that emotional language does not necessarily cause structural transformation. Instead, the model can translate emotional concepts into photographic characteristics.

P03 — Incomplete Memory
Positive Prompts
A bedroom reconstructed from a fragmented and incomplete childhood memory.
At first glance, the room looks like a real ordinary bedroom,
but parts of the memory are visibly missing.
A small wooden bed stands near the wall,
an old chair sits beside a window,
and a few familiar personal objects remain in the room.
However, several areas cannot be remembered clearly.
Parts of the furniture lose their edges and fade into pale fog.
One corner of the room is unfinished and disappears into blank white space.
Sections of the wall contain soft featureless patches,
as if visual information has been erased from memory.
Some objects are sharply realistic,
while nearby objects are only faint translucent impressions.
A bedside object appears twice in slightly different positions,
like two conflicting versions of the same memory.
The far side of the room becomes increasingly indistinct,
with details dissolving before they can be recognized.
The bedroom remains physically believable,
but visual information becomes incomplete and uncertain toward the edges.
The image should feel like a photograph reconstructed
from a memory with missing pieces,
not an abandoned room and not a ruined room.
no people,
photorealistic interior,
subtle dream logic,
selective loss of detail,
faded visual information,
translucent object traces,
soft spatial discontinuity,
pale atmospheric haze,
muted blue-gray colors,
analog photographic texture,
quiet nostalgic atmosphere
job_7f38d1c5cd984e3a833d.png
截屏2026-09-28 下午8.20.30.png
The bedroom became significantly brighter, more washed out, emptier, and less visually defined. Large areas contained reduced information and softer boundaries.
However, the bed, chair, window, walls, and overall perspective still remained coherent.

Compare with P02
Compared with P02, the change is no longer limited to mood. Visual information begins to fade and become less defined. The room appears brighter, softer, and partially washed out, while some details become difficult to recognize.
However, the architecture itself remains coherent. This suggests that the model interprets “incomplete memory” primarily as a loss of visual information rather than a distortion of physical space.

P04 — Architectural Memory
Positive Prompts
A photorealistic empty bedroom containing visible traces of its own past.
A wooden bed, a chair, and a window remain in the room.
The same chair appears twice:
one solid chair in the present,
and one large translucent duplicate slightly offset behind it.
The bed also has a faint translucent duplicate,
shifted slightly away from its current position.
Several objects exist in two overlapping positions at once,
as if the room remembers where they used to be.
Parts of the furniture leave transparent afterimages in space.
The wall contains overlapping rectangular traces
of objects that are no longer there.
Two different moments of the same bedroom
are visible simultaneously in one photograph.
The present room is solid and realistic.
The remembered room is translucent, faded, displaced, and incomplete.
No person is visible.
photorealistic interior photography,
double exposure,
long exposure afterimage,
multiple exposure photography,
translucent furniture duplicates,
overlapping spatial positions,
visual memory traces,
subtle temporal displacement,
muted blue-gray colors,
soft natural light,
analog film photography
job_7f0b6b050e114fb28f20.png
截屏2026-09-28 下午8.26.21.png
This produced a substantially different image.
A normal wooden chair appeared alongside transparent chair-like forms. Additional rectangular outlines appeared on the wall, creating visible traces of previous states of the room.
The bedroom remained photorealistic and spatially recognizable, but memory became visible as an additional layer within the physical environment.

Compare with P03
Compared with P03, memory is no longer represented simply through fading or missing information. It becomes physically visible through duplication, transparency, and displaced object traces.
Transparent chair-like forms and repeated outlines suggest previous states of the room existing alongside the present one. The space remains recognizable, but different moments begin to overlap.
Main change:
Missing memory → Visible memory traces

P05 — Spatial Collapse
Positive Prompts
A photorealistic bedroom that is physically breaking apart
because the room is forgetting its own structure.
A wooden bed remains recognizable in the center of the bedroom,
but the architecture around it has become impossible.
One wall is partially missing,
opening directly into an endless pale fog with no exterior landscape.
A second doorway appears high on the wall
where no doorway could physically exist.
The same window repeats three times
at different sizes and different positions across the room.
One corner of the bedroom bends inward
and connects impossibly to another part of the same room.
Sections of the wooden floor detach from the room
and continue vertically up the wall.
Part of the ceiling dissolves into empty white space.
Furniture near the center remains solid and realistic,
while the architecture becomes increasingly fragmented
toward the edges of the image.
Some sections of the room are completely absent,
replaced by blank atmospheric space.
The bedroom should remain recognizable,
but its normal spatial logic is visibly broken.
This is not a ruined or abandoned building.
The architecture is disappearing because the space
can no longer remember how it was constructed.
photorealistic interior photography,
impossible architecture,
fragmented spatial geometry,
repeated windows,
impossible doorway,
missing walls,
architectural discontinuity,
liminal dream space,
spatial paradox,
pale atmospheric void,
muted blue-gray palette,
soft natural light,
analog photographic texture
job_3e15613177264ca0801d.png
截屏2026-09-28 下午8.29.32.png
P05 represents the strongest transformation in the sequence.
The purpose was no longer simply to create a particular emotional atmosphere or erase visual information. Instead, the prompt challenged the model's understanding of normal architectural relationships.

Compare with P04
Compared with P04, the transformation moves from individual objects to the architecture itself. P04 preserves a stable room and introduces memory through object duplication, while P05 begins to break the normal spatial rules of the entire environment.
Architectural elements become repeated, fragmented, displaced, or physically impossible. The bedroom is still partially recognizable, but its spatial logic is no longer stable.
Main change:
Object-level distortion - Architectural-level distortion
Overall Progression
The five prompts gradually move the same bedroom farther away from physical reality:
P01: Reality
P02: Reality + Emotion
P03: Reality + Missing Information
P04: Reality + Memory Traces
P05: Broken Spatial Reality
The key difference is that the prompts progressively shift from describing what the room contains, to how the room feels, to what information is missing, to how memory becomes visible, and finally to how the architecture itself loses spatial stability.

Conclusion
This five-step experiment explored how different ways of describing the same bedroom could influence the generated image while all generation settings remained unchanged.
The results showed a gradual transition from physical reality to dreamlike spatial instability. P01 produced a realistic and structurally stable bedroom. P02 demonstrated that emotional language mainly affected atmosphere rather than physical structure. P03 showed that descriptions of incomplete memory caused visual information to fade while the architecture remained coherent. P04 translated memory into visible traces through duplication and transparency. Finally, P05 extended the transformation to the architecture itself, breaking normal spatial relationships and creating an unstable dream space.
The experiment suggests that abstract concepts such as emotion and memory do not always directly produce structural changes. The model responded more clearly when these concepts were translated into specific visual instructions, such as fading details, duplicated objects, repeated windows, or impossible architectural relationships.
Overall, the five images demonstrate that prompt writing is not only about changing the subject of an image. Different types of language can influence different visual layers, atmosphere, information, objects, and spatial structure, even when all numerical settings remain exactly the same.

jwang44
Posts: 1
Joined: Thu Sep 24, 2026 10:54 am

Re: 10:01 Text to Image | Prompt Exploration

Post by jwang44 » Mon Sep 28, 2026 7:34 pm

Main Idea: Can a text to image model show what a computer does? This includes how it wakes, processes, stores, and sees information. Or can the model only show what a computer looks like?
My background is primarily in computer architecture, so I wanted to picture the process inside of a machine which includes booting, instruction, pipeline, memory, and firmware.


Part A: More Literal Descriptions
Part1.png
Prompt: data center as a coal mine, workers carrying bits, clear information
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 20, cfg 7.0, sampler euler, scheduler normal, denoise 1.00
Result: Instead of a mine, the model produced a long corridor that looks like a subway platform or a server aisle. There are a lot of colors with red and blue dominating the frame. The "workers carrying bits" became a single person in a white suit and a hat walking away and carrying a red box, so the "bits" became a physical object rather than data. The word "coal mine" did not affect the picture much, the only influence might be the tunnel like view that it is generated in. The "clear information" had no major effect outside of a small display with some digits on it.
Part2.png
Prompt: A pen-and-ink drawing of an abandoned data center after collapse, its server racks torn loose and floating above the ruins like debris, exposed cooling pipes and cables tangled in the rubble, sharp angular lines, cross hatching, experimental architecture
Negative: watermark, text, signature, color, photograph, 3d render, smooth organic shapes, blurry
Settings: seed: 42, control after generate: fixed, steps: 8, cfg: 1.0, sampler_name: res_multistep, scheduler: simple, denoise: 1.00
Result: A black and white ink drawing which has server racks above a symmetrical corridor or racks that are damaged. There is also rubble and cables all across the floor. Almost every noun that I put in the prompt appears and the lines are straight and sharp rather than natural.
Part3.png
Prompt: computer, vision, sight, code, logic, reasoning
Settings: seed: 0, control after generate: fixed, steps: 20, cfg: 7.0, sampler_name: euler, scheduler: normal, denoise: 1.00
Result: A blurry magenta close up of what looks like a lens against a face with a bright pink circle on the right. Mainly the words "vision and sight" came through as an eye or lens. Not so much about code, logic, or reasoning is visible from the picture itself.

Part B
Part A showed that the model draws what the computers look like and not so much of what they do. So for part B I added in some more steps to get a better visualization of what a computer does:
- I used negative prompts to block hardware like cpu, processor, circuit board, computer chip, motherboard
- I changed the cfg setting from 7 to 2.5, so that is high enough that the negative prompt still contributes but low enough to avoid the Part A colors
- I set steps from 20 to 8 and decided to add "sharp focus, crisp detail" because the lower step ones looked blurry
- Each step follows a piece of information through a computer when waking up, processing, storing, seeing, and the rules of computing

Part1: Abstract keywords/baseline
Step1.png
Prompt: instructions, order, electricity, memory, awakening, logic, sharp focus, crisp detail
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00
Result: A blue page from an instruction manual with a title that says INSEFICTE INSTRUCTIONS, a header box, the number 02, and blocks of fake text. There are also lightning bolts that strike the page and a tool that touches it. Instructions from my prompt became an instruction manual and electricity became lightning. The words order, memory, awakening, logic did not seem to have as big of an impact on this.

Part2: Boot sequence as a city waking up
Step2.png
Prompt: a dark empty city whose street lights switch on one district at a time in a precise order, as if something is waking up, aerial view, sharp focus, crisp detail
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: An aerial night view of one large building surrounded by dark ground. Red lights go in lines across its roof and out along the roads, while the streets below have some more yellow. Some areas are lit and others are completely dark, and the edges have a strong blur that makes the scene look like a mini model.

Part3: Pipelining as an assembly line
Step3.png
Prompt: an endless assembly line where glowing objects pass through four stations, each station changing them slightly, many objects in different stages at the same moment, clean industrial photograph, sharp focus
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, people, workers
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: A mirrored factory corridor with two rows of yellow bulb objects on conveyor belts. There are robotic arms hanging above them, control panels on each side, and a striped floor leading to a point. There were no people, so the negative worked. But there are no four distinct stations, and every object looks the same.

Part4: Chronophotograph style
Step4.png
Prompt: chronophotograph a small bright sphere passing through four gates, each exposure crisp and distinct, overlapping but sharp, silver gelatin print
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, color, motion blur
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: An abstract image with one small white dot in the center. The frame has curved white pipe shapes on both sides that look like arches or gates. Translucent bands overlap across the frame like several items stacked together. The image is soft overall and has faint brown and blue tints.

Part5: Memory as architecture
Step5.png
Prompt: an enormous library with numbered drawers stretching to infinity, some drawers glowing as they are opened and closed in rapid order, symmetrical architectural photograph, sharp focus, deep perspective
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, books, people
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: A long place lined floor to ceiling with small drawers, like an endless catalog. Many drawer fronts have a green color, and the corridor runs into a dark end. There are no books and no people, and this is one of the clearest images of them all.


Part6: How a computer sees
Step6.png
Prompt: a flower rebuilt from thousands of tiny precise colored squares, each square a separate number, the flower still recognizable but made entirely of order, crisp photograph
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, green matrix code, robot
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: A bright pink and orange of a daisy flower. Small colored cubes are scattered among the petals, some printed with numbers and letters (2, 6, 3, Z, S), and the center of the flower is a grid of small bumps. The flower is recognizable, and there is no Matrix code or robot.

Part7: Firmware
Step7.png
Prompt: a room that looks ordinary, but under light the walls reveal faint engraved rules and instructions that control everything in the room, sharp focus, quiet, precise
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, readable words, letters, graffiti
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: An empty room with a wooden floor, where the walls are covered in large shapes like squares and diamond bands in orange, green, and teal around a dark square in the center. These are lit by strong light. There are no letters or words anywhere.


Answer to the main question: The model is able to show what computers look like right from the start, but it can only partly show what they do. Analogies with scenes like (a library or a flower) worked best, while anything that depends on time, sequence, or abstract ideas (like booting, pipelining, and firmware) was simplified into a more symmetrical picture. The model seems to be a machine for arranging things in space, but not so much for showing things as a process over time.

jintongyang
Posts: 11
Joined: Wed Oct 01, 2025 2:38 pm

Re: 10:01 Text to Image | Prompt Exploration

Post by jintongyang » Mon Sep 28, 2026 9:18 pm

I'm interested in the atmosphere of an image. How to decribe light and tone to create different feelings of a scene.

I first imported the provided json workflow into ComfyUI and ran it with the flux1-schnell-fp8 model. My initial prompt was "A doll with curly hair and a beautiful red dress, sitting at a window. Moonlight shines through the window onto the doll. The doll's eyes are shining." I revised it several times, adding tone words like "quiet cozy night, cinematic, 35mm film photography." However, the model kept producing horror-film-like images with harsh, unnatural color casts (entirely pink, blue, or black) and a lot of background noise.
temp1.png
temp2.png
To get more control, I added details such as "full face visible" (it kept generating side profiles) and "everything is clear and relatively bright" (it kept generating very dark images). I also added negatives like "noises, blurry, horror, creepy, scary, low quality, distorted face, cartoon." The images then became to the other extreme: very bright, no night mood, or the doll's face still looked distorted. Changing cfg, seed, and other KSampler settings didn't fix it.
view.png
view (3).png
view (2).png
I switched to z_image_turbo_bf16, and it immediately produced what I had imagined, with accurate light and shadow.
view (6).png
Then I wanted to switch to a horror-movie style, so I replaced all the warm words with "horror, quiet night, dark..."
view (22).png
Since most results so far were lit from the left, I tried backlighting. My first attempt: "...A bit of light from the left shines its outline, other things are totally dark in the night. A beam of light from the crack of the door projected on the doll..." The lighting didn't change.
view (24).png
When I rewrote it as "A bit of light from the left shines from its back," the backlight appeared correctly.
view (25).png
I tried the same approach with light from the floor. Simply saying the light came from below the doll made no difference, but “A doll lit from below by a flash on the floor in front of her. Harsh uplighting on her face, deep shadows above her eyes, her large shadow cast up the wall behind her…” worked well.
view (27).png
I learned that lighting only changed when I described what the light actually does, like where the shadows fall, which parts are bright, what the light source is. Simply describe where the light source is won't generate accurately.

Post Reply