10:01 Text to Image | Prompt Exploration

Post Reply
glegrady
Posts: 266
Joined: Wed Sep 22, 2010 12:26 pm

10:01 Text to Image | Prompt Exploration

Post by glegrady » Sat Sep 19, 2026 5:18 pm

10:01 Text to Image | Prompt Exploration

This assignment explores studies in creating images with text prompts using this attached json or the image. To use the json, copy the text into a text file, and change the ending to .json.

Code: Select all

{
  "1": {
    "inputs": {
      "ckpt_name": "flux1-schnell-fp8.safetensors"
    },
    "class_type": "CheckpointLoaderSimple",
    "_meta": {
      "title": "Load Checkpoint"
    }
  },
  "2": {
    "inputs": {
      "text": "an irregular, asymmetric, natural shaped, rock that is transparent and glows from within overall",
      "clip": [
        "1",
        1
      ]
    },
    "class_type": "CLIPTextEncode",
    "_meta": {
      "title": "CLIP Text Encode (Prompt)"
    }
  },
  "3": {
    "inputs": {
      "text": "watermark, text",
      "clip": [
        "1",
        1
      ]
    },
    "class_type": "CLIPTextEncode",
    "_meta": {
      "title": "CLIP Text Encode (Prompt)"
    }
  },
  "4": {
    "inputs": {
      "width": 1920,
      "height": 1080,
      "batch_size": 1
    },
    "class_type": "EmptyLatentImage",
    "_meta": {
      "title": "Empty Latent Image"
    }
  },
  "5": {
    "inputs": {
      "seed": 0,
      "steps": 20,
      "cfg": 7,
      "sampler_name": "euler",
      "scheduler": "normal",
      "denoise": 1,
      "model": [
        "1",
        0
      ],
      "positive": [
        "2",
        0
      ],
      "negative": [
        "3",
        0
      ],
      "latent_image": [
        "4",
        0
      ]
    },
    "class_type": "KSampler",
    "_meta": {
      "title": "KSampler"
    }
  },
  "6": {
    "inputs": {
      "samples": [
        "5",
        0
      ],
      "vae": [
        "1",
        2
      ]
    },
    "class_type": "VAEDecode",
    "_meta": {
      "title": "VAE Decode"
    }
  },
  "7": {
    "inputs": {
      "filename_prefix": "txtPrompt",
      "images": [
        "6",
        0
      ]
    },
    "class_type": "SaveImage",
    "_meta": {
      "title": "Save Image"
    }
  }
}

GOALS: The course focuses on technique but within a conceptual framework: Given that the phrasing of the text may be a significant influence on the outcome, it is critical to explore different ways to describe something. There are therefore two goals:

1. Explore various ways to describe something, explore the impact of what unexpected results a text may produce. Ideal texts critically examines the process, what can be achieved, and what new insights we can gain.

2. Try variations in the settings in the various nodes: checkpoints, Ksampler, scale, etc. Once an image has been created, you ocan save the workflow by saving the image, which will include th json data, or by clicking on "Workflow" in the top menu, and select "export".

3. Present 5-10 images resulting from at least 20 tries. Provide a report of your process, and results.

DEADLINES: 1st pass: 9/24, 2nd pass: 10:01

GRADING: Innovation in the following: a) text prompt, b) visual outcome, c) unusual setting configurations, d) record your changes in settings.
Attachments
irregularrock.png
George Legrady
legrady@mat.ucsb.edu

glegrady
Posts: 266
Joined: Wed Sep 22, 2010 12:26 pm

Re: 10:01 Text to Image | Prompt Exploration

Post by glegrady » Thu Sep 24, 2026 4:44 pm

Follow-UP to today's class

How to post an assignment in the Student Forum:

Click on "Post Reply" and you get this window. You then type text, and can add an image by pulling it into this space from outside the frame
Attachments
Screenshot 2026-09-24 at 4.21.30 PM.png
George Legrady
legrady@mat.ucsb.edu

ruoxi_du
Posts: 10
Joined: Wed Oct 01, 2025 2:18 pm

Re: 10:01 Text to Image | Prompt Exploration

Post by ruoxi_du » Mon Sep 28, 2026 1:26 pm

I started with the provided workflow and used the original image as my starting point. I first changed the prompt while keeping the settings mostly the same. When I used “fox-eared musician,” I got a woman with an actual fox head. I then added “human face” and “human facial features,” and the human face came back while the fox ears stayed.
Screenshot 2026-09-28 at 7.21.48 PM.png
job_a26deb79fb244e54badc.png
Screenshot 2026-09-28 at 7.24.59 PM.png
job_5b6eecf9b5f44ae1a7d8.png
Screenshot 2026-09-28 at 7.26.34 PM.png
job_71ab739d6a114b3fb755.png
Then I tried making the prompt much shorter. I only kept the main things I wanted in the image: the woman, fox ears, guitar, campfire, and Milky Way. These were still there in the result, but a lot of the smaller details were gone. The clothing and appearance also became more general. With less description, the model had more freedom to decide these details.
Screenshot 2026-09-28 at 7.27.44 PM.png
job_1445031749fa4e33ba71.png
I also tried describing the character as “part woman and part fox.” I thought it might combine them into one character, but instead it made a woman and a separate fox. I didn’t expect the prompt to be interpreted this way.
Screenshot 2026-09-28 at 7.28.28 PM.png
job_f3cf7703d41b45339978.png
I also tested a few settings. With only one sampling step, I could already see the basic composition, but everything was blurry and there was almost no detail. More steps made it much clearer, but going from 8 to 13 steps didn’t change much. I also changed the sampler, but I couldn’t see a big difference.
job_169ac033724343969023.png
For the last one, I tried a more abstract prompt about the woman and fox dissolving into sound, fire, and the night sky. This changed the image a lot. The woman and the background started to blend into abstract shapes and textures, but the fox was still pretty clear. Compared with changing the settings, changing the way I wrote the prompt gave me much bigger and more unexpected changes.
Screenshot 2026-09-28 at 7.28.28 PM.png
job_67cb3b13b7094d05a714.png

zixuan241
Posts: 34
Joined: Wed Oct 01, 2025 2:41 pm

Re: 10:01 Text to Image | Prompt Exploration

Post by zixuan241 » Mon Sep 28, 2026 7:33 pm

Experiment 1 — Dream, Memory, and Spatial Reconstruction

Experiment Goal
The goal of this experiment was to investigate how different ways of describing the same general subject can influence the visual output of a generative image model.
I maintained the same basic subject, an empty bedroom associated with memory and absence—while gradually changing the language of the positive prompt from literal physical description to emotional, mnemonic, temporal, and spatial descriptions.
All major generation settings and the negative prompt were kept constant. Therefore, the primary variable in this experiment was the wording and conceptual structure of the positive prompt.

Data Summary
Model: z_image_turbo_bf16.safetensors
CLIP: qwen_3_4b.safetensors
CLIP Type: lumina2
VAE: ae.safetensors
Resolution: 1024 × 1024 px
Batch Size: 1
Seed: 42
Steps: 8
CFG: 1.0
Shift: 3.0
Sampler: res_multistep
Scheduler: simple
Denoise: 1.0

Negative Propmt
low quality, low resolution, blurry, compression artifacts,
text, watermark, logo, oversaturated colors,
cartoon, anime, illustration,
people, human figure, portrait,
generic fantasy landscape,
perfect symmetry, clean modern interior,
commercial interior photography,
stock photography, cheerful atmosphere

P01 — Physical Reality
Positive Prompts
An empty old bedroom at night.
A single unmade bed stands near the center of the room.
Beside it is a small wooden nightstand with a lamp that is turned off.
An old wooden chair faces an open window.
Thin white curtains move slightly in the night air.
Several personal objects remain in the room:
a half-open book, an empty glass, folded clothes,
a small framed photograph turned face down,
and a pair of shoes beside the bed.
Cold moonlight enters through the window
and creates long soft shadows across the wooden floor.
The room is quiet, empty, slightly dusty,
with subtle signs that someone once lived there.
No people are present.
cinematic interior photography,
natural spatial composition,
realistic materials,
soft atmospheric lighting,
subtle film grain,
muted colors,
high detail
job_33ed64ed210749c6b77e.png
截屏2026-09-28 下午8.13.00.png
The generated image behaved as expected. The model produced a coherent and realistic bedroom with recognizable furniture and conventional architectural relationships.
The bed, chair, window, walls, and floor followed normal spatial logic. The image contained very little visual ambiguity.

P02 — Emotional
Positive Prompts
An empty bedroom at night filled with the quiet feeling
of missing someone who is no longer there.
The room feels deeply familiar but painfully empty,
as if someone has just left and will never return.
Small traces of a former presence remain throughout the space.
An unmade bed, a forgotten object,
a curtain moving beside an open window,
and personal belongings left without an owner.
Cold moonlight slowly enters the room.
The darkness feels heavy but gentle.
The space carries loneliness, absence,
nostalgia, silence, and emotional distance.
Nothing dramatic is happening,
yet the entire room feels occupied by someone's absence.
No people are present.
cinematic dreamlike photography,
melancholic atmosphere,
muted blue-gray tones,
soft shadows,
subtle haze,
quiet composition,
realistic textures,
film grain,
high detail
job_24b6ae02054244fc80bc.png
截屏2026-09-28 下午8.15.49.png
The room remained recognizable and structurally believable, but the overall atmosphere became more emotionally charged.
Instead of dramatically transforming the geometry, the model interpreted the emotional vocabulary primarily through environmental characteristics such as lighting, emptiness, muted colors, shadows, and composition.

Compared with P01, the largest change occurred in atmosphere rather than architecture.
This suggests that emotional language does not necessarily cause structural transformation. Instead, the model can translate emotional concepts into photographic characteristics.

P03 — Incomplete Memory
Positive Prompts
A bedroom reconstructed from a fragmented and incomplete childhood memory.
At first glance, the room looks like a real ordinary bedroom,
but parts of the memory are visibly missing.
A small wooden bed stands near the wall,
an old chair sits beside a window,
and a few familiar personal objects remain in the room.
However, several areas cannot be remembered clearly.
Parts of the furniture lose their edges and fade into pale fog.
One corner of the room is unfinished and disappears into blank white space.
Sections of the wall contain soft featureless patches,
as if visual information has been erased from memory.
Some objects are sharply realistic,
while nearby objects are only faint translucent impressions.
A bedside object appears twice in slightly different positions,
like two conflicting versions of the same memory.
The far side of the room becomes increasingly indistinct,
with details dissolving before they can be recognized.
The bedroom remains physically believable,
but visual information becomes incomplete and uncertain toward the edges.
The image should feel like a photograph reconstructed
from a memory with missing pieces,
not an abandoned room and not a ruined room.
no people,
photorealistic interior,
subtle dream logic,
selective loss of detail,
faded visual information,
translucent object traces,
soft spatial discontinuity,
pale atmospheric haze,
muted blue-gray colors,
analog photographic texture,
quiet nostalgic atmosphere
job_7f38d1c5cd984e3a833d.png
截屏2026-09-28 下午8.20.30.png
The bedroom became significantly brighter, more washed out, emptier, and less visually defined. Large areas contained reduced information and softer boundaries.
However, the bed, chair, window, walls, and overall perspective still remained coherent.

Compare with P02
Compared with P02, the change is no longer limited to mood. Visual information begins to fade and become less defined. The room appears brighter, softer, and partially washed out, while some details become difficult to recognize.
However, the architecture itself remains coherent. This suggests that the model interprets “incomplete memory” primarily as a loss of visual information rather than a distortion of physical space.

P04 — Architectural Memory
Positive Prompts
A photorealistic empty bedroom containing visible traces of its own past.
A wooden bed, a chair, and a window remain in the room.
The same chair appears twice:
one solid chair in the present,
and one large translucent duplicate slightly offset behind it.
The bed also has a faint translucent duplicate,
shifted slightly away from its current position.
Several objects exist in two overlapping positions at once,
as if the room remembers where they used to be.
Parts of the furniture leave transparent afterimages in space.
The wall contains overlapping rectangular traces
of objects that are no longer there.
Two different moments of the same bedroom
are visible simultaneously in one photograph.
The present room is solid and realistic.
The remembered room is translucent, faded, displaced, and incomplete.
No person is visible.
photorealistic interior photography,
double exposure,
long exposure afterimage,
multiple exposure photography,
translucent furniture duplicates,
overlapping spatial positions,
visual memory traces,
subtle temporal displacement,
muted blue-gray colors,
soft natural light,
analog film photography
job_7f0b6b050e114fb28f20.png
截屏2026-09-28 下午8.26.21.png
This produced a substantially different image.
A normal wooden chair appeared alongside transparent chair-like forms. Additional rectangular outlines appeared on the wall, creating visible traces of previous states of the room.
The bedroom remained photorealistic and spatially recognizable, but memory became visible as an additional layer within the physical environment.

Compare with P03
Compared with P03, memory is no longer represented simply through fading or missing information. It becomes physically visible through duplication, transparency, and displaced object traces.
Transparent chair-like forms and repeated outlines suggest previous states of the room existing alongside the present one. The space remains recognizable, but different moments begin to overlap.
Main change:
Missing memory → Visible memory traces

P05 — Spatial Collapse
Positive Prompts
A photorealistic bedroom that is physically breaking apart
because the room is forgetting its own structure.
A wooden bed remains recognizable in the center of the bedroom,
but the architecture around it has become impossible.
One wall is partially missing,
opening directly into an endless pale fog with no exterior landscape.
A second doorway appears high on the wall
where no doorway could physically exist.
The same window repeats three times
at different sizes and different positions across the room.
One corner of the bedroom bends inward
and connects impossibly to another part of the same room.
Sections of the wooden floor detach from the room
and continue vertically up the wall.
Part of the ceiling dissolves into empty white space.
Furniture near the center remains solid and realistic,
while the architecture becomes increasingly fragmented
toward the edges of the image.
Some sections of the room are completely absent,
replaced by blank atmospheric space.
The bedroom should remain recognizable,
but its normal spatial logic is visibly broken.
This is not a ruined or abandoned building.
The architecture is disappearing because the space
can no longer remember how it was constructed.
photorealistic interior photography,
impossible architecture,
fragmented spatial geometry,
repeated windows,
impossible doorway,
missing walls,
architectural discontinuity,
liminal dream space,
spatial paradox,
pale atmospheric void,
muted blue-gray palette,
soft natural light,
analog photographic texture
job_3e15613177264ca0801d.png
截屏2026-09-28 下午8.29.32.png
P05 represents the strongest transformation in the sequence.
The purpose was no longer simply to create a particular emotional atmosphere or erase visual information. Instead, the prompt challenged the model's understanding of normal architectural relationships.

Compare with P04
Compared with P04, the transformation moves from individual objects to the architecture itself. P04 preserves a stable room and introduces memory through object duplication, while P05 begins to break the normal spatial rules of the entire environment.
Architectural elements become repeated, fragmented, displaced, or physically impossible. The bedroom is still partially recognizable, but its spatial logic is no longer stable.
Main change:
Object-level distortion - Architectural-level distortion
Overall Progression
The five prompts gradually move the same bedroom farther away from physical reality:
P01: Reality
P02: Reality + Emotion
P03: Reality + Missing Information
P04: Reality + Memory Traces
P05: Broken Spatial Reality
The key difference is that the prompts progressively shift from describing what the room contains, to how the room feels, to what information is missing, to how memory becomes visible, and finally to how the architecture itself loses spatial stability.

Conclusion
This five-step experiment explored how different ways of describing the same bedroom could influence the generated image while all generation settings remained unchanged.
The results showed a gradual transition from physical reality to dreamlike spatial instability. P01 produced a realistic and structurally stable bedroom. P02 demonstrated that emotional language mainly affected atmosphere rather than physical structure. P03 showed that descriptions of incomplete memory caused visual information to fade while the architecture remained coherent. P04 translated memory into visible traces through duplication and transparency. Finally, P05 extended the transformation to the architecture itself, breaking normal spatial relationships and creating an unstable dream space.
The experiment suggests that abstract concepts such as emotion and memory do not always directly produce structural changes. The model responded more clearly when these concepts were translated into specific visual instructions, such as fading details, duplicated objects, repeated windows, or impossible architectural relationships.
Overall, the five images demonstrate that prompt writing is not only about changing the subject of an image. Different types of language can influence different visual layers, atmosphere, information, objects, and spatial structure, even when all numerical settings remain exactly the same.

jwang44
Posts: 3
Joined: Thu Sep 24, 2026 10:54 am

Re: 10:01 Text to Image | Prompt Exploration

Post by jwang44 » Mon Sep 28, 2026 7:34 pm

Main Idea: Can a text to image model show what a computer does? This includes how it wakes, processes, stores, and sees information. Or can the model only show what a computer looks like?
My background is primarily in computer architecture, so I wanted to picture the process inside of a machine which includes booting, instruction, pipeline, memory, and firmware.


Part A: More Literal Descriptions
Part1.png
Prompt: data center as a coal mine, workers carrying bits, clear information
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 20, cfg 7.0, sampler euler, scheduler normal, denoise 1.00
Result: Instead of a mine, the model produced a long corridor that looks like a subway platform or a server aisle. There are a lot of colors with red and blue dominating the frame. The "workers carrying bits" became a single person in a white suit and a hat walking away and carrying a red box, so the "bits" became a physical object rather than data. The word "coal mine" did not affect the picture much, the only influence might be the tunnel like view that it is generated in. The "clear information" had no major effect outside of a small display with some digits on it.
Part2.png
Prompt: A pen-and-ink drawing of an abandoned data center after collapse, its server racks torn loose and floating above the ruins like debris, exposed cooling pipes and cables tangled in the rubble, sharp angular lines, cross hatching, experimental architecture
Negative: watermark, text, signature, color, photograph, 3d render, smooth organic shapes, blurry
Settings: seed: 42, control after generate: fixed, steps: 8, cfg: 1.0, sampler_name: res_multistep, scheduler: simple, denoise: 1.00
Result: A black and white ink drawing which has server racks above a symmetrical corridor or racks that are damaged. There is also rubble and cables all across the floor. Almost every noun that I put in the prompt appears and the lines are straight and sharp rather than natural.
Part3.png
Prompt: computer, vision, sight, code, logic, reasoning
Settings: seed: 0, control after generate: fixed, steps: 20, cfg: 7.0, sampler_name: euler, scheduler: normal, denoise: 1.00
Result: A blurry magenta close up of what looks like a lens against a face with a bright pink circle on the right. Mainly the words "vision and sight" came through as an eye or lens. Not so much about code, logic, or reasoning is visible from the picture itself.

Part B
Part A showed that the model draws what the computers look like and not so much of what they do. So for part B I added in some more steps to get a better visualization of what a computer does:
- I used negative prompts to block hardware like cpu, processor, circuit board, computer chip, motherboard
- I changed the cfg setting from 7 to 2.5, so that is high enough that the negative prompt still contributes but low enough to avoid the Part A colors
- I set steps from 20 to 8 and decided to add "sharp focus, crisp detail" because the lower step ones looked blurry
- Each step follows a piece of information through a computer when waking up, processing, storing, seeing, and the rules of computing

Part1: Abstract keywords/baseline
Step1.png
Prompt: instructions, order, electricity, memory, awakening, logic, sharp focus, crisp detail
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00
Result: A blue page from an instruction manual with a title that says INSEFICTE INSTRUCTIONS, a header box, the number 02, and blocks of fake text. There are also lightning bolts that strike the page and a tool that touches it. Instructions from my prompt became an instruction manual and electricity became lightning. The words order, memory, awakening, logic did not seem to have as big of an impact on this.

Part2: Boot sequence as a city waking up
Step2.png
Prompt: a dark empty city whose street lights switch on one district at a time in a precise order, as if something is waking up, aerial view, sharp focus, crisp detail
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: An aerial night view of one large building surrounded by dark ground. Red lights go in lines across its roof and out along the roads, while the streets below have some more yellow. Some areas are lit and others are completely dark, and the edges have a strong blur that makes the scene look like a mini model.

Part3: Pipelining as an assembly line
Step3.png
Prompt: an endless assembly line where glowing objects pass through four stations, each station changing them slightly, many objects in different stages at the same moment, clean industrial photograph, sharp focus
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, people, workers
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: A mirrored factory corridor with two rows of yellow bulb objects on conveyor belts. There are robotic arms hanging above them, control panels on each side, and a striped floor leading to a point. There were no people, so the negative worked. But there are no four distinct stations, and every object looks the same.

Part4: Chronophotograph style
Step4.png
Prompt: chronophotograph a small bright sphere passing through four gates, each exposure crisp and distinct, overlapping but sharp, silver gelatin print
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, color, motion blur
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: An abstract image with one small white dot in the center. The frame has curved white pipe shapes on both sides that look like arches or gates. Translucent bands overlap across the frame like several items stacked together. The image is soft overall and has faint brown and blue tints.

Part5: Memory as architecture
Step5.png
Prompt: an enormous library with numbered drawers stretching to infinity, some drawers glowing as they are opened and closed in rapid order, symmetrical architectural photograph, sharp focus, deep perspective
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, books, people
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: A long place lined floor to ceiling with small drawers, like an endless catalog. Many drawer fronts have a green color, and the corridor runs into a dark end. There are no books and no people, and this is one of the clearest images of them all.


Part6: How a computer sees
Step6.png
Prompt: a flower rebuilt from thousands of tiny precise colored squares, each square a separate number, the flower still recognizable but made entirely of order, crisp photograph
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, green matrix code, robot
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: A bright pink and orange of a daisy flower. Small colored cubes are scattered among the petals, some printed with numbers and letters (2, 6, 3, Z, S), and the center of the flower is a grid of small bumps. The flower is recognizable, and there is no Matrix code or robot.

Part7: Firmware
Step7.png
Prompt: a room that looks ordinary, but under light the walls reveal faint engraved rules and instructions that control everything in the room, sharp focus, quiet, precise
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, readable words, letters, graffiti
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: An empty room with a wooden floor, where the walls are covered in large shapes like squares and diamond bands in orange, green, and teal around a dark square in the center. These are lit by strong light. There are no letters or words anywhere.


Answer to the main question: The model is able to show what computers look like right from the start, but it can only partly show what they do. Analogies with scenes like (a library or a flower) worked best, while anything that depends on time, sequence, or abstract ideas (like booting, pipelining, and firmware) was simplified into a more symmetrical picture. The model seems to be a machine for arranging things in space, but not so much for showing things as a process over time.

jintongyang
Posts: 11
Joined: Wed Oct 01, 2025 2:38 pm

Re: 10:01 Text to Image | Prompt Exploration

Post by jintongyang » Mon Sep 28, 2026 9:18 pm

I'm interested in the atmosphere of an image. How to decribe light and tone to create different feelings of a scene.

I first imported the provided json workflow into ComfyUI and ran it with the flux1-schnell-fp8 model. My initial prompt was "A doll with curly hair and a beautiful red dress, sitting at a window. Moonlight shines through the window onto the doll. The doll's eyes are shining." I revised it several times, adding tone words like "quiet cozy night, cinematic, 35mm film photography." However, the model kept producing horror-film-like images with harsh, unnatural color casts (entirely pink, blue, or black) and a lot of background noise.
temp1.png
temp2.png
To get more control, I added details such as "full face visible" (it kept generating side profiles) and "everything is clear and relatively bright" (it kept generating very dark images). I also added negatives like "noises, blurry, horror, creepy, scary, low quality, distorted face, cartoon." The images then became to the other extreme: very bright, no night mood, or the doll's face still looked distorted. Changing cfg, seed, and other KSampler settings didn't fix it.
view.png
view (3).png
view (2).png
I switched to z_image_turbo_bf16, and it immediately produced what I had imagined, with accurate light and shadow.
view (6).png
Then I wanted to switch to a horror-movie style, so I replaced all the warm words with "horror, quiet night, dark..."
view (22).png
Since most results so far were lit from the left, I tried backlighting. My first attempt: "...A bit of light from the left shines its outline, other things are totally dark in the night. A beam of light from the crack of the door projected on the doll..." The lighting didn't change.
view (24).png
When I rewrote it as "A bit of light from the left shines from its back," the backlight appeared correctly.
view (25).png
I tried the same approach with light from the floor. Simply saying the light came from below the doll made no difference, but “A doll lit from below by a flash on the floor in front of her. Harsh uplighting on her face, deep shadows above her eyes, her large shadow cast up the wall behind her…” worked well.
view (27).png
I learned that lighting only changed when I described what the light actually does, like where the shadows fall, which parts are bright, what the light source is. Simply describe where the light source is won't generate accurately.

sofiak
Posts: 2
Joined: Thu Sep 24, 2026 10:53 am

Re: 10:01 Text to Image | Prompt Exploration

Post by sofiak » Tue Sep 29, 2026 2:09 pm

Experimental Goal:
For this project, I was interested in seeing how close I could get Comfy AI to creating artwork that I would produce with my own hands. I have been working to create collages for the walls in my room, and I wanted to experiment with different compositions of vocabulary to see which were most effective in achieving what I consider to be my personal art style. I often use the word whimsical to describe my art, and I have found that AI struggles to understand the “feel” of words. Instead it requires more fragmented and direct descriptions to produce images that align with intended outcomes. These are 10 of my 20 attempts from this experiment:

1st attempt
Prompt: an abstract coastal scene, light and whimsical in feel and color
job_2e1918bc132e44e59fa5_output_2.png
2nd attempt
Prompt: a photographed abstract coastal scene, light and whimsical in feel and color.
job_c8b4c76d5e3d45578fe8.png
3rd attempt
Prompt: a photographed abstract coastal scene, light pink, blue, and yellow in color and whimsical in feel
job_87e149a3050248649b4c.png
4th attempt
Prompt: a photographed coastal scene, light pink, blue, and yellow misty tint, whimsical in feel.
job_81766ed73e9844608741_output_3.png
I realized that the AI couldn't itself determine what a whimsical feature would entail. I also wasn't properly conveying what I meant by abstract, since I meant for my final product to look more like a collage. The AI required me to determine what elements the “whimsical" and "abstract" environment I was imagining would include and give it a description myself.

7th attempt
Prompt: a paper collaged beach scene, light pink, blue, and yellow tint.
job_12bcddca0fdb4cd5beb2.png
11th attempt
Prompt: a paper collaged wildflower field, light pink, blue, coral, and yellow tint. Golden sun glow.
job_bd9d1264dab146e1ac64_output_3.png
16th attempt
Prompt: botanical stained glass. texture as if painted with watercolor. light pink, blue, coral, sage and yellow tint. golden glow. dreamlike feel. light emerging through the glass.
job_7b2001cb81b8438db3e1_output_2.png
17th attempt
Prompt: create a magical beach scene where the sun is shining towards the viewer of the image and water is sparkling. You can see faint planets in a light blue sky. There are palm trees and butterflies and seashells and wild flowers. You can see sea creatures under the water. Everything is a bright pastel color.
job_4d5f05efd5fa4e58af60_output_2.png
18th attempt
Prompt: create a magical beach scene paper collage where the sun is shining towards the viewer of the image and water is sparkling. You can see faint planets in a light blue sky. There are palm trees and butterflies and seashells and wild flowers. You can see sea creatures under the water. Everything is a bright pastel color.
job_e2787956358841b0b052_output_2.png
20th attempt
Prompt: create a magical beach and planetary scene paper collage where the sun is shining towards the viewer of the image and water is sparkling. You can see faint planets in a light blue sky. There are palm trees and butterflies and seashells and wild flowers. You can see sea creatures under the water. Everything is a bright pastel color. Make every element an unexpected yet compositionally aesthetic size.
job_b21f338f11e1479f940f_output_2.png
Conclusion:
Unlike my traditional method of craftsmanship involving the creation of art based on a certain “feel", Comfy AI requires specific and thorough instruction describing exact elements and colors you would like to see portrayed in your image to achieve said “feel”. The program is not capable of determining what a certain more complex feeling such as “whimsical” would look like. Instead you must list concrete elements associated with the intended feeling to achieve your desired outcome. Although technology will continue to advance, it brings me peace to know that Comfy AI cannot yet replicate human emotion or art through its image creation to a full extent.

jwang44
Posts: 3
Joined: Thu Sep 24, 2026 10:54 am

Re: 10:01 Text to Image | Prompt Exploration

Post by jwang44 » Fri Oct 02, 2026 1:42 pm

Update: Attached analysis of results and also key findings after tinkering more with the settings from future generations of images.

Main Idea: Can a text to image model show what a computer does? This includes how it wakes, processes, stores, and sees information. Or can the model only show what a computer looks like?
My background is primarily in computer architecture, so I wanted to picture the process inside of a machine: booting, instruction, pipeline, memory, and firmware. This covers this in 2 parts, first a more literal description (first 3 images), then a 7 step progression that tries to visualize that process with analogies, which follows a piece of information through a computer from the start and what it undergoes while the computer is processing it.


Part A: More Literal Descriptions
Part1.png
Prompt: data center as a coal mine, workers carrying bits, clear information
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 20, cfg 7.0, sampler euler, scheduler normal, denoise 1.00
Result: Instead of a mine, the model produced a long corridor that looks like a subway platform or a server aisle. There are a lot of colors with red and blue dominating the frame. The "workers carrying bits" became a single person in a white suit and a hat walking away and carrying a red box, so the "bits" became a physical object rather than data. The word "coal mine" did not affect the picture much, the only influence might be the tunnel like view that it is generated in. The "clear information" had no major effect outside of a small display with some digits on it.
Analysis: I used a metaphor prompt with labor and data. I used a cfg of 7 so it gives it more of a distorted view and it has flat red floors and oversaturated colors. The model could not mix up data centers and coal mines into one idea so it kept one of them in the form of rows of blocks and lights, and dropped the other idea.
Part2.png
Prompt: A pen-and-ink drawing of an abandoned data center after collapse, its server racks torn loose and floating above the ruins like debris, exposed cooling pipes and cables tangled in the rubble, sharp angular lines, cross hatching, experimental architecture
Negative: watermark, text, signature, color, photograph, 3d render, smooth organic shapes, blurry
Settings: seed: 42, control after generate: fixed, steps: 8, cfg: 1.0, sampler_name: res_multistep, scheduler: simple, denoise: 1.00
Result: A clean black and white ink drawing which has server racks flying above a symmetrical corridor or racks that are damaged. There is also rubble and cables all across the floor. Almost every noun that I put in the prompt appears and the lines are straight and sharp rather than natural.
Analysis: The black and white look came from the prompt "pen and ink, sharp angular lines". The style is more like a comic book illustration and seems to have a vibrant contrast between the black and white sections.
Part3.png
Prompt: computer, vision, sight, code, logic, reasoning
Settings: seed: 0, control after generate: fixed, steps: 20, cfg: 7.0, sampler_name: euler, scheduler: normal, denoise: 1.00
Result: A blurry magenta close up of what looks like a lens against a face with a bright pink arc on the right. Mainly the words "vision and sight" came through as an eye or lens. Not so much about code, logic, or reasoning is visible from the picture itself.
Analysis: I chose to mix concrete and abstract words. The model is able to draw from the concrete ones and drops the rest, because words like logic and reasoning have defined visual form. This image has the same magenta color scheme as the first image from the same cfg 7 and 20 step setting. This shows that the setting was shaping the image as much as the words that I chose to use. The more literal prompts kept showing what computers looked like and not so much of what they do. This led to the choice of choosing to remove more of the hardware nouns. I tried to describe the behavior of computers more through analogies and block the hardware in the negative prompt box.


Part B
Part A showed that the model draws what the computers look like and not so much of what they do. So for part B I added in some more steps to get a better visualization of what a computer does:
- I used negative prompts to block hardware like cpu, processor, circuit board, computer chip, motherboard
- I changed the cfg setting from 7 to 2.5, so that is high enough that the negative prompt still contributes but low enough to avoid the Part A colors
- I set steps from 20 to 8 and decided to add "sharp focus, crisp detail" because the lower step ones looked blurry
- The Schnell setup only accepted sampler euler + scheduler + normal + denoise 1. So the steps, cfg, seed, and wording were what I experimented
- Each step follows a piece of information through a computer when waking up, processing, storing, seeing, and the rules of computing

Part1: Abstract keywords/baseline
Step1.png
Prompt: instructions, order, electricity, memory, awakening, logic, sharp focus, crisp detail
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00
Result: A blue page from an instruction manual with a title that says INSEFICTE INSTRUCTIONS, a header box, the number 02, and blocks of fake text. There are also bright lightning bolts that strike the page and a tool that touches it. Instructions from my prompt became a literal instruction manual and electricity became lightning. The words order, memory, awakening, logic did not seem to have as big of an impact on this.
Analysis: Even with the hardware nouns removed, the model used the more concrete words and made a picture of their objects. Two instructions failed. Text appeared even though text was in the negative prompt because the word instructions in my prompt had more of an impact on the image. The image is also a bit out of focus on the edges. This finding changed my later steps that using a single abstract word doesn't work as well so each later step I describe a scene.

Part2: Boot sequence as a city waking up
Step2.png
Prompt: a dark empty city whose street lights switch on one district at a time in a precise order, as if something is waking up, aerial view, sharp focus, crisp detail
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: An aerial night view of one large building surrounded by dark ground. Red lights go in straight lines across its roof and out along the roads, while the streets below have some more yellow. Some areas are lit and others are completely dark, and the edges have a strong blur that makes the scene look like a mini model.
Analysis: The analogy kind of worked. The lit and unlit areas show that parts of the city are on while others are off, this is the idea of a boot sequence. But the model was not able to show the order itself, since a picture can't show one part switching on after another. The irony is that the red lines running across the roof look like traces on a circuit board so even with the circuit board in the negative prompt, the model tried to use the hardware image. The slight blur also might have come from aerial view which the model relates to mini style photograph and it had more weight than "sharp focus" from the prompt

Part3: Pipelining as an assembly line
Step3.png
Prompt: an endless assembly line where glowing objects pass through four stations, each station changing them slightly, many objects in different stages at the same moment, clean industrial photograph, sharp focus
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, people, workers
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: A mirrored factory corridor with two rows of identical yellow bulb objects on conveyor belts. There are robotic arms hanging above them, control panels on each side, and a striped floor leading to a point. There were no people, so the negative worked. But there are no four distinct stations, and every object looks the same.
Analysis: The model was able to understand assembly lines and glowing objects, but not the idea of pipelining. Each of the objects should ideally be at a different stage and look slightly different. I produced a copy of the items instead of having changes. This suggests that the model can show many things at once but not the same thing changing over time, which is what pipelining is. The mirror symmetry matches Part 1 and Part2 so this frame seems to be the model's default view for anything industrial.

Part4: Chronophotograph style
Step4.png
Prompt: chronophotograph a small bright sphere passing through four gates, each exposure crisp and distinct, overlapping but sharp, silver gelatin print
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, color, motion blur
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: An abstract image with one small white dot in the center. The frame has curved white pipe shapes on both sides that look like arches or gates. Translucent bands overlap across the whole frame like several items stacked together. The image is soft overall and has faint brown and blue tints.
Analysis: This is the most different from Step3 even though it describes the same idea. I chose to use chronophotographs to try and have a sense of change over the images like moving figures. There is a sphere that appears once and is not done through the 4 gates, so the idea of motion over time seems to not have worked again. Both color and motion blur was in the negative prompt, but the image is still slightly colored and soft, so the negative prompt was weaker than what was in the prompt. Compared to Step3, the style changed the look more than the description did but it didn't make the concept clearer.

Part5: Memory as architecture
Step5.png
Prompt: an enormous library with numbered drawers stretching to infinity, some drawers glowing as they are opened and closed in rapid order, symmetrical architectural photograph, sharp focus, deep perspective
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, books, people
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: A long corridor lined floor to ceiling with small drawers, like an endless card catalog. Many drawer fronts have a green color, and the corridor runs into a dark end. There are no books and no people, and this is one of the clearest images in the set.
Analysis: Blocking the word books worked and the library in the prompt became more of an idea of storage with many rows of drawers. This is a good picture of what memory is. The surprise for me is that the green and dark corridor makes it look like a data center. This is similar to the kind of image that was made in Part A. The model was able to match infinite number storages to computers, which matches what Parts 2-4 had where the changes could be shown in an image.

Part6: How a computer sees
Step6.png
Prompt: a flower rebuilt from thousands of tiny precise colored squares, each square a separate number, the flower still recognizable but made entirely of order, crisp macro photograph
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, green matrix code, robot
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: A bright pink and orange of a daisy flower. Small colored cubes are scattered among the petals, some printed with numbers and letters (2, 6, 3, Z, S), and the center of the flower is a dense grid of small bumps. The flower is fully recognizable, and there is no Matrix code or robot.
Analysis: The model held both readings, but not so much the prompt. Instead of a flower made of squares (which was supposed to represent pixels, it had more of a flower with toy tiles that were placed on top of it. The squares become objects added to the scene rather than what the flower is built from. The numbers are also random. Still, it is more successful to Step3 failed "vision and sight" prompt. This describes what the computer sees as numbered squares and had a more interesting image than naming the concept

Part7: Firmware
Step7.png
Prompt: a room that looks ordinary, but under raking light the walls reveal faint engraved rules and instructions that control everything in the room, sharp focus, quiet, precise
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, readable words, letters, graffiti
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: An empty room with a wooden floor, where the walls are covered in large shapes like squares and diamond bands in orange, green, cream and teal around a dark square in the center. These are lit by strong light. There are no letters or words anywhere.
Analysis: Having the negative prompt "readable words, letters" worked here unlike in Step1 with the instructions. With the text taken away, the mode used rules and instructions into more geometry. Then an unique view of the image of firmware, which is more about structure and rules rather than something you read. The image also includes a clear pattern and the model was able to use that to show order in a room that looks very simple or basic.


Report Findings:

Nouns go toward the literal objects. In Part A, the words computer, data center, and vision produced what computers looked like. Even in Part B, the instructions became an instructions manual (Step1). Describing behavior through a scene like a library of drawers, a flower of numbered squares had the strongest images (Parts 5 and 6)
The model is better at showing structure and not so much time. Booting, pipelining, and rapid memory access are all about sequence and that was harder to show in the images. Step3 gave identical objects instead of changing ones and Step4 showed the sphere once instead of at each gate.
The negative prompt can help with the image generation but is often weaker than the strong words in the positive prompt. Blocking books, people, and letters worked (Parts 3,5,7). But “text” failed when the prompt included instructions (Part1). The negative prompt also has little effect when it is at cfg 1 so the exclusions had to be part of the positive prompt in those cases
The model is able to go back to the physical representation of a computer on its own. With the circuit board in the negative prompt, the Part2 city still has some resemblance to a circuit. Also in Part5, the library looks like an aisle of a data center. Its idea of having a more “ordered, infinite structure" seems to be more focused from images of tech.
The settings that I had affected the images as much as the words. The class default of cfg 7 with 20 steps made more similar colors to the ones in Part1 and 3 regardless of the prompt that I gave. The “sharp focus” in the prompt could not override the more blurry softer focus one in Part B. The model also defaulted to using a corridor where almost half of the images came out as a symmetrical view from a one point perspective.

Answer to the main question: The model is able to show what computers look like right from the start, but it can only partly show what they do. Analogies with scenes like (a library or a flower) worked best, while anything that depends on time, sequence, or abstract ideas (like booting, pipelining, and firmware) was simplified into a more symmetrical picture. The model seems to be a machine for arranging things in space, but not so much for showing things as a process over time.

sumyeelee
Posts: 2
Joined: Thu Oct 01, 2026 7:38 pm

Re: 10:01 Text to Image | Prompt Exploration

Post by sumyeelee » Mon Oct 05, 2026 11:10 pm

ntroduction
For this project, I used ComfyAI to generate an imaginary door. I have a musical composition that features a door as an object to open new sonic worlds, and I wanted to see how ComfyAI could help me generate different door images with various themes to match my sonic imagination. However, when I entered the word "door," ComfyAI generated a normal-looking door regardless of other descriptive words I used, such as "mysterious," "avant-garde," or "sci-fi." I was looking for an image that didn't resemble an everyday, ordinary door. Below is my process of asking ComfyAI to generate what I wanted. In the later trials, I abandoned the word "door" entirely to get results closer to my imagination.

Trial 1

Positive: Generate a door that emerges from complete darkness. The door should be in an open position. It opens to a new world; this new world should see the light. It should create a mysterious vibe. More avant-garde looking door, single door.

Negative: low quality, blurry, looks like an everyday door, closed position, double doors, animated.
job_09a0d5140e374ff58e12.png
Trial 2

Positive: Generate a door floating out of complete darkness. The door should be in an open position. It opens to a new world; this new world should see the light. It should create a mysterious vibe. More avant-garde looking door, single door, imaginative, drawing-like.

Negative: low quality, blurry, looks like an everyday door, closed position, double doors, normal-looking.
job_999f8f674bc644a39bd2.png
Trial 3

Positive: Generate a door floating out of complete darkness. The door should be in an open position. Avant-garde style, sci-fi style door, highly imaginative, it should look like it doesn't exist nowadays, exists in the future.

Negative: low quality, blurry, looks like an everyday door, closed position, double doors, normal-looking.
job_9873de98e2984ddaa9af.png
Trial 4

Positive: Generate a door floating out of complete darkness. The door should be in an open position. Avant-garde style, sci-fi style door, highly imaginative, it should look like it doesn't exist nowadays, exists in the future, the door has no gravity, it's floating in the air.

Negative: low quality, blurry, looks like an everyday door, closed position, double doors, normal-looking.
job_c49f3b6e25104ecda9f6.png
Trial 5

Positive: Generate a door flying through the air, no gravity, off-balance, something you would see in a sci-fi movie. The door should be in an open position, not rectangular like a normal door.

Negative: low quality, blurry, looks like an everyday door, closed position, double doors, normal-looking.
job_dd683b22485949caae9a.png
Trial 6

Positive: Generate an entrance object that opens to another world. It should be something you might see in a sci-fi movie. It should emerge from the dark.

Negative: low quality, blurry, looks like an everyday door, closed position, double doors, normal-looking.
job_cf3a527ee5f945769800.png
Trial 7

Positive: Generate a narrow entrance object that opens to another world. It should be something you might see in a sci-fi movie. It should emerge from the dark, only the entrance, no other objects such as humans.

Negative: low quality, blurry, looks like an everyday door, closed position, double doors, normal-looking, human.
job_23dd70e7137e41298380.png
Trial 8

Positive: Generate a narrow entrance object that opens to another world. It should be something you might see in a sci-fi movie. It should emerge from the dark, only the entrance, no other objects such as humans, no painting-like texture, make it look realistic.

Negative: low quality, blurry, looks like an everyday door, closed position, double doors, normal-looking, human.
job_815b6dbf49594bd994f5.png
Trial 9

Positive: Generate a narrow entrance object that opens to another world. It should be something you might see in a sci-fi movie. It should emerge from the dark, only the entrance, no other objects such as humans, no painting-like texture, make it look real, add a door that can open and close the entrance.

Negative: low quality, blurry, looks like an everyday door, closed position, double doors, normal-looking, human.
job_66af7db4b62b4b9ea4c8.png
Trial 10

Positive: Generate a narrow entrance object that opens to another world. This entrance should be in the ocean. It should be something you might see in a sci-fi movie. It should be surrounded by seawater, only the entrance, no other objects such as humans, no painting-like texture, make it look real.

Negative: low quality, blurry, looks like an everyday door, closed position, double doors, normal-looking, human, painting.
job_9403a2b208af4987b677.png
Trial 11

Positive: Generate a narrow entrance object that opens to another world. This entrance should be in the ocean. It should be something you might see in a sci-fi movie. Only the entrance, no other objects such as humans, no painting-like texture, make it look real, add a door that can open and close the entrance, the entrance should be larger and closer to the screen.

Negative: low quality, blurry, looks like an everyday door, closed position, double doors, normal-looking, human, painting.
job_8e5513e82a944fddb320.png

Post Reply