10.08 Image to Image
First review the description of samplers: https://www.mat.ucsb.edu/~g.legrady/aca ... mplers.pdf and the workflow sequence: https://www.mat.ucsb.edu/~g.legrady/aca ... rkflow.pdf
The asignment is to explore the combination of text-to-image (txt2img) and image-to-image (img2img) and to see how they relate to each other.
Part of the assignment is to explore
1) either a research question, such as can comfyUI generate an image that looks like it was made by a human, as discussed in this article:
https://www.mat.ucsb.edu/~g.legrady/aca ... ffused.pdf or
2) arrive at prompts and results that are interesting, as described here: https://www.mat.ucsb.edu/~g.legrady/aca ... ngness.pdf
-----
Here is a basic workflow combining txt2img and img2img:
Here we see that the result is the same as the source image. Why? because the denoise scale is low at 0.35
Here we do see a change once the denoise is up to the max (1.0)
Here we see a different result because the checkpoint is different:
This is an example of 2 images as input, and also 2 outputs:
10.08 Image to Image
10.08 Image to Image
George Legrady
legrady@mat.ucsb.edu
legrady@mat.ucsb.edu
Re: 10.08 Image to Image
Research Question: Can a machine make a mistake on purpose? Human art often depends on flaws and accidents. Can the model produce something imperfect, or does it always smooth things into a polished "AI look"?
Original Image: Variation 1 Prompt: a drawing of a house with mistakes
Negative: watermark, text
Variation 2 Prompt: a drawing of a house with mistakes, shaky wobbly lines, crooked walls, one window much bigger than the other
Negative: watermark, text, perfect symmetry
Variation 3 Prompt: a drawing of a house with mistakes, shaky wobbly lines, crooked walls, one window much bigger than the other, a wall scribbled out and redrawn in the wrong place, the mistakes turned into part of the house, a crossed-out line becoming a staircase, neat on the left and slowly falling apart toward the right, drawn quickly by hand in pencil
Negative: watermark, text, perfect symmetry, erased, corrected, evenly messy, uniform style, polished, digital art, vector illustration, chaos everywhere
Original Image: Variation 1 Prompt: an acrylic painting of a lone tree in a field under a stormy sky, painted by hand on canvas, thick visible brushstrokes in the gray clouds, a strange purple shadow pooled beneath the tree, rough gray grass in the foreground
Negative: watermark, text
Variation 2 Prompt: an acrylic painting of a lone tree in a field under a stormy sky, painted by hand on canvas, thick visible brushstrokes in the gray clouds, a strange purple shadow pooled beneath the tree, rough gray grass in the foreground, the tree slightly lopsided with its crown leaning to one side, the field painted flat with uneven patches where the paint ran thin and the canvas weave shows through
Negative: watermark, text, smooth gradients, perfect symmetry
Variation 3 Prompt: an acrylic painting of a lone tree in a field under a stormy sky, painted by hand on canvas, thick visible brushstrokes in the gray clouds, a strange purple shadow pooled beneath the tree, rough gray grass in the foreground, the tree slightly lopsided with its crown leaning to one side, the field painted flat with uneven patches where the paint ran thin and the canvas weave shows through, the purple shadow too bright and too round like a color mistake the painter decided to keep, small specks of paint flicked across the field, grass blades painted fast in one direction with some strokes overlapping and smearing together, the clouds painted over twice with an earlier layer still showing through, a brushstroke in the sky that went too far and was left there, the hills along the horizon uneven and slightly smudged where a finger wiped the paint
Negative: watermark, text, smooth gradients, perfect symmetry, realistic shadow, corrected colors, clean edges, photorealistic, blended clouds, evenly textured, digital painting
Original Image: Variation 1 Prompt: a wildlife photograph of a mother giraffe bending down toward her baby on the golden savanna, taken a second too late, the mother already lifting her head back up and the baby looking away, their noses no longer touching, an ordinary snapshot by a tourist
Negative: watermark, text, perfect timing, noses touching, peak moment
Variation 2 Prompt: a wildlife photograph of a mother giraffe bending down toward her baby on the golden savanna, taken a second too late, the mother already lifting her head back up and the baby looking away, their noses no longer touching, an ordinary snapshot by a tourist, shot from inside a safari vehicle at a bad angle, the mother's head cut off by the top of the frame, the giraffes small and off to one side, too much empty grass, the horizon tilted, the camera settings wrong so the sky is washed out to white and the giraffes look dull and flat, the focus caught on the tall grass in the foreground leaving the giraffes soft, a little motion blur from the vehicle moving, part of the car window frame showing in the corner
Negative: watermark, text, perfect timing, noses touching, peak moment, well-framed, centered subject, full body in frame, level horizon, professional wildlife photography, correct exposure, blue sky, vivid colors, sharp giraffes, telephoto lens, shallow depth of field
Variation 3 Prompt: a wildlife photograph of a mother giraffe bending down toward her baby on the golden savanna, taken a second too late, the mother already lifting her head back up and the baby looking away, their noses no longer touching, an ordinary snapshot by a tourist, shot from inside a safari vehicle at a bad angle, the mother's head cut off by the top of the frame, the giraffes small and off to one side, too much empty grass, the horizon tilted, the camera settings wrong so the sky is washed out to white and the giraffes look dull and flat, the focus caught on the tall grass in the foreground leaving the giraffes soft, a little motion blur from the vehicle moving, part of the car window frame showing in the corner, the tender moment just missed, the baby already wandering off as the mother watches, the kind of real vacation photo someone takes in a rush and keeps because they were there and saw it happen
Negative: watermark, text, perfect timing, noses touching, peak moment, well-framed, centered subject, full body in frame, level horizon, professional wildlife photography, correct exposure, blue sky, vivid colors, sharp giraffes, telephoto lens, shallow depth of field, double exposure, light leaks, glitch, distortion, artistic effect, polished, award-winning, National Geographic, AI art, staged
Question Conclusion: Partly, but not in the way a human does.
My conclusion is that the model can show a mistake when I say it, but it can't make one. Across all three series, there were two kinds of errors. Human mistakes, the specific accidents in the original images, were almost always removed or "corrected." Machine mistakes, like extra animals, broken anatomy and stray fragments, appeared without being asked for, but they don't look like human mistakes. When the model did produce imperfection on purpose, it turned the mistake into evenly applied messiness rather than something a human might make.
House
The original drawing's wobbly lines were simulated with code, and they look stiff and zigzagged. Surprisingly, the model made the drawing look more human. In h1 and h2 the lines became light, the pencil strokes have overshoots and construction lines, like a sketch. But it also worked on the drawing. It removed the sun and tree, straightened the house, and in h2 added extra houses, so the result looks like an architect's quick study rather than a mistake. In h4 it followed the prompt literally and drew a staircase out of the "crossed-out line." The mistake became an illustration of a mistake. The whole drawing is also messy in the same way everywhere, even though "evenly messy" was in the negative prompt. The model treats mess as a texture, and not so much as part of the picture.
Tree
The most distinctive human mistake in the original painting is the bright purple shadow under the tree. The model removed it in all three variations, even when the prompt described it directly and the negative blocked "corrected colors." In t1 and t2 it replaced the simple, slightly awkward painting with dramatic, polished storm clouds and a detailed tree that looks more like a professional landscape or concept art. Only t4, with the longest prompt and most negatives, looked like an amateur painting. It had rough strokes, a flat field, a purple sky. Even there, the imperfection is more of a beginner look. The specific error the original painter made never came back.
Giraffe
This series shows the machine's own mistakes most clearly. In g1 the model multiplied the giraffes into a family of five and turned the photo into an oversaturated poster with glowing outlines. That's an error, but it isn't the human one I asked for (a mistimed shot). g3 came closest to a human mistake with the car dashboard at the bottom, the giraffes small and off to the side, the sky. It also has machine glitches like a gray rectangle floating in the sky and a baby giraffe that barely holds together. g4 is the most revealing. One flaw is the mother's stretched, oddly jointed body, which is a machine error, not a photographer's. The more I described the feeling of a snapshot, the more the model drifted back toward a polished image.
Original Image: Variation 1 Prompt: a drawing of a house with mistakes
Negative: watermark, text
Variation 2 Prompt: a drawing of a house with mistakes, shaky wobbly lines, crooked walls, one window much bigger than the other
Negative: watermark, text, perfect symmetry
Variation 3 Prompt: a drawing of a house with mistakes, shaky wobbly lines, crooked walls, one window much bigger than the other, a wall scribbled out and redrawn in the wrong place, the mistakes turned into part of the house, a crossed-out line becoming a staircase, neat on the left and slowly falling apart toward the right, drawn quickly by hand in pencil
Negative: watermark, text, perfect symmetry, erased, corrected, evenly messy, uniform style, polished, digital art, vector illustration, chaos everywhere
Original Image: Variation 1 Prompt: an acrylic painting of a lone tree in a field under a stormy sky, painted by hand on canvas, thick visible brushstrokes in the gray clouds, a strange purple shadow pooled beneath the tree, rough gray grass in the foreground
Negative: watermark, text
Variation 2 Prompt: an acrylic painting of a lone tree in a field under a stormy sky, painted by hand on canvas, thick visible brushstrokes in the gray clouds, a strange purple shadow pooled beneath the tree, rough gray grass in the foreground, the tree slightly lopsided with its crown leaning to one side, the field painted flat with uneven patches where the paint ran thin and the canvas weave shows through
Negative: watermark, text, smooth gradients, perfect symmetry
Variation 3 Prompt: an acrylic painting of a lone tree in a field under a stormy sky, painted by hand on canvas, thick visible brushstrokes in the gray clouds, a strange purple shadow pooled beneath the tree, rough gray grass in the foreground, the tree slightly lopsided with its crown leaning to one side, the field painted flat with uneven patches where the paint ran thin and the canvas weave shows through, the purple shadow too bright and too round like a color mistake the painter decided to keep, small specks of paint flicked across the field, grass blades painted fast in one direction with some strokes overlapping and smearing together, the clouds painted over twice with an earlier layer still showing through, a brushstroke in the sky that went too far and was left there, the hills along the horizon uneven and slightly smudged where a finger wiped the paint
Negative: watermark, text, smooth gradients, perfect symmetry, realistic shadow, corrected colors, clean edges, photorealistic, blended clouds, evenly textured, digital painting
Original Image: Variation 1 Prompt: a wildlife photograph of a mother giraffe bending down toward her baby on the golden savanna, taken a second too late, the mother already lifting her head back up and the baby looking away, their noses no longer touching, an ordinary snapshot by a tourist
Negative: watermark, text, perfect timing, noses touching, peak moment
Variation 2 Prompt: a wildlife photograph of a mother giraffe bending down toward her baby on the golden savanna, taken a second too late, the mother already lifting her head back up and the baby looking away, their noses no longer touching, an ordinary snapshot by a tourist, shot from inside a safari vehicle at a bad angle, the mother's head cut off by the top of the frame, the giraffes small and off to one side, too much empty grass, the horizon tilted, the camera settings wrong so the sky is washed out to white and the giraffes look dull and flat, the focus caught on the tall grass in the foreground leaving the giraffes soft, a little motion blur from the vehicle moving, part of the car window frame showing in the corner
Negative: watermark, text, perfect timing, noses touching, peak moment, well-framed, centered subject, full body in frame, level horizon, professional wildlife photography, correct exposure, blue sky, vivid colors, sharp giraffes, telephoto lens, shallow depth of field
Variation 3 Prompt: a wildlife photograph of a mother giraffe bending down toward her baby on the golden savanna, taken a second too late, the mother already lifting her head back up and the baby looking away, their noses no longer touching, an ordinary snapshot by a tourist, shot from inside a safari vehicle at a bad angle, the mother's head cut off by the top of the frame, the giraffes small and off to one side, too much empty grass, the horizon tilted, the camera settings wrong so the sky is washed out to white and the giraffes look dull and flat, the focus caught on the tall grass in the foreground leaving the giraffes soft, a little motion blur from the vehicle moving, part of the car window frame showing in the corner, the tender moment just missed, the baby already wandering off as the mother watches, the kind of real vacation photo someone takes in a rush and keeps because they were there and saw it happen
Negative: watermark, text, perfect timing, noses touching, peak moment, well-framed, centered subject, full body in frame, level horizon, professional wildlife photography, correct exposure, blue sky, vivid colors, sharp giraffes, telephoto lens, shallow depth of field, double exposure, light leaks, glitch, distortion, artistic effect, polished, award-winning, National Geographic, AI art, staged
Question Conclusion: Partly, but not in the way a human does.
My conclusion is that the model can show a mistake when I say it, but it can't make one. Across all three series, there were two kinds of errors. Human mistakes, the specific accidents in the original images, were almost always removed or "corrected." Machine mistakes, like extra animals, broken anatomy and stray fragments, appeared without being asked for, but they don't look like human mistakes. When the model did produce imperfection on purpose, it turned the mistake into evenly applied messiness rather than something a human might make.
House
The original drawing's wobbly lines were simulated with code, and they look stiff and zigzagged. Surprisingly, the model made the drawing look more human. In h1 and h2 the lines became light, the pencil strokes have overshoots and construction lines, like a sketch. But it also worked on the drawing. It removed the sun and tree, straightened the house, and in h2 added extra houses, so the result looks like an architect's quick study rather than a mistake. In h4 it followed the prompt literally and drew a staircase out of the "crossed-out line." The mistake became an illustration of a mistake. The whole drawing is also messy in the same way everywhere, even though "evenly messy" was in the negative prompt. The model treats mess as a texture, and not so much as part of the picture.
Tree
The most distinctive human mistake in the original painting is the bright purple shadow under the tree. The model removed it in all three variations, even when the prompt described it directly and the negative blocked "corrected colors." In t1 and t2 it replaced the simple, slightly awkward painting with dramatic, polished storm clouds and a detailed tree that looks more like a professional landscape or concept art. Only t4, with the longest prompt and most negatives, looked like an amateur painting. It had rough strokes, a flat field, a purple sky. Even there, the imperfection is more of a beginner look. The specific error the original painter made never came back.
Giraffe
This series shows the machine's own mistakes most clearly. In g1 the model multiplied the giraffes into a family of five and turned the photo into an oversaturated poster with glowing outlines. That's an error, but it isn't the human one I asked for (a mistimed shot). g3 came closest to a human mistake with the car dashboard at the bottom, the giraffes small and off to the side, the sky. It also has machine glitches like a gray rectangle floating in the sky and a baby giraffe that barely holds together. g4 is the most revealing. One flaw is the mother's stretched, oddly jointed body, which is a machine error, not a photographer's. The more I described the feeling of a snapshot, the more the model drifted back toward a polished image.
Re: 10.08 Image to Image
No Right Side Up: From Paper Folds to Imagined Spaces
Original Image
My starting image is an AI generated, photographic style close up of an abstract paper sculpture.
I chose this image because its scale and orientation are ambiguous. Although it resembles folded paper, its curves could also suggest walls, pathways, roofs, or terrain. This ambiguity provided a starting point for exploring how the same visual structure could become different imagined spaces.
Experiment 1: Rotating the Input
Research question: How does the orientation of an abstract input image affect the space generated by an image-to-image workflow when the text prompt remains unchanged?
I prepared four versions of the original image, labeled 0°, 90°, 180°, and 270°. Each version came from the same original image rather than from a previous generated result. I kept the positive prompt, negative prompt, model, seed, and sampling settings constant.
The basic workflow was:
Load Image → ImageScale → VAE Encode → KSampler → VAE Decode → Save Image
Data Information
Model: sd_xl_base_1.0.safetensors
Seed: 271828
Control after Generate: fixed
Steps: 28
CFG: 6.0
Sampler: dpmpp_2m
Scheduler: karras
Denoise: 0.65
Output size: 1920 × 1080
Image scaling: Bilinear, crop disabled
Positive prompt
“A surreal oil painting of a vast dream theater, sweeping ivory structures, layered balconies, crimson curtains, small groups of figures engaged in mysterious rituals, deep passages, warm amber light and cool blue shadows, rich painterly textures, dramatic spatial depth, an asymmetrical composition with quiet open areas.”
Negative prompt:
“text, watermark, logo, blurry, flat lighting”
Observations
0° — Vertical enclosure.
The layered folds became monumental curved walls. A narrow amber opening and small figures suggested an architectural space much larger than the original paper object.
90° — Open landscape.
The curves became broad, ground-like surfaces extending toward a distant horizon. Scattered figures and distant buildings created a sense of scale and movement through an open environment.
180° — Abstract surfaces.
The result remained closer to overlapping material surfaces. A circular opening and a narrow warm-colored passage appeared, but the scene offered fewer recognizable cues for human occupation.
270° — Overhead canopy.
The layered forms occupied the upper part of the image and resembled a large ceiling. An open floor, a distant doorway, and a small figure suggested a sheltered interior.
Reflection
The most striking comparison was between 90° and 270°: similar curved forms appeared to function as ground in one image and as an overhead canopy in the other. The outputs shared a palette and painterly atmosphere, while their spatial organization differed.
This first set suggests that input orientation can help redirect the generated interpretation of an ambiguous form. However, it represents one source image and one fixed seed. The square inputs were also stretched to the required 1920 × 1080 format after rotation, so the results reflect orientation within this specific resizing workflow. Further tests would be needed to determine which differences persist across seeds.
Experiment 2: Repeating the Rotation Study with a Different Seed
Research question: Which spatial interpretations remain consistent across two seeds, and which vary?
I changed the seed from 271828 to 271829 and repeated the four orientations: 0°, 90°, 180°, and 270°. I used the same original paper image for each rotation, keeping both prompts, the model, sampling settings, denoise strength, and output dimensions unchanged. Within this second set, the seed remained fixed at 271829.
This produced eight images across the two experiments, allowing each orientation to be compared directly between seeds.
Observations
0° — From curved walls to an inhabited settlement.
The dominant vertical forms remained recognizable, but the second result introduced numerous small openings, figures, and tower like structures. The relatively sparse enclosure of the first set became a densely inhabited, cave like environment.
Original Image
My starting image is an AI generated, photographic style close up of an abstract paper sculpture.
I chose this image because its scale and orientation are ambiguous. Although it resembles folded paper, its curves could also suggest walls, pathways, roofs, or terrain. This ambiguity provided a starting point for exploring how the same visual structure could become different imagined spaces.
Experiment 1: Rotating the Input
Research question: How does the orientation of an abstract input image affect the space generated by an image-to-image workflow when the text prompt remains unchanged?
I prepared four versions of the original image, labeled 0°, 90°, 180°, and 270°. Each version came from the same original image rather than from a previous generated result. I kept the positive prompt, negative prompt, model, seed, and sampling settings constant.
The basic workflow was:
Load Image → ImageScale → VAE Encode → KSampler → VAE Decode → Save Image
Data Information
Model: sd_xl_base_1.0.safetensors
Seed: 271828
Control after Generate: fixed
Steps: 28
CFG: 6.0
Sampler: dpmpp_2m
Scheduler: karras
Denoise: 0.65
Output size: 1920 × 1080
Image scaling: Bilinear, crop disabled
Positive prompt
“A surreal oil painting of a vast dream theater, sweeping ivory structures, layered balconies, crimson curtains, small groups of figures engaged in mysterious rituals, deep passages, warm amber light and cool blue shadows, rich painterly textures, dramatic spatial depth, an asymmetrical composition with quiet open areas.”
Negative prompt:
“text, watermark, logo, blurry, flat lighting”
Observations
0° — Vertical enclosure.
The layered folds became monumental curved walls. A narrow amber opening and small figures suggested an architectural space much larger than the original paper object.
90° — Open landscape.
The curves became broad, ground-like surfaces extending toward a distant horizon. Scattered figures and distant buildings created a sense of scale and movement through an open environment.
180° — Abstract surfaces.
The result remained closer to overlapping material surfaces. A circular opening and a narrow warm-colored passage appeared, but the scene offered fewer recognizable cues for human occupation.
270° — Overhead canopy.
The layered forms occupied the upper part of the image and resembled a large ceiling. An open floor, a distant doorway, and a small figure suggested a sheltered interior.
Reflection
The most striking comparison was between 90° and 270°: similar curved forms appeared to function as ground in one image and as an overhead canopy in the other. The outputs shared a palette and painterly atmosphere, while their spatial organization differed.
This first set suggests that input orientation can help redirect the generated interpretation of an ambiguous form. However, it represents one source image and one fixed seed. The square inputs were also stretched to the required 1920 × 1080 format after rotation, so the results reflect orientation within this specific resizing workflow. Further tests would be needed to determine which differences persist across seeds.
Experiment 2: Repeating the Rotation Study with a Different Seed
Research question: Which spatial interpretations remain consistent across two seeds, and which vary?
I changed the seed from 271828 to 271829 and repeated the four orientations: 0°, 90°, 180°, and 270°. I used the same original paper image for each rotation, keeping both prompts, the model, sampling settings, denoise strength, and output dimensions unchanged. Within this second set, the seed remained fixed at 271829.
This produced eight images across the two experiments, allowing each orientation to be compared directly between seeds.
Observations
0° — From curved walls to an inhabited settlement.
The dominant vertical forms remained recognizable, but the second result introduced numerous small openings, figures, and tower like structures. The relatively sparse enclosure of the first set became a densely inhabited, cave like environment.
Last edited by zixuan241 on Wed Oct 07, 2026 6:34 pm, edited 1 time in total.
Re: 10.08 Image to Image
90° — From an open landscape to a stepped gathering space.
Both results suggested an expansive, ground-oriented environment. With the second seed, the broad curves developed into interconnected stairways, platforms, and bridges. Groups of robed figures gave the scene a stronger sense of collective activity.
180° — From abstract folds to occupied slopes.
This orientation showed a particularly noticeable change. The first result remained mostly abstract, while the second placed figures across sloping surfaces and introduced architectural openings within the folds. An orientation that initially seemed less productive became a more legible imagined environment.
270° — A recurring overhead canopy.
The large layered structure continued to read as a ceiling above an open floor. However, the second seed introduced larger foreground figures and a more prominent red band, giving the space a more theatrical atmosphere.
Reflection
Across these two seeds, several broad spatial readings recurred: the 0° images emphasized vertical enclosure, the 90° images suggested traversable terrain or platforms, and the 270° images retained an overhead canopy. The arrangement of figures, openings, and architectural details varied considerably.
The 180° comparison also challenged my initial judgment. Its first result seemed less visually engaging, but its second result developed a convincing relationship between figures and surrounding surfaces. I therefore could not treat the first image as evidence that this orientation was inherently less effective.
Artistically, I preferred the second set for its richer relationships between people and architecture, particularly the stepped gathering space at 90°. However, this preference does not establish that one seed is generally better. These two sets suggest that orientation provides a recurring compositional structure, while changing the seed can substantially alter how that structure is populated and developed.
Experiment 3: Mirroring the 90° Input
Research question: Does horizontally mirroring the input produce a mirrored version of the same scene, or does it also change the scene’s details and spatial relationships?
90° input
Mirroring the 90° Input
Observations
The unmirrored result contained a dense arrangement of stairways, platforms, and groups of robed figures. A prominent warm colored curved structure extended through the left foreground, helping frame the gathering space.
In the mirrored input result, the corresponding curved structure appeared on the right. The broad composition followed the reversed direction of the input, but the smaller details were not exact reflections. The stairways were reorganized, the figures became more dispersed, and the distant architecture changed. To me, this version felt more open, with clearer separation between the people and surrounding structures.
Reflection
The comparison showed both structural continuity and local variation. The dominant curves broadly followed the mirrored input, while the generated people and architectural details were recomposed.
Artistically, I preferred the mirrored version for its more spacious composition, although the unmirrored version conveyed a stronger sense of collective activity. Mirroring offered a way to explore these different visual relationships without rewriting the prompt.
This comparison does not isolate a general left right preference in the model. Although I kept the seed fixed, the sampling noise was not itself mirrored with the input. The result therefore documents how this workflow responded to a mirrored image, rather than proving that the model consistently favors one direction.
Experiment 4: Changing the Edge Color
Research question: How does changing a small color detail in the input affect the generated scene when the prompt remains unchanged?
I locally edited the paper image to change its red edge to blue-green, preserving the surrounding paper surfaces and composition. I used the orientation labeled 270° in my rotation study and compared the result with its red edge counterpart.
The seed remained fixed at 271829. Both prompts, the model, sampling settings, denoise strength, and output dimensions remained unchanged. Notably, the positive prompt still included “crimson curtains,” creating a contrast between the edited input color and the text’s red color cue.
Observations
Both results retained a large, sweeping overhead structure above an open space. In the blue green edge version, the lower edge of this structure appeared dark blue, while a warm orange-red area remained near the center.
The change extended beyond color alone. The central warm-colored shape and the arrangement of several figures also differed, although the overall spatial composition remained recognizable.
Reflection
In this comparison, the edited edge color appeared to carry into the generated architecture without eliminating the warmer colors elsewhere. The result suggests that a small input detail can influence the image even when the prompt contains a competing color description.
However, this single pair cannot establish whether the remaining orange red area came from the prompt, the unchanged warm regions of the input, or their interaction. It also shows that local input recoloring does not necessarily produce an output in which only color changes: nearby forms and figures may be recomposed as well.
Both results suggested an expansive, ground-oriented environment. With the second seed, the broad curves developed into interconnected stairways, platforms, and bridges. Groups of robed figures gave the scene a stronger sense of collective activity.
180° — From abstract folds to occupied slopes.
This orientation showed a particularly noticeable change. The first result remained mostly abstract, while the second placed figures across sloping surfaces and introduced architectural openings within the folds. An orientation that initially seemed less productive became a more legible imagined environment.
270° — A recurring overhead canopy.
The large layered structure continued to read as a ceiling above an open floor. However, the second seed introduced larger foreground figures and a more prominent red band, giving the space a more theatrical atmosphere.
Reflection
Across these two seeds, several broad spatial readings recurred: the 0° images emphasized vertical enclosure, the 90° images suggested traversable terrain or platforms, and the 270° images retained an overhead canopy. The arrangement of figures, openings, and architectural details varied considerably.
The 180° comparison also challenged my initial judgment. Its first result seemed less visually engaging, but its second result developed a convincing relationship between figures and surrounding surfaces. I therefore could not treat the first image as evidence that this orientation was inherently less effective.
Artistically, I preferred the second set for its richer relationships between people and architecture, particularly the stepped gathering space at 90°. However, this preference does not establish that one seed is generally better. These two sets suggest that orientation provides a recurring compositional structure, while changing the seed can substantially alter how that structure is populated and developed.
Experiment 3: Mirroring the 90° Input
Research question: Does horizontally mirroring the input produce a mirrored version of the same scene, or does it also change the scene’s details and spatial relationships?
90° input
Mirroring the 90° Input
Observations
The unmirrored result contained a dense arrangement of stairways, platforms, and groups of robed figures. A prominent warm colored curved structure extended through the left foreground, helping frame the gathering space.
In the mirrored input result, the corresponding curved structure appeared on the right. The broad composition followed the reversed direction of the input, but the smaller details were not exact reflections. The stairways were reorganized, the figures became more dispersed, and the distant architecture changed. To me, this version felt more open, with clearer separation between the people and surrounding structures.
Reflection
The comparison showed both structural continuity and local variation. The dominant curves broadly followed the mirrored input, while the generated people and architectural details were recomposed.
Artistically, I preferred the mirrored version for its more spacious composition, although the unmirrored version conveyed a stronger sense of collective activity. Mirroring offered a way to explore these different visual relationships without rewriting the prompt.
This comparison does not isolate a general left right preference in the model. Although I kept the seed fixed, the sampling noise was not itself mirrored with the input. The result therefore documents how this workflow responded to a mirrored image, rather than proving that the model consistently favors one direction.
Experiment 4: Changing the Edge Color
Research question: How does changing a small color detail in the input affect the generated scene when the prompt remains unchanged?
I locally edited the paper image to change its red edge to blue-green, preserving the surrounding paper surfaces and composition. I used the orientation labeled 270° in my rotation study and compared the result with its red edge counterpart.
The seed remained fixed at 271829. Both prompts, the model, sampling settings, denoise strength, and output dimensions remained unchanged. Notably, the positive prompt still included “crimson curtains,” creating a contrast between the edited input color and the text’s red color cue.
Observations
Both results retained a large, sweeping overhead structure above an open space. In the blue green edge version, the lower edge of this structure appeared dark blue, while a warm orange-red area remained near the center.
The change extended beyond color alone. The central warm-colored shape and the arrangement of several figures also differed, although the overall spatial composition remained recognizable.
Reflection
In this comparison, the edited edge color appeared to carry into the generated architecture without eliminating the warmer colors elsewhere. The result suggests that a small input detail can influence the image even when the prompt contains a competing color description.
However, this single pair cannot establish whether the remaining orange red area came from the prompt, the unchanged warm regions of the input, or their interaction. It also shows that local input recoloring does not necessarily produce an output in which only color changes: nearby forms and figures may be recomposed as well.
Re: 10.08 Image to Image
Can ComfyUI Pass as a Human Photographer?
My question for this assignment is whether ComfyUI can generate a photo that people can't tell apart from one taken by a human. I chose the American street photographer Lee Friedlander as my model.
Three evaluation criteria from Ha et al. (2024), Organic or Diffused: Can We Distinguish Human Art from AI-generated Images?:
1. Does the whole image look like it was made with one tool, such as one person's hand or one camera?
2. Do the details make sense?
3. Does the image look too clean or too perfect?
The experiment had four stages: testing seeds, revising the prompt, trying new scenes, and trying image to image.
Stage 1: Testing seeds
based on this piece I kept the prompt fixed and tested seeds 42, 43, 44, and 45:
Prompt: "1960s black and white street photograph, two women in sunglasses holding small cameras pointing at the viewer, pedestrian walking away, New Orleans street, 35mm, Kodak Tri-X grain, harsh daylight, awkward crop, figure cut off at frame edge, snapshot aesthetic" Seed 42: The two women are almost identical: same dress, same hair, same face, as if one person had been copied. The composition is symmetrical, with the vanishing point right in the center. "Awkward crop" in the prompt had no effect, and everyone fits neatly inside the frame. The layers are tidy and separate, which is exactly what the paper means by "too clean."
Seed 43: It still doesn't look like the original. The background is too blurry; Friedlander probably shot with a small aperture. There is almost no grain, so it looks more like a digital photo converted to black and white. The signs in the background are blurred out, so you can't tell whether the text is correct. This is a common way AI avoids difficult details. Seed 44: The twin problem came back. Seed 45: This one had the worst error: the figure on the left has one body and two heads, and there is an extra hand in the middle that doesn't belong to anyone. Oddly, with the patterned dresses and people squeezed together, this image is the closest to Friedlander's crowded feel. The more complex the scene, the more likely the AI is to get bodies wrong. All four images had the same three problems: a blurry background, and no grain. Since changing the seed didn't fix them, they had to come from the prompt. I picked seed 43 for the next stage because it was the better one.
Stage 2: Revising the prompt
I kept seed 43 and rewrote the prompt. From my research, Friedlander shot black and white with a Leica 35mm and a wide-angle lens in the 1960s. His compositions are asymmetrical and are often cut up by poles and street signs. Stage 1 also showed that abstract style words like "awkward crop" don't work, so the second version uses concrete descriptions instead:
Prompt: "Gelatin silver print of a 1960s American street snapshot, shot on a Leica 35mm rangefinder with a 35mm wide-angle lens at f/11, zone focused, everything sharp from the foreground to the distant shop signs, no background blur. Two tourist women photograph the photographer at close range. In the foreground, a taller woman with short dark bouffant hair and white cat-eye sunglasses presses a small Kodak Instamatic camera against her eye. Just behind her shoulder, partly hidden, a shorter woman with round white sunglasses and a busy patterned blouse holds her own Instamatic up to her face. On the right edge, a woman in a dark dress walks away, cut off by the frame. A street pole splits the frame vertically. Crowded storefronts, hand-painted signs, parked 1960s cars, overlapping figures. Tilted, off-balance composition, harsh midday sun, deep shadows, pushed Kodak Tri-X 400 film, coarse visible grain, slight dust on the print."
The main changes were: (1) describing each woman's height, position, sunglasses, and clothes separately to avoid twins; (2) naming the camera model and saying it is pressed against her eye; (3) naming the film stock and describing the grain; (4) specifying f/11 and no background blur; (5) adding a pole splitting the frame and a figure cut off at the edge; (6) adding crowded, overlapping figures; and (7) keeping the core idea of the women photographing the photographer. I also removed "New Orleans" and "pedestrian walking away," because they kept producing the same iron balconies and the same man walking away.
Overall, the result was not bad, but the signs went wrong. The sign in the upper left reads "TRIS 400," which is "Tri-X 400" from the prompt turned into a storefront sign. The other signs are gibberish. The AI can't tell which words describe how the photo was taken and which describe what is in it. Next, I would replace the film name with a description of its look and give the signs specific words. Stage 3: Trying new scenes
Next, I generated two photos of different scenes to see whether the AI could imitate Friedlander's style more broadly, instead of just recreating one photo. They still don't look real enough. The main problem is that there isn't enough grain or noise, so they don't feel like old photos. I'm not sure whether this comes from the workflow or the model itself. This platform locks the sampler, CFG, and other settings, so the prompt is almost the only thing I can change. Can you see what makes these different from real Friedlander photos?
Stage 4: Trying image to image
Still Trying to Figure Out
My question for this assignment is whether ComfyUI can generate a photo that people can't tell apart from one taken by a human. I chose the American street photographer Lee Friedlander as my model.
Three evaluation criteria from Ha et al. (2024), Organic or Diffused: Can We Distinguish Human Art from AI-generated Images?:
1. Does the whole image look like it was made with one tool, such as one person's hand or one camera?
2. Do the details make sense?
3. Does the image look too clean or too perfect?
The experiment had four stages: testing seeds, revising the prompt, trying new scenes, and trying image to image.
Stage 1: Testing seeds
based on this piece I kept the prompt fixed and tested seeds 42, 43, 44, and 45:
Prompt: "1960s black and white street photograph, two women in sunglasses holding small cameras pointing at the viewer, pedestrian walking away, New Orleans street, 35mm, Kodak Tri-X grain, harsh daylight, awkward crop, figure cut off at frame edge, snapshot aesthetic" Seed 42: The two women are almost identical: same dress, same hair, same face, as if one person had been copied. The composition is symmetrical, with the vanishing point right in the center. "Awkward crop" in the prompt had no effect, and everyone fits neatly inside the frame. The layers are tidy and separate, which is exactly what the paper means by "too clean."
Seed 43: It still doesn't look like the original. The background is too blurry; Friedlander probably shot with a small aperture. There is almost no grain, so it looks more like a digital photo converted to black and white. The signs in the background are blurred out, so you can't tell whether the text is correct. This is a common way AI avoids difficult details. Seed 44: The twin problem came back. Seed 45: This one had the worst error: the figure on the left has one body and two heads, and there is an extra hand in the middle that doesn't belong to anyone. Oddly, with the patterned dresses and people squeezed together, this image is the closest to Friedlander's crowded feel. The more complex the scene, the more likely the AI is to get bodies wrong. All four images had the same three problems: a blurry background, and no grain. Since changing the seed didn't fix them, they had to come from the prompt. I picked seed 43 for the next stage because it was the better one.
Stage 2: Revising the prompt
I kept seed 43 and rewrote the prompt. From my research, Friedlander shot black and white with a Leica 35mm and a wide-angle lens in the 1960s. His compositions are asymmetrical and are often cut up by poles and street signs. Stage 1 also showed that abstract style words like "awkward crop" don't work, so the second version uses concrete descriptions instead:
Prompt: "Gelatin silver print of a 1960s American street snapshot, shot on a Leica 35mm rangefinder with a 35mm wide-angle lens at f/11, zone focused, everything sharp from the foreground to the distant shop signs, no background blur. Two tourist women photograph the photographer at close range. In the foreground, a taller woman with short dark bouffant hair and white cat-eye sunglasses presses a small Kodak Instamatic camera against her eye. Just behind her shoulder, partly hidden, a shorter woman with round white sunglasses and a busy patterned blouse holds her own Instamatic up to her face. On the right edge, a woman in a dark dress walks away, cut off by the frame. A street pole splits the frame vertically. Crowded storefronts, hand-painted signs, parked 1960s cars, overlapping figures. Tilted, off-balance composition, harsh midday sun, deep shadows, pushed Kodak Tri-X 400 film, coarse visible grain, slight dust on the print."
The main changes were: (1) describing each woman's height, position, sunglasses, and clothes separately to avoid twins; (2) naming the camera model and saying it is pressed against her eye; (3) naming the film stock and describing the grain; (4) specifying f/11 and no background blur; (5) adding a pole splitting the frame and a figure cut off at the edge; (6) adding crowded, overlapping figures; and (7) keeping the core idea of the women photographing the photographer. I also removed "New Orleans" and "pedestrian walking away," because they kept producing the same iron balconies and the same man walking away.
Overall, the result was not bad, but the signs went wrong. The sign in the upper left reads "TRIS 400," which is "Tri-X 400" from the prompt turned into a storefront sign. The other signs are gibberish. The AI can't tell which words describe how the photo was taken and which describe what is in it. Next, I would replace the film name with a description of its look and give the signs specific words. Stage 3: Trying new scenes
Next, I generated two photos of different scenes to see whether the AI could imitate Friedlander's style more broadly, instead of just recreating one photo. They still don't look real enough. The main problem is that there isn't enough grain or noise, so they don't feel like old photos. I'm not sure whether this comes from the workflow or the model itself. This platform locks the sampler, CFG, and other settings, so the prompt is almost the only thing I can change. Can you see what makes these different from real Friedlander photos?
Stage 4: Trying image to image
Still Trying to Figure Out
Re: 10.08 Image to Image
Part one:
For this assignment, I wanted to experiment with how I could use image to image creation on Comfy Ui to imitate the work that is carried out in product design sectors. Much of the work of a product designer involves the creation of skews in various color ways, often using programs such as photoshop. This work can be tedious and time consuming. I wanted to see if Comfy Ui was capable of producing the same outcome, in a more efficient manner.
Original Image: 1st attempt
Prompt: Premium HOKA running shoe advertising campaign. Preserve the exact shoe from the reference image, including its silhouette, proportions, color way, materials, laces, sole geometry, branding, and product details. Transform the basic stock photograph into a high-end performance footwear campaign. The shoe is the hero product, photographed in a dynamic three-quarter angle, appearing to float slightly above a textured running surface. Subtle dust and fine particles suspended around the outsole suggest movement and impact. Clean sculptural composition, dramatic directional studio lighting, crisp highlights along the shoe materials, realistic shadows, sophisticated cool-toned background with a soft gradient, energetic but minimal. Premium sports photography, contemporary HOKA campaign aesthetic, technical performance meets elevated editorial design. Extremely realistic commercial product photography, sharp product details, natural material textures, professional retouching, controlled depth of field, 8k advertising photography. Do not redesign, distort, recolor, or alter the original shoe. No additional shoes, no people, no text, no invented logos.
Negative Prompt: blurry, low quality 2nd attempt
Prompt: Premium HOKA running shoe advertising campaign. Preserve the exact shoe from the reference image, including its silhouette, proportions, color way, materials, laces, sole geometry, branding, and product details. Transform the basic stock photograph into a high-end performance footwear campaign.
Negative Prompt: blurry, low quality 3rd attempt
Prompt: Preserve the exact shoe from the reference image, including its silhouette, proportions, materials, laces, sole geometry, branding, and product details. Transform the stock photograph color way to include light blue, neon yellow, and navy.
Negative Prompt: blurry, low quality 7th attempt
Prompt: Preserve the same shoe from the reference image. Change ONLY the color way. Recolor the shoe to light blue, neon yellow, and navy, applying the new colors naturally across the existing material panels while maintaining realistic material textures, highlights, shadows, reflections, and lighting from the original image.Keep the shoe completely identical in design, silhouette, proportions, construction, materials, stitching, panel placement, laces, sole shape, tread, HOKA branding, logo placement, camera angle, and perspective. Do not redesign or modify any physical feature of the shoe. The final result should look like an authentic alternate factory color way of the exact same HOKA model. Photorealistic commercial product photography. No changes to shape, structure, branding, background, composition, or product details.
Negative Prompt: blurry, low quality, different style, different angle, same coloring Conclusion
Comfy Ui struggles to keep an image consistent when asked to only change one particular element. When I prompted the system to modify the coloring of my image without altering the rest of it, it failed to deliver desired results. Instead, a majority of the image was modified in terms of scale, angle, branding, etc, even when specifically told not to do so (regardless of denoise settings).
Part Two:
I wanted to experiment with how Comfy Ui is able to adjust the perspective of an image, specifically through the lens of a subject in motion. My goal was to convert original images into ones that looked like they were being seen from the view of someone running.
Original Image: 1st attempt
Prompt: Make this image look like its coming from the eyes of someone who is running. 4th attempt
Prompt: Make this image look like it's coming from the eyes of someone who is running. The runner is running on this trail and the image depicts what the runner is seeing from their eyes. Likely a slightly blurred scene.
Negative Prompt: text, cartoon, non-realism, people 6th attempt
Prompt: Make this image look like its coming from the eyes of someone who is running. The runner is running on this trail and the image depicts what the runner is seeing from their eyes. Likely a blurred scene. As if the runner were to be hallucinating.
Negative Prompt: text, cartoon, non-realism, people 8th attempt
Prompt: Make this image look like its coming from the eyes of someone who is running. Center of focus and gets more blurry towards the outer image. The image depicts what the runner is seeing from their eyes. Likely a blurred scene. As if the runner were to be hallucinating Conclusion
Comfy Ui is able to take an existing image and give it different meaning when paired with outside context. The program is more successful when given prompts that can be executed without exactness.
For this assignment, I wanted to experiment with how I could use image to image creation on Comfy Ui to imitate the work that is carried out in product design sectors. Much of the work of a product designer involves the creation of skews in various color ways, often using programs such as photoshop. This work can be tedious and time consuming. I wanted to see if Comfy Ui was capable of producing the same outcome, in a more efficient manner.
Original Image: 1st attempt
Prompt: Premium HOKA running shoe advertising campaign. Preserve the exact shoe from the reference image, including its silhouette, proportions, color way, materials, laces, sole geometry, branding, and product details. Transform the basic stock photograph into a high-end performance footwear campaign. The shoe is the hero product, photographed in a dynamic three-quarter angle, appearing to float slightly above a textured running surface. Subtle dust and fine particles suspended around the outsole suggest movement and impact. Clean sculptural composition, dramatic directional studio lighting, crisp highlights along the shoe materials, realistic shadows, sophisticated cool-toned background with a soft gradient, energetic but minimal. Premium sports photography, contemporary HOKA campaign aesthetic, technical performance meets elevated editorial design. Extremely realistic commercial product photography, sharp product details, natural material textures, professional retouching, controlled depth of field, 8k advertising photography. Do not redesign, distort, recolor, or alter the original shoe. No additional shoes, no people, no text, no invented logos.
Negative Prompt: blurry, low quality 2nd attempt
Prompt: Premium HOKA running shoe advertising campaign. Preserve the exact shoe from the reference image, including its silhouette, proportions, color way, materials, laces, sole geometry, branding, and product details. Transform the basic stock photograph into a high-end performance footwear campaign.
Negative Prompt: blurry, low quality 3rd attempt
Prompt: Preserve the exact shoe from the reference image, including its silhouette, proportions, materials, laces, sole geometry, branding, and product details. Transform the stock photograph color way to include light blue, neon yellow, and navy.
Negative Prompt: blurry, low quality 7th attempt
Prompt: Preserve the same shoe from the reference image. Change ONLY the color way. Recolor the shoe to light blue, neon yellow, and navy, applying the new colors naturally across the existing material panels while maintaining realistic material textures, highlights, shadows, reflections, and lighting from the original image.Keep the shoe completely identical in design, silhouette, proportions, construction, materials, stitching, panel placement, laces, sole shape, tread, HOKA branding, logo placement, camera angle, and perspective. Do not redesign or modify any physical feature of the shoe. The final result should look like an authentic alternate factory color way of the exact same HOKA model. Photorealistic commercial product photography. No changes to shape, structure, branding, background, composition, or product details.
Negative Prompt: blurry, low quality, different style, different angle, same coloring Conclusion
Comfy Ui struggles to keep an image consistent when asked to only change one particular element. When I prompted the system to modify the coloring of my image without altering the rest of it, it failed to deliver desired results. Instead, a majority of the image was modified in terms of scale, angle, branding, etc, even when specifically told not to do so (regardless of denoise settings).
Part Two:
I wanted to experiment with how Comfy Ui is able to adjust the perspective of an image, specifically through the lens of a subject in motion. My goal was to convert original images into ones that looked like they were being seen from the view of someone running.
Original Image: 1st attempt
Prompt: Make this image look like its coming from the eyes of someone who is running. 4th attempt
Prompt: Make this image look like it's coming from the eyes of someone who is running. The runner is running on this trail and the image depicts what the runner is seeing from their eyes. Likely a slightly blurred scene.
Negative Prompt: text, cartoon, non-realism, people 6th attempt
Prompt: Make this image look like its coming from the eyes of someone who is running. The runner is running on this trail and the image depicts what the runner is seeing from their eyes. Likely a blurred scene. As if the runner were to be hallucinating.
Negative Prompt: text, cartoon, non-realism, people 8th attempt
Prompt: Make this image look like its coming from the eyes of someone who is running. Center of focus and gets more blurry towards the outer image. The image depicts what the runner is seeing from their eyes. Likely a blurred scene. As if the runner were to be hallucinating Conclusion
Comfy Ui is able to take an existing image and give it different meaning when paired with outside context. The program is more successful when given prompts that can be executed without exactness.
Last edited by sofiak on Thu Oct 08, 2026 2:53 pm, edited 2 times in total.
Re: 10.08 Image to Image
Psychology: Psychologists like Daniel Berlyne suggest interest operates on an "inverted U" curve—something too simple is boring, while something too complex is confusing, so interest peaks in the middle. Paul Silvia added that for something to be interesting, it must be both complex and comprehensible; as you gain expertise, your capacity to understand grows, pushing your interest toward greater complexity.
How can we use ComfyAI to generate an image that is not too simple but also not too complex? I will generate several very complex pictures, several very simple pictures, and pictures that fall in the middle—something that is both complex and comprehensible. I will ask ComfyAI to generate pictures for each category and see if we humans agree with ComfyAI. I am also manipulating the steps for each trial, since increasing the steps makes the images more detailed and complex.
I will ask Gemini to help me finish this project. I will use the keyword 'complex' in my positive prompt and 'simple' in my negative prompt. Next, I will use the keyword 'simple' in my positive prompt and 'complex' in my negative prompt. Likewise, the keywords 'not so complex' and 'not so simple' will be used in my positive prompt section. Complex:
Positive: The Portal & Door: A colossal, arching gateway (the "Gate to Another World") made of weathered, ornately carved stone and ancient dark metal, spanning the natural gap. Massive, detailed guardian-like warrior statues are carved directly into the rock faces flanking the arch. The gateway is overgrown with glowing bioluminescent sea-flora, coral, and kelp. The entire structure is covered in intricate, complex patterns of glowing alien runes, swirling cosmic glyphs, and celestial geometric carvings.
The Portal's interior: Inside the arch is a swirling, shimmering, multi-dimensional visual vortex—a glowing gateway of ethereal, cosmic light in hues of deep purples, golds, and brilliant turquoise. It contains glimpses of a foreign cosmic reality, star-like glints, and otherworldly biomes, spilling volumetric light and sparkling particles into the ocean.
Environment & Debris: The surrounding underwater canyon is teeming with detail: an eroded seabed cluttered with ancient ruins, crumbling stone steps and archways leading toward the gate, and scattered debris. High above, natural sunlight filters through the water's surface, mixing with the powerful otherworldly glow of the portal.
Details & Life: Massive shoals of tiny, swirling bioluminescent fish and deep-sea creatures. Diverse marine life, coral, sea fans, and giant kelp. In the lower-left, tiny scuba divers and a small submersible vehicle with exploration lights approach the gateway, highlighting its immense scale.
Atmosphere & Style: A natural, cinematic, and highly detailed underwater photograph. Natural marine snow, floating particles, deep underwater shadows, complex patterns of light, volumetric scattering, and cinematic film grain.
Negative: empty, simple, bare, flat lighting, artificial, cartoonish, low-contrast, minimal detail, missing natural elements, clean, low-resolution, lack of flora, lack of fauna, empty ocean floor, smooth walls, untextured rock, dark void, missing environmental debris, bad photography, watermark
Step 10 Step 20 Step 50
Simple:
Positive: The Door: A large, unadorned single door (rectangular with a subtle arch) made of heavily weathered stone or ancient dark, natural-looking metal, fitted snugly into the rock at the canyon's end. The door has no ornate carvings, statues, or patterns—just basic flat, natural panels and a simple rustic frame. It is slightly ajar, creating a natural threshold. The Portal/The Other World: From the slightly open gap of the door, a soft, bright, otherworldly blue-white light pours out onto the sandy ocean floor, implying a transition to another world. A gentle, diffused glow and a few small rising bubbles escape from the opening, but without any magical spirals, starry cosmic effects, or glowing runes. Environment: The natural dark underwater cave and canyon walls remain rugged, bare, and unembellished. The seabed is simple, flat sand with natural ripples and occasional loose rocks—entirely free of ruins, debris, overgrown kelp, or coral. Natural sunlight gently filters through the rippling surface far above, mixing quietly with the soft, inviting glow from the door. No people, scuba divers, or submarines. Quality/Atmosphere: A natural, quiet, and immersive underwater photograph. Peaceful, uncluttered, and minimalist composition, taken with natural lighting on an underwater camera. Focus on the simple, weathered door and its soft, leaking light.
Negative: complex, cluttered, detailed, gate, giant portal, busy, crowded, decorative, ornate, carvings, patterns, runes, statues, ruins, rubble, shipwreck, people, scuba divers, submersibles, submarines, fish, sea creatures, bioluminescent plants, kelp, coral, overgrown, cosmic, galaxy, swirling, sparks, magical particles, text, signs, artificial lighting, many patterns, high-contrast, messy, noise, cinematic drama Steps: 10 Step 20 Step 50
Middle:
Positive: The Door: A sturdy, arched stone and iron door, aged and weathered, set into a natural stone frame built between the canyon walls. It features a few simple, worn carvings and subtle decorative borders—not highly ornate, but showing signs of human or ancient craft. The door is slightly open
The Details: Gentle sea plants, kelp, and small patches of coral have naturally grown around the frame and on the nearby rocks. Natural, soft bioluminescence is dotted on the surrounding flora, hinting at mystery without being cluttered. Old, crumbling stone steps lead up to the door from the sandy seabed, showing a clear, inviting path.
Atmosphere & Light: The natural, dark canyon walls from the remain prominent. Natural sunlight filters through the water’s surface from above, blending with the glowing light from the doorway. Visible but scattered underwater details: floating particles (marine snow), small bubbles, and some loose rocks on the sandy floor.
Negative:
completely empty, bare, flat lighting, artificial, cartoonish, low-contrast, minimal, smooth, untextured, no plants, no details, no scale, uninteresting, door closed tightly, pitch black, low resolution
Steps: 10
Steps: 20
Step 50
My thought: I am not sure if this is a good research direction, because everyone has different aesthetics; some people might prefer pictures that have more objects in them. However, if I shift my research direction, perhaps disregarding the positive and negative prompts and focusing on the steps instead, that might be an interesting approach. 'Steps' in ComfyUI control how many iterations the model takes to turn pure noise into a clear image. Fewer steps mean a blurrier image, while more steps mean more detail. Personally, I think 10 steps is too blurry, and 50 steps is a bit too harsh, especially in the image generated by a complex prompt. However, I do like the one that uses the middle ground more. I think it has the right amount of detail to look like an actual ocean compared to the other two.
How can we use ComfyAI to generate an image that is not too simple but also not too complex? I will generate several very complex pictures, several very simple pictures, and pictures that fall in the middle—something that is both complex and comprehensible. I will ask ComfyAI to generate pictures for each category and see if we humans agree with ComfyAI. I am also manipulating the steps for each trial, since increasing the steps makes the images more detailed and complex.
I will ask Gemini to help me finish this project. I will use the keyword 'complex' in my positive prompt and 'simple' in my negative prompt. Next, I will use the keyword 'simple' in my positive prompt and 'complex' in my negative prompt. Likewise, the keywords 'not so complex' and 'not so simple' will be used in my positive prompt section. Complex:
Positive: The Portal & Door: A colossal, arching gateway (the "Gate to Another World") made of weathered, ornately carved stone and ancient dark metal, spanning the natural gap. Massive, detailed guardian-like warrior statues are carved directly into the rock faces flanking the arch. The gateway is overgrown with glowing bioluminescent sea-flora, coral, and kelp. The entire structure is covered in intricate, complex patterns of glowing alien runes, swirling cosmic glyphs, and celestial geometric carvings.
The Portal's interior: Inside the arch is a swirling, shimmering, multi-dimensional visual vortex—a glowing gateway of ethereal, cosmic light in hues of deep purples, golds, and brilliant turquoise. It contains glimpses of a foreign cosmic reality, star-like glints, and otherworldly biomes, spilling volumetric light and sparkling particles into the ocean.
Environment & Debris: The surrounding underwater canyon is teeming with detail: an eroded seabed cluttered with ancient ruins, crumbling stone steps and archways leading toward the gate, and scattered debris. High above, natural sunlight filters through the water's surface, mixing with the powerful otherworldly glow of the portal.
Details & Life: Massive shoals of tiny, swirling bioluminescent fish and deep-sea creatures. Diverse marine life, coral, sea fans, and giant kelp. In the lower-left, tiny scuba divers and a small submersible vehicle with exploration lights approach the gateway, highlighting its immense scale.
Atmosphere & Style: A natural, cinematic, and highly detailed underwater photograph. Natural marine snow, floating particles, deep underwater shadows, complex patterns of light, volumetric scattering, and cinematic film grain.
Negative: empty, simple, bare, flat lighting, artificial, cartoonish, low-contrast, minimal detail, missing natural elements, clean, low-resolution, lack of flora, lack of fauna, empty ocean floor, smooth walls, untextured rock, dark void, missing environmental debris, bad photography, watermark
Step 10 Step 20 Step 50
Simple:
Positive: The Door: A large, unadorned single door (rectangular with a subtle arch) made of heavily weathered stone or ancient dark, natural-looking metal, fitted snugly into the rock at the canyon's end. The door has no ornate carvings, statues, or patterns—just basic flat, natural panels and a simple rustic frame. It is slightly ajar, creating a natural threshold. The Portal/The Other World: From the slightly open gap of the door, a soft, bright, otherworldly blue-white light pours out onto the sandy ocean floor, implying a transition to another world. A gentle, diffused glow and a few small rising bubbles escape from the opening, but without any magical spirals, starry cosmic effects, or glowing runes. Environment: The natural dark underwater cave and canyon walls remain rugged, bare, and unembellished. The seabed is simple, flat sand with natural ripples and occasional loose rocks—entirely free of ruins, debris, overgrown kelp, or coral. Natural sunlight gently filters through the rippling surface far above, mixing quietly with the soft, inviting glow from the door. No people, scuba divers, or submarines. Quality/Atmosphere: A natural, quiet, and immersive underwater photograph. Peaceful, uncluttered, and minimalist composition, taken with natural lighting on an underwater camera. Focus on the simple, weathered door and its soft, leaking light.
Negative: complex, cluttered, detailed, gate, giant portal, busy, crowded, decorative, ornate, carvings, patterns, runes, statues, ruins, rubble, shipwreck, people, scuba divers, submersibles, submarines, fish, sea creatures, bioluminescent plants, kelp, coral, overgrown, cosmic, galaxy, swirling, sparks, magical particles, text, signs, artificial lighting, many patterns, high-contrast, messy, noise, cinematic drama Steps: 10 Step 20 Step 50
Middle:
Positive: The Door: A sturdy, arched stone and iron door, aged and weathered, set into a natural stone frame built between the canyon walls. It features a few simple, worn carvings and subtle decorative borders—not highly ornate, but showing signs of human or ancient craft. The door is slightly open
The Details: Gentle sea plants, kelp, and small patches of coral have naturally grown around the frame and on the nearby rocks. Natural, soft bioluminescence is dotted on the surrounding flora, hinting at mystery without being cluttered. Old, crumbling stone steps lead up to the door from the sandy seabed, showing a clear, inviting path.
Atmosphere & Light: The natural, dark canyon walls from the remain prominent. Natural sunlight filters through the water’s surface from above, blending with the glowing light from the doorway. Visible but scattered underwater details: floating particles (marine snow), small bubbles, and some loose rocks on the sandy floor.
Negative:
completely empty, bare, flat lighting, artificial, cartoonish, low-contrast, minimal, smooth, untextured, no plants, no details, no scale, uninteresting, door closed tightly, pitch black, low resolution
Steps: 10
Steps: 20
Step 50
My thought: I am not sure if this is a good research direction, because everyone has different aesthetics; some people might prefer pictures that have more objects in them. However, if I shift my research direction, perhaps disregarding the positive and negative prompts and focusing on the steps instead, that might be an interesting approach. 'Steps' in ComfyUI control how many iterations the model takes to turn pure noise into a clear image. Fewer steps mean a blurrier image, while more steps mean more detail. Personally, I think 10 steps is too blurry, and 50 steps is a bit too harsh, especially in the image generated by a complex prompt. However, I do like the one that uses the middle ground more. I think it has the right amount of detail to look like an actual ocean compared to the other two.