Page 1 of 1

10.08 Image to Image

Posted: Sat Sep 19, 2026 5:17 pm
by glegrady
10.08 Image to Image

First review the description of samplers: https://www.mat.ucsb.edu/~g.legrady/aca ... mplers.pdf and the workflow sequence: https://www.mat.ucsb.edu/~g.legrady/aca ... rkflow.pdf

The asignment is to explore the combination of text-to-image (txt2img) and image-to-image (img2img) and to see how they relate to each other.

Part of the assignment is to explore
1) either a research question, such as can comfyUI generate an image that looks like it was made by a human, as discussed in this article:
https://www.mat.ucsb.edu/~g.legrady/aca ... ffused.pdf or

2) arrive at prompts and results that are interesting, as described here: https://www.mat.ucsb.edu/~g.legrady/aca ... ngness.pdf
-----
Here is a basic workflow combining txt2img and img2img:
basic_imgtoimg.png
Here we see that the result is the same as the source image. Why? because the denoise scale is low at 0.35
Screenshot 2026-09-30 at 2.35.28 PM.png
Here we do see a change once the denoise is up to the max (1.0)
Screenshot 2026-09-30 at 2.36.21 PM.png
Here we see a different result because the checkpoint is different:
Screenshot 2026-09-30 at 2.39.04 PM.png
This is an example of 2 images as input, and also 2 outputs:
2imgtoimg.png

Re: 10.08 Image to Image

Posted: Tue Oct 06, 2026 9:17 pm
by jwang44
Research Question: Can a machine make a mistake on purpose? Human art often depends on flaws and accidents. Can the model produce something imperfect, or does it always smooth things into a polished "AI look"?

Original Image:
HouseOriginal.png
Variation 1
h1.png
Prompt: a drawing of a house with mistakes
Negative: watermark, text

Variation 2
h2.png
Prompt: a drawing of a house with mistakes, shaky wobbly lines, crooked walls, one window much bigger than the other
Negative: watermark, text, perfect symmetry

Variation 3
h4.png
Prompt: a drawing of a house with mistakes, shaky wobbly lines, crooked walls, one window much bigger than the other, a wall scribbled out and redrawn in the wrong place, the mistakes turned into part of the house, a crossed-out line becoming a staircase, neat on the left and slowly falling apart toward the right, drawn quickly by hand in pencil
Negative: watermark, text, perfect symmetry, erased, corrected, evenly messy, uniform style, polished, digital art, vector illustration, chaos everywhere

Original Image:
TreeOriginal.jpg
Variation 1
t1.png
Prompt: an acrylic painting of a lone tree in a field under a stormy sky, painted by hand on canvas, thick visible brushstrokes in the gray clouds, a strange purple shadow pooled beneath the tree, rough gray grass in the foreground
Negative: watermark, text

Variation 2
t2.png
Prompt: an acrylic painting of a lone tree in a field under a stormy sky, painted by hand on canvas, thick visible brushstrokes in the gray clouds, a strange purple shadow pooled beneath the tree, rough gray grass in the foreground, the tree slightly lopsided with its crown leaning to one side, the field painted flat with uneven patches where the paint ran thin and the canvas weave shows through
Negative: watermark, text, smooth gradients, perfect symmetry

Variation 3
t4.png
Prompt: an acrylic painting of a lone tree in a field under a stormy sky, painted by hand on canvas, thick visible brushstrokes in the gray clouds, a strange purple shadow pooled beneath the tree, rough gray grass in the foreground, the tree slightly lopsided with its crown leaning to one side, the field painted flat with uneven patches where the paint ran thin and the canvas weave shows through, the purple shadow too bright and too round like a color mistake the painter decided to keep, small specks of paint flicked across the field, grass blades painted fast in one direction with some strokes overlapping and smearing together, the clouds painted over twice with an earlier layer still showing through, a brushstroke in the sky that went too far and was left there, the hills along the horizon uneven and slightly smudged where a finger wiped the paint
Negative: watermark, text, smooth gradients, perfect symmetry, realistic shadow, corrected colors, clean edges, photorealistic, blended clouds, evenly textured, digital painting


Original Image:
GiraffeOriginal.jpg
Variation 1
g1.png
Prompt: a wildlife photograph of a mother giraffe bending down toward her baby on the golden savanna, taken a second too late, the mother already lifting her head back up and the baby looking away, their noses no longer touching, an ordinary snapshot by a tourist
Negative: watermark, text, perfect timing, noses touching, peak moment

Variation 2
g3.png
Prompt: a wildlife photograph of a mother giraffe bending down toward her baby on the golden savanna, taken a second too late, the mother already lifting her head back up and the baby looking away, their noses no longer touching, an ordinary snapshot by a tourist, shot from inside a safari vehicle at a bad angle, the mother's head cut off by the top of the frame, the giraffes small and off to one side, too much empty grass, the horizon tilted, the camera settings wrong so the sky is washed out to white and the giraffes look dull and flat, the focus caught on the tall grass in the foreground leaving the giraffes soft, a little motion blur from the vehicle moving, part of the car window frame showing in the corner
Negative: watermark, text, perfect timing, noses touching, peak moment, well-framed, centered subject, full body in frame, level horizon, professional wildlife photography, correct exposure, blue sky, vivid colors, sharp giraffes, telephoto lens, shallow depth of field

Variation 3
g4.png
Prompt: a wildlife photograph of a mother giraffe bending down toward her baby on the golden savanna, taken a second too late, the mother already lifting her head back up and the baby looking away, their noses no longer touching, an ordinary snapshot by a tourist, shot from inside a safari vehicle at a bad angle, the mother's head cut off by the top of the frame, the giraffes small and off to one side, too much empty grass, the horizon tilted, the camera settings wrong so the sky is washed out to white and the giraffes look dull and flat, the focus caught on the tall grass in the foreground leaving the giraffes soft, a little motion blur from the vehicle moving, part of the car window frame showing in the corner, the tender moment just missed, the baby already wandering off as the mother watches, the kind of real vacation photo someone takes in a rush and keeps because they were there and saw it happen
Negative: watermark, text, perfect timing, noses touching, peak moment, well-framed, centered subject, full body in frame, level horizon, professional wildlife photography, correct exposure, blue sky, vivid colors, sharp giraffes, telephoto lens, shallow depth of field, double exposure, light leaks, glitch, distortion, artistic effect, polished, award-winning, National Geographic, AI art, staged

Question Conclusion: Partly, but not in the way a human does.

My conclusion is that the model can show a mistake when I say it, but it can't make one. Across all three series, there were two kinds of errors. Human mistakes, the specific accidents in the original images, were almost always removed or "corrected." Machine mistakes, like extra animals, broken anatomy and stray fragments, appeared without being asked for, but they don't look like human mistakes. When the model did produce imperfection on purpose, it turned the mistake into evenly applied messiness rather than something a human might make.

House
The original drawing's wobbly lines were simulated with code, and they look stiff and zigzagged. Surprisingly, the model made the drawing look more human. In h1 and h2 the lines became light, the pencil strokes have overshoots and construction lines, like a sketch. But it also worked on the drawing. It removed the sun and tree, straightened the house, and in h2 added extra houses, so the result looks like an architect's quick study rather than a mistake. In h4 it followed the prompt literally and drew a staircase out of the "crossed-out line." The mistake became an illustration of a mistake. The whole drawing is also messy in the same way everywhere, even though "evenly messy" was in the negative prompt. The model treats mess as a texture, and not so much as part of the picture.

Tree
The most distinctive human mistake in the original painting is the bright purple shadow under the tree. The model removed it in all three variations, even when the prompt described it directly and the negative blocked "corrected colors." In t1 and t2 it replaced the simple, slightly awkward painting with dramatic, polished storm clouds and a detailed tree that looks more like a professional landscape or concept art. Only t4, with the longest prompt and most negatives, looked like an amateur painting. It had rough strokes, a flat field, a purple sky. Even there, the imperfection is more of a beginner look. The specific error the original painter made never came back.

Giraffe
This series shows the machine's own mistakes most clearly. In g1 the model multiplied the giraffes into a family of five and turned the photo into an oversaturated poster with glowing outlines. That's an error, but it isn't the human one I asked for (a mistimed shot). g3 came closest to a human mistake with the car dashboard at the bottom, the giraffes small and off to the side, the sky. It also has machine glitches like a gray rectangle floating in the sky and a baby giraffe that barely holds together. g4 is the most revealing. One flaw is the mother's stretched, oddly jointed body, which is a machine error, not a photographer's. The more I described the feeling of a snapshot, the more the model drifted back toward a polished image.

Re: 10.08 Image to Image

Posted: Wed Oct 07, 2026 6:06 pm
by zixuan241
No Right Side Up: From Paper Folds to Imagined Spaces
job_cacb57b3cc8843218f38_副本.png
Original Image

My starting image is an AI generated, photographic style close up of an abstract paper sculpture.
I chose this image because its scale and orientation are ambiguous. Although it resembles folded paper, its curves could also suggest walls, pathways, roofs, or terrain. This ambiguity provided a starting point for exploring how the same visual structure could become different imagined spaces.

Experiment 1: Rotating the Input

Research question: How does the orientation of an abstract input image affect the space generated by an image-to-image workflow when the text prompt remains unchanged?

I prepared four versions of the original image, labeled 0°, 90°, 180°, and 270°. Each version came from the same original image rather than from a previous generated result. I kept the positive prompt, negative prompt, model, seed, and sampling settings constant.

The basic workflow was:
Load Image → ImageScale → VAE Encode → KSampler → VAE Decode → Save Image

Data Information

Model: sd_xl_base_1.0.safetensors
Seed: 271828
Control after Generate: fixed
Steps: 28
CFG: 6.0
Sampler: dpmpp_2m
Scheduler: karras
Denoise: 0.65
Output size: 1920 × 1080
Image scaling: Bilinear, crop disabled

Positive prompt

“A surreal oil painting of a vast dream theater, sweeping ivory structures, layered balconies, crimson curtains, small groups of figures engaged in mysterious rituals, deep passages, warm amber light and cool blue shadows, rich painterly textures, dramatic spatial depth, an asymmetrical composition with quiet open areas.”

Negative prompt:

“text, watermark, logo, blurry, flat lighting”

Observations
job_06d5a2648f3f4962b7ff.png
截屏2026-10-07 下午4.52.27.png
0° — Vertical enclosure.
The layered folds became monumental curved walls. A narrow amber opening and small figures suggested an architectural space much larger than the original paper object.

job_b9c18e5d03b54512a6bb.png
截屏2026-10-07 下午4.55.47.png
90° — Open landscape.
The curves became broad, ground-like surfaces extending toward a distant horizon. Scattered figures and distant buildings created a sense of scale and movement through an open environment.

job_fe5fdab69c4c465daa12.png
截屏2026-10-07 下午4.57.53.png
180° — Abstract surfaces.
The result remained closer to overlapping material surfaces. A circular opening and a narrow warm-colored passage appeared, but the scene offered fewer recognizable cues for human occupation.

job_70e37106b7aa45f583dd.png
截屏2026-10-07 下午4.59.48.png
270° — Overhead canopy.
The layered forms occupied the upper part of the image and resembled a large ceiling. An open floor, a distant doorway, and a small figure suggested a sheltered interior.

Reflection

The most striking comparison was between 90° and 270°: similar curved forms appeared to function as ground in one image and as an overhead canopy in the other. The outputs shared a palette and painterly atmosphere, while their spatial organization differed.

This first set suggests that input orientation can help redirect the generated interpretation of an ambiguous form. However, it represents one source image and one fixed seed. The square inputs were also stretched to the required 1920 × 1080 format after rotation, so the results reflect orientation within this specific resizing workflow. Further tests would be needed to determine which differences persist across seeds.

Experiment 2: Repeating the Rotation Study with a Different Seed

Research question: Which spatial interpretations remain consistent across two seeds, and which vary?

I changed the seed from 271828 to 271829 and repeated the four orientations: 0°, 90°, 180°, and 270°. I used the same original paper image for each rotation, keeping both prompts, the model, sampling settings, denoise strength, and output dimensions unchanged. Within this second set, the seed remained fixed at 271829.

This produced eight images across the two experiments, allowing each orientation to be compared directly between seeds.

Observations
job_6eec68ddfc9644b387b7.png
截屏2026-10-07 下午5.08.09.png
0° — From curved walls to an inhabited settlement.
The dominant vertical forms remained recognizable, but the second result introduced numerous small openings, figures, and tower like structures. The relatively sparse enclosure of the first set became a densely inhabited, cave like environment.

Re: 10.08 Image to Image

Posted: Wed Oct 07, 2026 6:24 pm
by zixuan241
job_d05ba907b9d14830a902.png
截屏2026-10-07 下午5.10.53.png
90° — From an open landscape to a stepped gathering space.
Both results suggested an expansive, ground-oriented environment. With the second seed, the broad curves developed into interconnected stairways, platforms, and bridges. Groups of robed figures gave the scene a stronger sense of collective activity.

job_b905638a30c5452ea5e3.png
截屏2026-10-07 下午5.20.24.png
180° — From abstract folds to occupied slopes.
This orientation showed a particularly noticeable change. The first result remained mostly abstract, while the second placed figures across sloping surfaces and introduced architectural openings within the folds. An orientation that initially seemed less productive became a more legible imagined environment.

job_8025884093b54c839c90.png
截屏2026-10-07 下午5.22.20.png
270° — A recurring overhead canopy.
The large layered structure continued to read as a ceiling above an open floor. However, the second seed introduced larger foreground figures and a more prominent red band, giving the space a more theatrical atmosphere.

Reflection

Across these two seeds, several broad spatial readings recurred: the 0° images emphasized vertical enclosure, the 90° images suggested traversable terrain or platforms, and the 270° images retained an overhead canopy. The arrangement of figures, openings, and architectural details varied considerably.

The 180° comparison also challenged my initial judgment. Its first result seemed less visually engaging, but its second result developed a convincing relationship between figures and surrounding surfaces. I therefore could not treat the first image as evidence that this orientation was inherently less effective.

Artistically, I preferred the second set for its richer relationships between people and architecture, particularly the stepped gathering space at 90°. However, this preference does not establish that one seed is generally better. These two sets suggest that orientation provides a recurring compositional structure, while changing the seed can substantially alter how that structure is populated and developed.

Experiment 3: Mirroring the 90° Input

Research question: Does horizontally mirroring the input produce a mirrored version of the same scene, or does it also change the scene’s details and spatial relationships?
job_cacb57b3cc8843218f38.png
90° input
job_cacb57b3cc8843218f38_副本.png
Mirroring the 90° Input


Observations

The unmirrored result contained a dense arrangement of stairways, platforms, and groups of robed figures. A prominent warm colored curved structure extended through the left foreground, helping frame the gathering space.

In the mirrored input result, the corresponding curved structure appeared on the right. The broad composition followed the reversed direction of the input, but the smaller details were not exact reflections. The stairways were reorganized, the figures became more dispersed, and the distant architecture changed. To me, this version felt more open, with clearer separation between the people and surrounding structures.
job_961accd45b4f4090abb0.png

Reflection

The comparison showed both structural continuity and local variation. The dominant curves broadly followed the mirrored input, while the generated people and architectural details were recomposed.

Artistically, I preferred the mirrored version for its more spacious composition, although the unmirrored version conveyed a stronger sense of collective activity. Mirroring offered a way to explore these different visual relationships without rewriting the prompt.

This comparison does not isolate a general left right preference in the model. Although I kept the seed fixed, the sampling noise was not itself mirrored with the input. The result therefore documents how this workflow responded to a mirrored image, rather than proving that the model consistently favors one direction.

Experiment 4: Changing the Edge Color

Research question: How does changing a small color detail in the input affect the generated scene when the prompt remains unchanged?
download.png
I locally edited the paper image to change its red edge to blue-green, preserving the surrounding paper surfaces and composition. I used the orientation labeled 270° in my rotation study and compared the result with its red edge counterpart.

The seed remained fixed at 271829. Both prompts, the model, sampling settings, denoise strength, and output dimensions remained unchanged. Notably, the positive prompt still included “crimson curtains,” creating a contrast between the edited input color and the text’s red color cue.

Observations

Both results retained a large, sweeping overhead structure above an open space. In the blue green edge version, the lower edge of this structure appeared dark blue, while a warm orange-red area remained near the center.

The change extended beyond color alone. The central warm-colored shape and the arrangement of several figures also differed, although the overall spatial composition remained recognizable.
job_437004cc7cc249309eb1.png

Reflection

In this comparison, the edited edge color appeared to carry into the generated architecture without eliminating the warmer colors elsewhere. The result suggests that a small input detail can influence the image even when the prompt contains a competing color description.

However, this single pair cannot establish whether the remaining orange red area came from the prompt, the unchanged warm regions of the input, or their interaction. It also shows that local input recoloring does not necessarily produce an output in which only color changes: nearby forms and figures may be recomposed as well.

Re: 10.08 Image to Image

Posted: Wed Oct 07, 2026 8:42 pm
by ruoxi_du
Can ComfyUI Pass as a Human Photographer?

My question for this assignment is whether ComfyUI can generate a photo that people can't tell apart from one taken by a human. I chose the American street photographer Lee Friedlander as my model.

Three evaluation criteria from Ha et al. (2024), Organic or Diffused: Can We Distinguish Human Art from AI-generated Images?:
1. Does the whole image look like it was made with one tool, such as one person's hand or one camera?
2. Do the details make sense?
3. Does the image look too clean or too perfect?

The experiment had four stages: testing seeds, revising the prompt, trying new scenes, and trying image to image.

Stage 1: Testing seeds
based on this piece
Dbz9zs0VAAAGZt5.jpg
I kept the prompt fixed and tested seeds 42, 43, 44, and 45:

Prompt: "1960s black and white street photograph, two women in sunglasses holding small cameras pointing at the viewer, pedestrian walking away, New Orleans street, 35mm, Kodak Tri-X grain, harsh daylight, awkward crop, figure cut off at frame edge, snapshot aesthetic"
A_s42.png
Seed 42: The two women are almost identical: same dress, same hair, same face, as if one person had been copied. The composition is symmetrical, with the vanishing point right in the center. "Awkward crop" in the prompt had no effect, and everyone fits neatly inside the frame. The layers are tidy and separate, which is exactly what the paper means by "too clean."
A_s422.png

Seed 43: It still doesn't look like the original. The background is too blurry; Friedlander probably shot with a small aperture. There is almost no grain, so it looks more like a digital photo converted to black and white. The signs in the background are blurred out, so you can't tell whether the text is correct. This is a common way AI avoids difficult details.
A_s433.png
Seed 44: The twin problem came back.
A_s444.png
Seed 45: This one had the worst error: the figure on the left has one body and two heads, and there is an extra hand in the middle that doesn't belong to anyone. Oddly, with the patterned dresses and people squeezed together, this image is the closest to Friedlander's crowded feel. The more complex the scene, the more likely the AI is to get bodies wrong.
A_s45.png
All four images had the same three problems: a blurry background, and no grain. Since changing the seed didn't fix them, they had to come from the prompt. I picked seed 43 for the next stage because it was the better one.

Stage 2: Revising the prompt

I kept seed 43 and rewrote the prompt. From my research, Friedlander shot black and white with a Leica 35mm and a wide-angle lens in the 1960s. His compositions are asymmetrical and are often cut up by poles and street signs. Stage 1 also showed that abstract style words like "awkward crop" don't work, so the second version uses concrete descriptions instead:

Prompt: "Gelatin silver print of a 1960s American street snapshot, shot on a Leica 35mm rangefinder with a 35mm wide-angle lens at f/11, zone focused, everything sharp from the foreground to the distant shop signs, no background blur. Two tourist women photograph the photographer at close range. In the foreground, a taller woman with short dark bouffant hair and white cat-eye sunglasses presses a small Kodak Instamatic camera against her eye. Just behind her shoulder, partly hidden, a shorter woman with round white sunglasses and a busy patterned blouse holds her own Instamatic up to her face. On the right edge, a woman in a dark dress walks away, cut off by the frame. A street pole splits the frame vertically. Crowded storefronts, hand-painted signs, parked 1960s cars, overlapping figures. Tilted, off-balance composition, harsh midday sun, deep shadows, pushed Kodak Tri-X 400 film, coarse visible grain, slight dust on the print."

The main changes were: (1) describing each woman's height, position, sunglasses, and clothes separately to avoid twins; (2) naming the camera model and saying it is pressed against her eye; (3) naming the film stock and describing the grain; (4) specifying f/11 and no background blur; (5) adding a pole splitting the frame and a figure cut off at the edge; (6) adding crowded, overlapping figures; and (7) keeping the core idea of the women photographing the photographer. I also removed "New Orleans" and "pedestrian walking away," because they kept producing the same iron balconies and the same man walking away.

Overall, the result was not bad, but the signs went wrong. The sign in the upper left reads "TRIS 400," which is "Tri-X 400" from the prompt turned into a storefront sign. The other signs are gibberish. The AI can't tell which words describe how the photo was taken and which describe what is in it. Next, I would replace the film name with a description of its look and give the signs specific words.
新peomot 43.png
Stage 3: Trying new scenes

Next, I generated two photos of different scenes to see whether the AI could imitate Friedlander's style more broadly, instead of just recreating one photo. They still don't look real enough. The main problem is that there isn't enough grain or noise, so they don't feel like old photos. I'm not sure whether this comes from the workflow or the model itself. This platform locks the sampler, CFG, and other settings, so the prompt is almost the only thing I can change. Can you see what makes these different from real Friedlander photos?
Screenshot 2026-10-07 at 5.08.44 PM.png
Screenshot 2026-10-07 at 5.12.23 PM.png
job_d56e241804924ad5bccc.png
job_32a600e7f28c474697df.png

Stage 4: Trying image to image
Still Trying to Figure Out
job_42f9081e139840118368.png