Update: Attached analysis of results and also key findings after tinkering more with the settings from future generations of images.
Main Idea: Can a text to image model show what a computer does? This includes how it wakes, processes, stores, and sees information. Or can the model only show what a computer looks like?
My background is primarily in computer architecture, so I wanted to picture the process inside of a machine: booting, instruction, pipeline, memory, and firmware. This covers this in 2 parts, first a more literal description (first 3 images), then a 7 step progression that tries to visualize that process with analogies, which follows a piece of information through a computer from the start and what it undergoes while the computer is processing it.
Part A: More Literal Descriptions
Prompt: data center as a coal mine, workers carrying bits, clear information
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 20, cfg 7.0, sampler euler, scheduler normal, denoise 1.00
Result: Instead of a mine, the model produced a long corridor that looks like a subway platform or a server aisle. There are a lot of colors with red and blue dominating the frame. The "workers carrying bits" became a single person in a white suit and a hat walking away and carrying a red box, so the "bits" became a physical object rather than data. The word "coal mine" did not affect the picture much, the only influence might be the tunnel like view that it is generated in. The "clear information" had no major effect outside of a small display with some digits on it.
Analysis: I used a metaphor prompt with labor and data. I used a cfg of 7 so it gives it more of a distorted view and it has flat red floors and oversaturated colors. The model could not mix up data centers and coal mines into one idea so it kept one of them in the form of rows of blocks and lights, and dropped the other idea.
Prompt: A pen-and-ink drawing of an abandoned data center after collapse, its server racks torn loose and floating above the ruins like debris, exposed cooling pipes and cables tangled in the rubble, sharp angular lines, cross hatching, experimental architecture
Negative: watermark, text, signature, color, photograph, 3d render, smooth organic shapes, blurry
Settings: seed: 42, control after generate: fixed, steps: 8, cfg: 1.0, sampler_name: res_multistep, scheduler: simple, denoise: 1.00
Result: A clean black and white ink drawing which has server racks flying above a symmetrical corridor or racks that are damaged. There is also rubble and cables all across the floor. Almost every noun that I put in the prompt appears and the lines are straight and sharp rather than natural.
Analysis: The black and white look came from the prompt "pen and ink, sharp angular lines". The style is more like a comic book illustration and seems to have a vibrant contrast between the black and white sections.
Prompt: computer, vision, sight, code, logic, reasoning
Settings: seed: 0, control after generate: fixed, steps: 20, cfg: 7.0, sampler_name: euler, scheduler: normal, denoise: 1.00
Result: A blurry magenta close up of what looks like a lens against a face with a bright pink arc on the right. Mainly the words "vision and sight" came through as an eye or lens. Not so much about code, logic, or reasoning is visible from the picture itself.
Analysis: I chose to mix concrete and abstract words. The model is able to draw from the concrete ones and drops the rest, because words like logic and reasoning have defined visual form. This image has the same magenta color scheme as the first image from the same cfg 7 and 20 step setting. This shows that the setting was shaping the image as much as the words that I chose to use. The more literal prompts kept showing what computers looked like and not so much of what they do. This led to the choice of choosing to remove more of the hardware nouns. I tried to describe the behavior of computers more through analogies and block the hardware in the negative prompt box.
Part B
Part A showed that the model draws what the computers look like and not so much of what they do. So for part B I added in some more steps to get a better visualization of what a computer does:
- I used negative prompts to block hardware like cpu, processor, circuit board, computer chip, motherboard
- I changed the cfg setting from 7 to 2.5, so that is high enough that the negative prompt still contributes but low enough to avoid the Part A colors
- I set steps from 20 to 8 and decided to add "sharp focus, crisp detail" because the lower step ones looked blurry
- The Schnell setup only accepted sampler euler + scheduler + normal + denoise 1. So the steps, cfg, seed, and wording were what I experimented
- Each step follows a piece of information through a computer when waking up, processing, storing, seeing, and the rules of computing
Part1: Abstract keywords/baseline
Prompt: instructions, order, electricity, memory, awakening, logic, sharp focus, crisp detail
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00
Result: A blue page from an instruction manual with a title that says INSEFICTE INSTRUCTIONS, a header box, the number 02, and blocks of fake text. There are also bright lightning bolts that strike the page and a tool that touches it. Instructions from my prompt became a literal instruction manual and electricity became lightning. The words order, memory, awakening, logic did not seem to have as big of an impact on this.
Analysis: Even with the hardware nouns removed, the model used the more concrete words and made a picture of their objects. Two instructions failed. Text appeared even though text was in the negative prompt because the word instructions in my prompt had more of an impact on the image. The image is also a bit out of focus on the edges. This finding changed my later steps that using a single abstract word doesn't work as well so each later step I describe a scene.
Part2: Boot sequence as a city waking up
Prompt: a dark empty city whose street lights switch on one district at a time in a precise order, as if something is waking up, aerial view, sharp focus, crisp detail
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: An aerial night view of one large building surrounded by dark ground. Red lights go in straight lines across its roof and out along the roads, while the streets below have some more yellow. Some areas are lit and others are completely dark, and the edges have a strong blur that makes the scene look like a mini model.
Analysis: The analogy kind of worked. The lit and unlit areas show that parts of the city are on while others are off, this is the idea of a boot sequence. But the model was not able to show the order itself, since a picture can't show one part switching on after another. The irony is that the red lines running across the roof look like traces on a circuit board so even with the circuit board in the negative prompt, the model tried to use the hardware image. The slight blur also might have come from aerial view which the model relates to mini style photograph and it had more weight than "sharp focus" from the prompt
Part3: Pipelining as an assembly line
Prompt: an endless assembly line where glowing objects pass through four stations, each station changing them slightly, many objects in different stages at the same moment, clean industrial photograph, sharp focus
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, people, workers
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: A mirrored factory corridor with two rows of identical yellow bulb objects on conveyor belts. There are robotic arms hanging above them, control panels on each side, and a striped floor leading to a point. There were no people, so the negative worked. But there are no four distinct stations, and every object looks the same.
Analysis: The model was able to understand assembly lines and glowing objects, but not the idea of pipelining. Each of the objects should ideally be at a different stage and look slightly different. I produced a copy of the items instead of having changes. This suggests that the model can show many things at once but not the same thing changing over time, which is what pipelining is. The mirror symmetry matches Part 1 and Part2 so this frame seems to be the model's default view for anything industrial.
Part4: Chronophotograph style
Prompt: chronophotograph a small bright sphere passing through four gates, each exposure crisp and distinct, overlapping but sharp, silver gelatin print
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, color, motion blur
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: An abstract image with one small white dot in the center. The frame has curved white pipe shapes on both sides that look like arches or gates. Translucent bands overlap across the whole frame like several items stacked together. The image is soft overall and has faint brown and blue tints.
Analysis: This is the most different from Step3 even though it describes the same idea. I chose to use chronophotographs to try and have a sense of change over the images like moving figures. There is a sphere that appears once and is not done through the 4 gates, so the idea of motion over time seems to not have worked again. Both color and motion blur was in the negative prompt, but the image is still slightly colored and soft, so the negative prompt was weaker than what was in the prompt. Compared to Step3, the style changed the look more than the description did but it didn't make the concept clearer.
Part5: Memory as architecture
Prompt: an enormous library with numbered drawers stretching to infinity, some drawers glowing as they are opened and closed in rapid order, symmetrical architectural photograph, sharp focus, deep perspective
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, books, people
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: A long corridor lined floor to ceiling with small drawers, like an endless card catalog. Many drawer fronts have a green color, and the corridor runs into a dark end. There are no books and no people, and this is one of the clearest images in the set.
Analysis: Blocking the word books worked and the library in the prompt became more of an idea of storage with many rows of drawers. This is a good picture of what memory is. The surprise for me is that the green and dark corridor makes it look like a data center. This is similar to the kind of image that was made in Part A. The model was able to match infinite number storages to computers, which matches what Parts 2-4 had where the changes could be shown in an image.
Part6: How a computer sees
Prompt: a flower rebuilt from thousands of tiny precise colored squares, each square a separate number, the flower still recognizable but made entirely of order, crisp macro photograph
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, green matrix code, robot
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: A bright pink and orange of a daisy flower. Small colored cubes are scattered among the petals, some printed with numbers and letters (2, 6, 3, Z, S), and the center of the flower is a dense grid of small bumps. The flower is fully recognizable, and there is no Matrix code or robot.
Analysis: The model held both readings, but not so much the prompt. Instead of a flower made of squares (which was supposed to represent pixels, it had more of a flower with toy tiles that were placed on top of it. The squares become objects added to the scene rather than what the flower is built from. The numbers are also random. Still, it is more successful to Step3 failed "vision and sight" prompt. This describes what the computer sees as numbered squares and had a more interesting image than naming the concept
Part7: Firmware
Prompt: a room that looks ordinary, but under raking light the walls reveal faint engraved rules and instructions that control everything in the room, sharp focus, quiet, precise
Negative: watermark, text, cpu, processor, circuit board, computer chip, motherboard, blurry, soft focus, readable words, letters, graffiti
Settings: checkpoint flux1-schnell-fp8, seed 0, control after generate fixed, steps 8, cfg 2.5, sampler euler, scheduler normal, denoise 1.00, 1344x768
Result: An empty room with a wooden floor, where the walls are covered in large shapes like squares and diamond bands in orange, green, cream and teal around a dark square in the center. These are lit by strong light. There are no letters or words anywhere.
Analysis: Having the negative prompt "readable words, letters" worked here unlike in Step1 with the instructions. With the text taken away, the mode used rules and instructions into more geometry. Then an unique view of the image of firmware, which is more about structure and rules rather than something you read. The image also includes a clear pattern and the model was able to use that to show order in a room that looks very simple or basic.
Report Findings:
Nouns go toward the literal objects. In Part A, the words computer, data center, and vision produced what computers looked like. Even in Part B, the instructions became an instructions manual (Step1). Describing behavior through a scene like a library of drawers, a flower of numbered squares had the strongest images (Parts 5 and 6)
The model is better at showing structure and not so much time. Booting, pipelining, and rapid memory access are all about sequence and that was harder to show in the images. Step3 gave identical objects instead of changing ones and Step4 showed the sphere once instead of at each gate.
The negative prompt can help with the image generation but is often weaker than the strong words in the positive prompt. Blocking books, people, and letters worked (Parts 3,5,7). But “text” failed when the prompt included instructions (Part1). The negative prompt also has little effect when it is at cfg 1 so the exclusions had to be part of the positive prompt in those cases
The model is able to go back to the physical representation of a computer on its own. With the circuit board in the negative prompt, the Part2 city still has some resemblance to a circuit. Also in Part5, the library looks like an aisle of a data center. Its idea of having a more “ordered, infinite structure" seems to be more focused from images of tech.
The settings that I had affected the images as much as the words. The class default of cfg 7 with 20 steps made more similar colors to the ones in Part1 and 3 regardless of the prompt that I gave. The “sharp focus” in the prompt could not override the more blurry softer focus one in Part B. The model also defaulted to using a corridor where almost half of the images came out as a symmetrical view from a one point perspective.
Answer to the main question: The model is able to show what computers look like right from the start, but it can only partly show what they do. Analogies with scenes like (a library or a flower) worked best, while anything that depends on time, sequence, or abstract ideas (like booting, pipelining, and firmware) was simplified into a more symmetrical picture. The model seems to be a machine for arranging things in space, but not so much for showing things as a process over time.