Collect About

The ark we did not build

On provenance, type and structure in generative image models

1. What a model takes a person to be

Ask a model for a face and I am given skin. The skin is pore-less, symmetrical, young and unambiguously gendered. I ask for something else: a freckle, an asymmetry, an age. I am given the same skin again, draped more convincingly.

Hito Steyerl found the most compact formula for images of this kind: “They replace likenesses with likelinesses.”1 What a model returns is not a person but the statistical middle of every image it has gathered under that word, and those images had already been smoothed before they ever reached the internet. Retouching, campaign, filter. The model does not average a way of seeing. It averages an industry.

With categories the same process becomes easier to see. Kate Crawford and Trevor Paglen spent two years uncovering the training sets behind machine vision, and called their work an archaeology of datasets.2 ImageNet, still the most cited set of its kind, gathers millions of photographs from search engines and holiday albums, taken without the knowledge of the people in them, and sorts those people into classes asserting origin, occupation, character and morality.3 A woman in a bikini becomes a slattern, a child in sunglasses a failure. This is more than offensive. It is the claim that a person can be read off their surface.

Both follow from a single cause, and it sits one level below the charge of bias.

2. The type comes from missing provenance

An image without provenance is an image without particulars. There is no record of when it was made, or where, or by whom, under what light, at what distance, for what occasion. Such an image can be sorted by one criterion only: how it looks.

With that, the type is already decided. A vessel from Kyoto, photographed in a museum, has a date, a workshop, a material and a function; scraped from an image search it has an appearance and nothing more, and that appearance is averaged with a thousand others under the word “Asian”. A painting in a collection has an attribution, a date and a technique; in a training set it is a style. Antonio Somaini describes where this leads: what survives of historical material is decontextualised surface, atmosphere, an indeterminate pastness.4 A person from a particular city, a particular decade, with a particular life becomes a type by the same route. Not because anyone set out to typify them, but because the system had nothing else to go on.

Somaini calls the mathematical space at the centre of these models a latent space: a compressed field in which everything a model has seen is reduced to coordinates of statistical distance.5 What matters is his observation that the compression is not a side effect of scale. The latent space is built so that provenance, material and date of origin fall out of it.6 He quotes Kate Crawford:

“All the content in latent space comes from training data, so that data becomes the Weltanschauung of the model: It sets the parameters of the possible.”7

There is a practical consequence. A system that knows a thing only by its appearance can only be limited by its appearance. My work “Remains of Shadow” (64.2 × 80.4 cm, July 2026, Midjourney 8.1) shows a figure almost entirely in shadow, a strap, a little skin. Neither Midjourney nor Photoshop will extend this image at the edge; both classify it as a concern, and neither gives a reason. Whether the classification is warranted I leave open. What is certain is that the systems cannot judge it. They see an area of skin and a strap. What kind of image this is, in what format, to what end, in what tradition, is recorded nowhere, because it was never recorded in the training data. Where origin and kind were known, limits could be drawn where they belong, and the reasoning could be supplied along with them.

A figure almost entirely in shadow against a blue-grey ground, arms resting on a pale edge, synthetic image by KUENZELZELLER
KUENZELZELLER, “Remains of Shadow”, 64.2 × 80.4 cm, July 2026. Midjourney 8.1.

3. The same absence beneath the skin

Ask further and there is no bone under the skin either. There never was one.

For years the best-known form of this problem was hands. Six fingers, fused knuckles, a thumb on the wrong side. Hands now mostly come out right. The fix came from more pictures of hands, better text encoders and higher resolution, not from any model learning that a hand has twenty-seven bones. The statistics grew denser. The anatomy is still missing.

Mostly. In my work “Ông Tshian-hoh, peeling lychee” (58.2 × 81.28 cm, June 2026) a hand holds a peeled lychee beside the broken shell. That hand has two thumbs. I did not notice; a friend pointed it out long after the work was finished. The error itself is trivial. The second part is not: the surface is plausible enough that the anatomy beneath it goes unchecked, including by the person who composed the image.

Two hands in close-up, black and white. The right hand holds a peeled lychee beside the broken shell, synthetic image by KUENZELZELLER
KUENZELZELLER, “Ông Tshian-hoh, peeling lychee”, 58.2 × 81.28 cm, June 2026. Midjourney 8.1. The hand on the right has two thumbs.

Other symptoms remain. A strap disappears behind a shoulder and comes out displaced. The arm of a pair of glasses ends at the hairline. A chain lies in front of a neck and behind it at once. The clearest case came in autumn 2022, when the text-to-3D system DreamFusion was being trialled. Its objects often had several faces pointing in different directions; a squirrel with three fronts. The Google Brain researcher Ben Poole named the fault the Janus problem.8 Faces are massively over-represented in training data, backs of things barely at all. The system did not know an animal has a back.

The type and the missing bone are the same omission, once facing outward and once facing in. Collect surfaces and you get probabilities about surfaces.

4. The ark

There was another way to go about it. Instead of skimming every reachable surface of the internet in the shortest possible time, one could have collected slowly and in order. Three steps, in this sequence.

First: institutions volunteer what they already hold. Museum collections, herbaria, scientific measurement series, radiographic and scan data, catalogued objects, each item with its provenance, scale and recording conditions.

Second: photographers and technical staff systematically record what has not yet been recorded. Every kind of object, every animal, people from every region, photographed or scanned under identical conditions from four or five directions. Not the most beautiful image, the comparable one.

Third: render databases are tied to these holdings, so that a generated object consists not of images of itself but of a geometry to which images are bound.

The usual objection is that a holding of this kind would never have reached the order of magnitude at which these models begin to work at all. That assumes the order of magnitude was an achievement. It was a compensation. A comparison makes this plain. In 1999 Volker Blanz and Thomas Vetter built a morphable model from two hundred laser scans of real faces, with which a face could be turned, lit and altered, because its geometry was known.9 LAION-5B, the dataset behind Stable Diffusion, contains 5.8 billion image-text pairs.10 Two hundred measured faces yield a face that can be turned. Five point eight billion scraped ones yield a face that cannot. How much material is needed is set by the quality of the record, not the other way round.

The morphable model can only do faces, and only within the range of its two hundred. That is exactly the point: what should have been scaled was the recording, not the pile. A second example has been available for thirty years. In 1994 and 1995 the US National Library of Medicine published the Visible Human Project: two bodies willed to science, frozen and milled away in layers, 1,878 cross-sections for the man, 5,190 for the woman, with CT and MRI scans alongside, in the public domain.11 Two bodies, completely documented. No diffusion model has anything comparable.

5. Why it was not built

Because skimming costs almost nothing. It requires storage and the underpaid attention of those who have to sort out whatever comes in with it; Steyerl has described this workforce, content moderators working through the worst of the internet so that a model learns what not to show.12 A curated archive of every object, animal and person, recorded to a common standard, with consent and provenance, costs decades and staff who have to be paid.

And because opacity pays. Somaini points out that the few companies who own the dominant latent spaces are accountable to almost nobody for what those spaces are made of, and have every reason to keep it that way.13 An ordered holding carries an index with it, in both senses: a register of what it contains, and a trace leading back to the origin. A holding that never had one acquires neither after the fact.

6. What is happening now

My own work takes on the first half of this problem. I work against the data, against the trained, retouched picture of the human being. What I try to set against it are unflattered afterimages in which age, deviation and gender are treated as neutral, and to get them I bend the model as far as it will bend. This remains work inside the system I am criticising here; no other is available. Why I consider it necessary all the same I have set out in I work with words, and the position behind it in What I believe I am doing.14

The second half the industry is now taking on itself, backwards. Research into world models, systems generating navigable environments that stay stable for minutes at a time, has run into the Janus problem in enlarged form: models trained on images and video alone hold no spatial priors and lose consistency as soon as the viewing angle changes.15 The answer has been to fit geometry afterwards, by way of three-dimensional representations, camera poses, depth estimation, spatial memory. Genie 3 holds a scene together for roughly a minute, its physical stability coming out of training rather than out of any rule.16

The direction is right, the sequence is not. What is being built is a reconstruction after the fact: structure derived from images, instead of images derived from structure. A growing share of the material for it already comes from generated data. This can work, and it almost certainly will. What it does not produce is indexicality. Somaini recalls Barthes’s formula for photography, “ça a été”, it has been: the image as the trace of something that took place.17 A model computing geometry back out of its own outputs produces views of something that never stood anywhere. They will be correct. They will point at nothing.

Notes

  1. Hito Steyerl, “Mean Images”, New Left Review 140/141 (March–June 2023), p. 82.
  2. Kate Crawford and Trevor Paglen, “Excavating AI: The Politics of Training Sets for Machine Learning”, AI Now Institute, New York University, 19 September 2019, excavating.ai.
  3. Ibid., sections “Taxonomy” and “Categories”.
  4. Antonio Somaini, “Latent Spaces: AI, Art, and the Archive”, October 196 (Spring 2026), pp. 19–60, DOI: 10.1162/OCTO.a.545, section 4.1, where Somaini draws on Roland Meyer and, through him, on Fredric Jameson’s notion of pastness.
  5. Ibid., section 2.
  6. Ibid., section 4.1.
  7. Kate Crawford in “A Questionnaire on Art and Machine Learning”, October 189 (Summer 2024), p. 22; quoted in Somaini, “Latent Spaces”, n. 40.
  8. See Steyerl, “Mean Images”, pp. 84–85 and n. 2. Poole posted the finding in October 2022.
  9. Volker Blanz and Thomas Vetter, “A Morphable Model for the Synthesis of 3D Faces”, in Proceedings of SIGGRAPH '99 (ACM Press, 1999), pp. 187–194. The model rests on two hundred scans of young adults.
  10. Steyerl, “Mean Images”, p. 94.
  11. National Library of Medicine, The Visible Human Project, datasets released 1994 (male) and 1995 (female), nlm.nih.gov/research/visible. No licence has been required since 2019.
  12. Steyerl, “Mean Images”, section “The means of mean production”, pp. 90–93.
  13. Somaini, “Latent Spaces”, section 4.2.
  14. KUENZELZELLER, “I work with words”, published 10th June 2026; and “What I believe I am doing”, published 4th August 2026.
  15. See for instance “Terra: Explorable Native 3D World Model with Point Latents”, arXiv:2510.14977 (2025), and the research on geometrically consistent world models surveyed there.
  16. Google DeepMind, Genie 3, presented August 2025.
  17. Somaini, “Latent Spaces”, section 7; see Roland Barthes, La chambre claire (Paris, 1980).

Read more

"Thousands of ghosts"

Published
29th August 2026
Format
Essay

"What I believe I am doing"

Published
4th August 2026
Format
Essay

"I work with words"

Published
10th June 2026
Format
Essay
クンゼル=ゼラー쿠엔첼 첼러
Instagram Threads
Legal