
Unmaking seeing. A design inquiry into refusing algorithmic standardisation
Antonella Autuori
9 March 2026
What does it mean for artificial intelligence to see, to sort, and to decide who I’m?
When classification becomes automated, whose histories, bodies, and values are folded into its logic? What might it mean to unmake classification and to design ways of seeing that refuse standardisation, opening space for ambiguity and multiplicity?

00. One, No One and One Hundred Thousand
I’m Antonella Autuori.
But that is only one of the many names through which I exist.
My mother calls me Ella. She only discovered two years ago that I had always wished someone would choose this name for me, instead of the predictable and convenient “Anto.” For the past two years, she has called me nothing else but Ella.

For my brother and some of my cousins, I am Toony. I no longer remember how this nickname came into being, but over the years it functioned as a kind of operational password within our family, an informal authentication code layered with its own quiet security logic.

For my partner, I am Nene, a nickname that emerged naturally in the intimacy of everyday life. It stands for something little to care about and it exists only within that relational space. Outside of it, it dissolves.

For others, I am Tonia, a name that appeared within a friendship as my behavioural nuances shifted and multiplied. It feels slightly sharper, slightly more assertive.

And for some, unfortunately and rather quickly, I am simply Anto. A shortcut. The most immediate and conventional reduction.

Each of these names refers to me and coexists within me.
Each holds a fragment, a memory, a relation, a situated version of myself. None of them is false. None of them is entirely true. None of them is sufficient.
They are multiple, at times contradictory, yet they share the same interior space, returning its irreducible complexity.
As Antonella, Ella, Toony, Nene, Tonia and Anto, I am a PhD student in Design at RMIT School of Design in Melbourne and at the SUPSI Institute of Design in Switzerland. I am also a visual communicator. I am also deeply passionate about food and wine. I am not sure whether I am more fascinated by the semiotics of wine labels, by the shared space of reflection that a bottle can open between people, or simply by the pleasure of a good glass of wine.
I am also an information designer. I am drawn to data in their most disordered, indisciplined, and organic states, before they are cleaned, structured, and formatted into functional and-or beautiful infographics.
I am also a information design teacher. What I find most rewarding is the moment when students fall in love with data not because they can impose order on them, but because they learn something through engaging with their messiness.
And finally, or perhaps only finally for now, I am also a white, heterosexual woman from the Global North. I am aware that this position shapes how I see, know, and interpret the world. Growing up in Southern Italy, within contexts marked by rigid gender norms and persistent socio-economic disparities, sharpened my sensitivity to the ways identities, labour, and agency are never neutral, but culturally and historically produced. This background informs a way of working grounded in plurality and situated knowledge, and in a sustained critique of technological systems that stabilise and confine identity within predefined regimes of knowledge.
I am one. I am no one. I am one hundred thousand.
I’m irreducible to a single category. I can not collapse into a hierarchical system of knowledge.00.1 Me and-with-through-against technology
I have always been fascinated by technology. Not for its promise of efficiency or control, but for the moments when it hesitates, breaks, or behaves in unexpected ways. It is often in these moments of friction that technical systems become legible. Working through misuse, breakdown, and experimentation becomes a way of engaging with technology that brings to the surface what is usually taken for granted, making it possible to sense how systems shape perception, behaviour, and meaning.

In 2023, with the widespread introduction of generative artificial intelligence systems, and in particular the possibility of producing images from text, I began to explore the regimes of representation embedded in these machines, starting from the female body. The mechanism was straightforward. Write a prompt. Receive an image. But what kinds of visual histories were being activated behind that simplicity. What would emerge if the machine were asked to visualise what has long been silenced?
Out of this inquiry emerged Female Pleasure (2023). The project examines how the historical repression of female pleasure continues to be operationalised within contemporary visual generation systems. Using Adobe’s generative suite, I prompted the model with phrases such as “woman pleasure,” “female pleasure,” and “woman self-pleasure.” The distortions that emerged were not surprising.

They were structurally consistent with the visual cultures on which these systems are trained. Pleasure was frequently displaced into metaphor. It appeared as food, as abstraction, as flowers or fruit, or as decorative elements that subtly redirected attention away from embodied experience. When bodies were rendered, they often appeared fragmented, stylised, or flattened into aesthetic tropes. The system did not invent these associations. It recomposed them from a visual archive where female pleasure has long been mediated, euphemised, aestheticised, or made palatable. What surfaced was less a failure of the algorithm than a reflection of the cultural conditions embedded within its training data. Where is the beauty of our pleasure?

The problem with technologically mediated representation does not begin with what today seems to shout from the outputs of generative AI systems. Those images are not autonomous inventions. They are just accumulations. They are the agglomeration of what these systems have ingested from web browsers, image repositories, and our collective visual habits.
If you search for “vulva” on Google Images, a recurring pattern appears. Diagrams. Infographics. The same white-rose vulva, perfectly pink, symmetrical, hygienic, surrounded by neat lines and explanatory labels emerging from every side. A standardised anatomy. A clean version. A pedagogical version.
Girls, I am speaking to you. Do we really all have the same vulva? Is it always perfectly pink, perfectly symmetrical, annotated and diagram-ready?

The Impossible Anatomy (2024) emerges from this repetition. In this work, artificial intelligence reveals itself in two intertwined ways. On the one hand, there is the impossibility of representing the female sexual apparatus in its lived, embodied variability. On the other, there is the impossibility of reading it within the very images that claim to depict it. When such images are processed through machine vision systems, they sometimes return labels such as “banana” or “cake.” The anatomical becomes edible. The intimate becomes dessert. The body is reclassified through resemblance rather than recognition.

By 2025, my relationship with technology had become increasingly personal. The investigation moved inward. I began placing myself at the centre of the exploration.
How do machines see me?

Classificatory Machines (2025) is an ongoing experiment born from the desire to gather multiple answers to these questions rather than settle for a single one. My image became input data for several automated vision systems, including Clarifai, Google Vision API, Imagga and Microsoft Azure. What came back were not interpretations, but probabilities. Percentages. Decimal points. Scaffoldings of labels, categories, emotional states, qualitative adjectives translated into numerical confidence scores.
I am 13.68% “sexy” according to Imagga and 0.86756426 according to Clarifai.
What does 13.68 percent sexy even mean? Which fragment of my body has been converted into that number? The curve of a glass, a patch of skin, a configuration of pixels? Or am I simply being measured against a template of desirability that was defined long before I entered the frame?
In this translation, I become measurable. Fragmented. Quantified without a reason. I am distributed across decimals, broken down into percentages that suggest precision without offering understanding. The numbers imply objectivity, but they do not explain anything. They assign confidence scores, not meaning. They simulate certainty, yet remain detached from context, relation, and experience. The decimal promises accuracy. It delivers abstraction.
And yet, none of those numbers explains who I am.

recognition is never innocent it scans / it extracts / it reduces bodies flattened into receipts labels stamped as truth | sexy / attractive / other each word a cut each tag a cage each receipt a fiction of value | to be seen is to be sorted to be measured is to be sold classification becomes currency perception becomes commodity | vision is never neutral not seeing but seizing not reading but rewriting a politics of desire a machinery of control
00.2 My last ongoing journey
I began my PhD at the RMIT University of Melbourne in 2024; at the same time, generative AI entered mainstream discourse through a sudden, obvious exposure of bias. Text-to-image models were widely criticised for producing stereotypical representations of people, professions, and identities1 [1][2][3]. At the time, almost everyone, myself included, seemed genuinely shocked and quick to complain about how these systems were generating biased representations of both human and non-human actors.

Yet this reaction was, in many ways, misleading. To frame the issue as a malfunction was to misunderstand the system. This was not simply a problem of AI functionality. Artificial intelligence does not exist in a vacuum [4]. These systems were doing little more than reflecting the stereotypes that Western societies have reinforced for centuries and have continuously fed into digital platforms, image repositories, and online archives for decades. In that sense, they were not failing. They were operating exactly as designed.
To speak about biased AI outputs, then, is to move beyond the surface of the image and toward the structures that make it possible. It is to encounter datasets. Datasets are not neutral collections of information; they function as the worldview of AI models. They are infrastructures of knowledge. Through selection, labelling, and categorisation, they encode particular ways of seeing the world, deciding what matters, what can be compared, and what is rendered irrelevant [5][14] .
And once datasets, understood as archives and repositories of heterogeneous things enter a socio-cultural-technical system, classification inevitably follows.

Classification is not an abstract or purely technical operation. It is a deeply human activity. We categorise constantly, when we decide what we like or dislike, what we want to watch, how we combine clothes, where we store plates and glasses, or why we separate white laundry from coloured clothes. Categorisation helps us navigate complexity, but it is never a neutral act. It is shaped by politics, habits, culture, memory, and desire [16]. For this and many other reasons, classification cannot be reduced to a pursuit of objectivity [6]. It should remain open to subjective deliberation, negotiation, and contestation.
In artificial intelligence, however, classification becomes formalised, stabilised, and scaled. Through categories, labels, and taxonomies, datasets define what becomes visible, comparable, and meaningful for machines [15]. By the time something becomes data, it has already been classified. Someone has decided what it is, and what it is not. As Pierre Bourdieu reminds us.
“There’s a kind of sorcery that goes into the creation of categories. To create a category or to name things is to divide an almost infinitely complex universe into separate phenomena. To impose order onto an undifferentiated mass, to ascribe phenomena to a category–that is, to name a thing—is in turn a means of reifying the existence of that category.”
Pierre Bourdieu. 1992. Language and Symbolic Power (new ed.). Blackwell Publishers, Cambridge.
This logic has deep historical roots. From Lombroso’s physiognomy to contemporary computer vision systems, classification has repeatedly linked appearance to moral, social, and political judgment. Generative AI does not depart from this history. It intensifies it. What changes is scale. Classificatory decisions become infrastructural, embedded within models that quietly govern what can be generated, recognised, and valued.

A key moment where this becomes operationalised in the building of AI systems is the practice of annotation. Annotation is the process of adding labels, descriptions, or markings to data such as images, text, or audio, so that AI systems can learn to recognise patterns and supposedly make sense of the world.
In contemporary AI practices, annotation is rarely an abstract computational process. It is performed by human workers, often outsourced through platforms such as Amazon Mechanical Turk [7]. Thousands of images are labelled every day from private homes, under conditions where speed frequently determines payment. Annotation becomes piecework. The more images processed, the more one earns. To annotate data, human annotators must follow fixed instructions through rigid interface and their interpretive labour is constrained by fixed ontologies and limited options [18]. Some are asked to rate how sad a face is on a scale from one to seven, as if emotional complexity could be reduced to a single number, as documented in Cleaning Emotional Data (2020) by Elisa Giardina Papa [8].
Others draw bounding boxes around objects and assign universal labels such as “car” or “person” compressing visual ambiguity into standardised categories.

Others, working through interfaces such as those used for COCO Captions, are asked to describe the “important parts” of an image. The instruction may appear straightforward, even self-evident. Yet it raises questions that remain largely underexplored and unanswered. Important for whom? According to which worldview? And above all, what does “important parts” actually mean?

And in some cases, annotators are asked to make arbitrary and binary decisions, such as choosing which of two faces appears “more straight.” This is not a fictional example. It was an actual interface used at Stanford to train a model that claimed to predict whether someone is gay or not. (OK, WAIT A MOMENT, DO WE REALLY NEED AN AI MODEL FOR THIS????)

Across these tasks, the underlying logic remains the same. Limited ontologies, predefined taxonomies, fixed instructions, and the systematic elimination of ambiguity and plurality.
Annotation becomes a site where complexity is flattened to fit computational requirements.
Through labelled data, machines learn what counts as a person, a home, or even as “beauty”. Now, talking about categories and annotations, one of the core assumptions and also pitfalls behind dataset construction, taxonomies, and AI models is that abstract concepts can be treated as universal and have recognizable visual cues that can be found in images of people.
In 2019, Kate Crawford and Trevor Paglen revealed how the Person category in ImageNet contained hundreds of offensive labels applied to human subjects [9]. ImageNet is one of the first and largest datasets for AI, developed in 2006 by Fei-Fei Li, then released to the public in 2009, and maintained by both Stanford and Princeton University [10][11]. Now, what was the reaction of these institutions to the discovery of these offensive labels? Rather than questioning why such categories existed, the decision was to simply delete the entire “person” category from the dataset. The bias ceased to be visible, while the classificatory logic that produced it continued to govern the underlying infrastructure.
A similar response occurred with Google Photos, when the algorithm notoriously labelled Black people as “gorillas.” Was the solution to confront the racial assumptions embedded in the system, or simply to erase “gorilla” from the classificatory system altogether?
This kind of response has become a recurring shortcut in institutional and corporate contexts. Instead of interrogating the structures that generate harm, uncomfortable outcomes are deleted, silenced, or sanitised. In this process, invisibility is repeatedly mistaken for accountability. The problem does not cease to exist; it simply becomes harder to see.

Now, I’ve a series of provocations.
Why should we remove bias from AI if the same biases continue to structure our societies? Is hiding, flattening, and erasing difference not itself a form of harm?
And what if bias were approached not as a technical error, but as a trace of positionality? What if, instead of erasing it, we exposed bias as a site of negotiation, responsibility, and situated knowledge?
These questions led me to classification itself. Rather than asking how to fix AI outputs, I began asking how AI learns to see. In an experiment conducted in 2023 using Krea’s webcam-to-image tool, I prompted myself into different professional roles.
The results were, as expected, strikingly consistent:
If I were a nurse, I would be a woman. If I were a CEO, I would be a man. If I were a teacher, I would wear glasses. If I were a cashier, I would be a woman. If I were a politician, I would be a man.
What emerged was not randomness, but an infrastructural grammar through which bodies, roles, and value were being associated. This raised a further question.
Is it possible to intervene in the predictive closure of these categories? Can classification be opened to ambiguity without being replaced by another rigid system decided in advance?
At this point, I chose to start, again, from myself.
Not as a passive observer, but as a situated subject within these systems. I began asking how I would annotate my own image. Which labels would I choose for myself? What would it mean to be both the one who sees and the one who is seen?
This inquiry unfolded through a practice-based journey into annotation as a form of self-knowledge and reflection. I built a small dataset composed of thirty photographs of myself, taken over several years by people close to me. Restricting the dataset to my own image was both an ethical choice and a methodological strategy, positioning the self as the site of inquiry. In doing so, I simultaneously inhabited three roles: researcher, annotator, and subject of the annotations.

Using Roboflow, a commonly used open-source platform for dataset annotation, I began the process of semantic segmentation and manual labelling. Unlike conventional annotation, which relies on predefined ontologies and rigid instructions, I adopted a subjective, reflective, embodied, and relational approach. Annotating just thirty images took approximately five hours. This slowness was intentional. It stood in direct contrast to the speed and efficiency expected in industrial AI pipelines.

Working slowly allowed me to reflect on how each label emerged, to notice when and why labels changed, and to become aware of what I was highlighting or omitting. This reflective groundwork later shaped how I interpreted the behaviour of the model trained on this dataset.

By reflexive, subjective, and intimate labels, I mean labels that exceed object description. A body part was no longer just an arm or a leg, but became “feeling well with my body” or “finally loving my legs.” Some annotations extended beyond the visible frame, such as “my boyfriend looking at me” or “happy with my mum,” showing how relationships and memory shape meaning, as also argued by Gillian Rose [12]. Even ordinary objects shifted through interpretation. A glass became “too small” because I wanted more coffee. The floor turned into “Angry Birds,” recalling the frantic, spasmodic movements of small birds stealing sugar from café tables during a Sunday morning breakfast.

I then trained a small model on this dataset to explore how affective, subjective, and interpretive descriptors could be processed computationally.
When tested on unseen images, the model generated unstable and ambiguous predictions. Rather than interpreting this instability as error, I treated it as a relational encounter between my annotations and the machine’s logic.
They reflected the partial, changing logic I had embedded into the dataset through my moods, associations, and relational cues.

In this sense, the model acted as a mirror. It reflected back the way I see, describe, and interpret myself, but only through the structure of the annotations I had provided. The categories I chose, the associations I made, the elements I highlighted or omitted, all became part of its learning process. When the model produced unstable or partial predictions, it was not revealing a defect in the system alone. It was revealing the situated perspective embedded in the dataset I had constructed. What surfaced was not an external flaw, but my own bias, sedimented through annotation and returned as a relational trace of how identity and meaning emerge.

This experiment suggests a different role for design.
Rather than optimising AI systems or correcting outputs, design can operate as an epistemic practice that works from within classification to open it up, interrogate it, and reconfigure it.
Datasets can become dialogical spaces rather than fixed truths. Annotation can shift from enforcing agreement to enabling encounter. Disagreement does not need to be resolved. It needs to be made visible.
If, as Hanna Davis suggests, a dataset is a worldview [13], then the question is not how to make it neutral, but whose perspective it carries, and who is accountable for it. Classification is undeniably a political act. But it can also become an act of care.
Care toward oneself, in acknowledging positionality. Care toward others, in making space for difference to exist without being reduced, erased, or forced into resolution.

What might it mean to design classificatory systems that listen rather than predict? Systems that hold ambiguity rather than eliminate it. Systems that treat knowledge not as a final statement, but as an ongoing, situated negotiation.
These are not questions meant to be resolved quickly. They are questions worth inhabiting.
References