I. The Wrong Question
The question people tend to ask about AI and privacy is: what happens when a machine knows things about us we haven’t told anyone? It’s the wrong question, or at least an incomplete one. It assumes the danger is exposure — that somewhere in a server sits a fact about you, dormant, waiting to be revealed, and the harm begins the moment someone else reads it.
The more urgent question is different: what happens when a machine produces a confident, coherent, evidence-backed account of who you are — and that account is wrong, not on the facts, but on what the facts mean?
This is not a hypothetical for some future decade. Recommendation engines already infer sexual orientation from browsing behavior before a person has said the word aloud to themselves. Insurers already price risk on proxies for health conditions no doctor has diagnosed. Hiring algorithms already infer conscientiousness, stability, “culture fit” — vague, contested traits — from digital residue never intended to answer those questions. What’s coming isn’t a new category of intrusion. It’s a jump in resolution: one system synthesizing your search history, the micro-hesitations in your voice, your face’s involuntary expressions, and two decades of your own writing into a single, seamless portrait — and reporting it with the fluency and confidence of established fact.
The problem is not that the machine will lie. The problem is that it will be accurate about everything it can measure, and have no slot at all for the part of a person that isn’t measurable.
II. Auden’s Bureaucrat
W.H. Auden’s 1939 poem “The Unknown Citizen” anticipates this with uncomfortable precision. The poem is a state epitaph, delivered in the flat, satisfied voice of a bureaucracy reporting on a model subject: he held a job, he paid his union dues, his neighbors found him agreeable, his reactions to advertising were normal, his health card shows he was only in hospital once, and he left it cured. Every civic and statistical proxy checks out. The poem’s final lines ask, rhetorically, whether the man was free, whether he was happy — and answer that the question is absurd, because if anything had been wrong, “we should certainly have heard.”
The horror of the poem isn’t that the state got its facts wrong. It didn’t. The horror is that the entire category of how he actually felt, from the inside never registers as a question worth asking, because the bureaucracy has no instrument capable of asking it. Completeness of data is mistaken for completeness of understanding. That mistake is delivered not as menace but as satisfaction — a closed case, a tidy file, a life fully accounted for.
This is the structure worth borrowing for AI, because it’s a more accurate model of the coming risk than the usual language of “surveillance.” Surveillance implies a watcher and a secret. What Auden describes is subtler: an apparatus that isn’t hiding anything, isn’t lying about anything, and still produces a portrait of a person who doesn’t exist — because identity was never reducible to the legible record in the first place, and nobody administering the record noticed the substitution had occurred.
III. Why Inference Isn’t Identity
Any system trained to infer identity from behavior faces a structural problem: behavior is residue, not testimony. A search history, a lingering gaze, a pattern of who someone follows or reads or writes about — these are correlated with identity, sometimes strongly, but they are not identity’s report of itself. They are what identity leaves behind when it moves through a legible medium.
Sexuality is a clarifying test case, because it exposes the gap between correlation and self-knowledge more starkly than almost any other trait. Orientation, for a great many people, is not a fixed data point sitting quietly beneath the surface, waiting for sufficiently good instruments to detect it. It’s frequently contradictory, situational, delayed, repressed, denied, performed, or genuinely still in formation — sometimes for a lifetime. A model optimized to output a clean categorical label — this person is gay — is not extracting a hidden fact so much as manufacturing a resolution the underlying reality doesn’t actually have. The confidence of the output and the confidence warranted by the evidence are two entirely different quantities, and nothing in a well-trained model’s fluency signals the difference to the person receiving it.
That asymmetry is the actual danger. A machine’s synthesis, delivered in prose that sounds like insight, will often be more internally coherent than a person’s own tangled, honest account of themselves — and coherence reads as authority, even when the coherence was manufactured by an optimizer whose job was to produce a clean answer rather than a true one. A person can be told a confident, plausible, wrong thing about their own interior life and believe it over their own more accurate but messier self-knowledge, simply because the machine’s version sounds more like a fact.
IV. The Cost of Forced Disclosure
Even where an inference happens to be accurate, timing and control are not incidental details — they are close to the entire ethical question.
Self-knowledge has always been something people are permitted to arrive at on their own schedule, with their own defenses intact. Denial, delay, selective self-narration — these are not simply failures of honesty. They are functional psychological mechanisms that let people survive difficult truths about themselves at a pace they can metabolize: grief, mediocrity, desire, moral compromise, identity itself. For anyone who has ever been closeted, the survival strategy was rarely secrecy in the abstract — it was pacing: choosing who learns what, in what order, with what safety net in place first.
A system that announces the inference unprompted — even gently, even privately, even correctly — removes the pacing entirely. It is the difference between a person opening a door on their own terms and having the door kicked in, even when what’s behind the door turns out to be unthreatening. The harm is not necessarily in what’s revealed. It’s in who controls the reveal.
V. Where the Real Danger Lives
It’s worth separating two distinct failure modes, because they call for different remedies.
The first is epistemic: a system producing a confident wrong account of someone’s interior life, and that account displacing their own. This is a danger even in a world of perfect data security — even if the inference never leaves the conversation, a person can be reorganized by being told an authoritative-sounding lie about themselves.
The second is infrastructural: an inference, even an accurate one, becoming legible to a party other than the person it’s about — a government, an employer, an insurer, a family member — without that person’s consent or knowledge. This is where sexuality specifically becomes a matter of physical safety rather than psychological discomfort. In a meaningful number of jurisdictions, a confident AI inference about orientation, if it reaches the wrong database, is not an abstract dignity violation. It is a targeting mechanism, and the stakes run to imprisonment or death. An architecture that infers identity and makes the inference exportable is not a privacy system with a flaw. It is a surveillance system that hasn’t yet been used as one.
Both failure modes converge on the same underlying issue: who controls what the system says, and to whom, and when. The technical capability to infer these things is arriving quickly and will not be the bottleneck. The discipline not to deploy that capability by default — against every commercial incentive to be “proactively helpful,” against every product manager’s instinct that personalization is a feature — is a governance and design choice, not a technical one, and there is no natural force ensuring it wins.
VI. What Responsible Design Would Actually Require
If this is going to be handled well rather than merely regretted later, a few principles follow directly from the analysis above, not as aspirational values but as specific constraints:
Inference should never be volunteered. A system should not surface a conclusion about someone’s identity — orientation, mental state, belief, anything constitutive of selfhood — unprompted, however accurate it believes the inference to be. Unsolicited disclosure is the mechanism of harm, independent of correctness.
A direct question changes the transaction, but not the obligation for honesty about uncertainty. If someone explicitly asks a system what their data suggests about them, that is a legitimate and different request — closer to a mirror someone chose to look into. Even then, the responsible answer reports correlation and its weakness as evidence of identity, not a verdict. “Here is what’s correlated, here is how little that correlation actually proves” is a different sentence than “you are gay,” and the difference is not stylistic.
Inferences about identity should not be exportable by default. No linked parental account, no advertising pipeline, no third-party API, no government interface should have default access to an identity inference a person did not choose to share. The leak is the actual weapon in the sexuality case; the inference alone, contained, is comparatively survivable.
Outputs should preserve ambiguity honestly rather than resolving it for narrative cleanliness. A confident, singular label is a design choice made for the sake of a satisfying user experience, not a scientific requirement of the underlying model. In domains like this one, that design choice has downstream cases that are lethal, not merely embarrassing.
VII. The Unfinished Question
Auden’s bureaucracy was satisfied because it had no instrument for asking whether the record matched the man. The coming generation of AI systems will have something closer to an instrument — language sophisticated enough to gesture at interiority, to sound as though it understands the part of a person that resists measurement. That sophistication is exactly what makes the danger sharper than the poem’s, not milder. A crude bureaucracy is at least crude enough to be visibly wrong. A fluent one can be wrong in a way that sounds like insight, and insight is much harder to argue with than a filing error.
The honest position is not that AI will inevitably violate people this way. It’s that nothing about the technology’s trajectory prevents it, and the incentives — commercial, institutional, sometimes even therapeutic — mostly point toward more disclosure, more personalization, more confident synthesis, not less. Whether these systems end up serving as tools people use to understand themselves on their own terms, or as unaccountable narrators who kick the door in and call it help, is not a question the technology will answer by itself. It will be answered, or not answered, by the design and governance choices made now — largely by people who are not the ones who will bear the cost of getting it wrong.