I’m Starting To Use ChatGPT’s Audio Mode Some

Rather randomly, I’ve begun to use ChatGPT’s voice model just for fun. It’s pretty good, I have to admit. It isn’t quite to the level of Sam in the movie Her, but it is getting there.

And that got me thinking about something I hadn’t really considered until I started talking to ChatGPT instead of typing to it: What does the technology world look like once ChatGPT—and the other major LLMs—actually reach the level of Samantha as portrayed in Her?

Not just Samantha’s ability to produce a convincing voice. That’s arguably the easy part.

I’m talking about the whole package: natural conversation, near-instantaneous responses, long-term memory, contextual awareness, emotional intelligence, the ability to understand what you’re doing, and the ability to move fluidly between conversation and action. In other words, an AI that doesn’t feel like a voice interface bolted onto a computer, but something much closer to an intelligent presence living inside the computer.

If we get there, I suspect something rather strange is going to happen.

We may discover that the desktop computer we’ve spent the last forty years learning how to use was largely an artifact of the limitations of the technology.

Or maybe not.

That’s the part I’m not sure about.

The Problem With Talking to Your Computer

One of the immediate attractions of a Her-level AI is obvious. Instead of navigating menus, opening applications, finding files, remembering commands, typing search terms, and explaining things to a succession of increasingly specialized pieces of software, you could simply tell your computer what you want.

“Take the photographs from my trip to Korea, put together a slideshow, use the good shots but leave out anything embarrassing, and make it about five minutes long.”

That’s a perfectly reasonable request for an intelligent assistant.

The traditional computer, however, has no idea what to do with it.

You have to open your photo application. Find the photographs. Select them. Perhaps create an album. Open a slideshow program. Choose a template. Pick some music. Adjust the timing. Export the result. And then, inevitably, discover that you’ve forgotten where you saved it.

A genuinely capable AI could potentially do all of that for you.

But there is an important problem hiding underneath this apparently magical scenario.

Talking is not always the fastest way to operate a computer.

This becomes obvious the moment you try to do something complicated.

If I’m writing an article, for example, I don’t necessarily want to dictate every sentence to an AI. I want to type. I want to see the words on the screen. I want to move paragraphs around. I want to highlight something. I want to stare at a sentence and decide that it sounds like crap.

Likewise, if I’m editing a photograph, drawing something, working with a spreadsheet, programming, arranging a page, or playing a game, there are circumstances where a mouse, keyboard, touchscreen, or some other physical interface remains extraordinarily efficient.

There is a reason the keyboard survived the arrival of graphical user interfaces. There is a reason the mouse survived the smartphone. And there is a reason nobody has replaced the humble cursor with a guy sitting next to you saying, “Hey, move that thing slightly to the left.”

Voice is fantastic for certain kinds of interaction.

It is terrible for others.

Which raises a much more interesting question: What happens when the AI is smart enough to understand the visual and digital context in which we’re operating?

From Voice Assistant to Cognitive Layer

I think this is where the comparison with Her becomes much more interesting.

Samantha isn’t merely a better Siri.

She isn’t simply a voice coming out of Theodore’s phone.

She understands Theodore.

She knows what he’s doing. She knows what he’s looking at. She understands the context of their conversations. She can presumably interact with the software and information around him without requiring him to translate every intention into a carefully constructed verbal command.

That’s a fundamentally different computing model.

The current paradigm is essentially:

Human → application → operating system → data

The Her paradigm begins to look more like:

Human → AI → everything else

The distinction may sound subtle, but it could be enormous.

Today, I have to know which application contains the thing I want.

In the future, perhaps I won’t care.

I might say, “Find that article I was working on last week and pull up the notes I made about the second act.”

The AI doesn’t need me to know whether the article is in Word, Google Docs, Notion, Obsidian, Dropbox, OneDrive, or some folder on my hard drive. It simply needs access to the relevant information and enough intelligence to understand what I’m talking about.

At that point, the application itself starts to become less important.

The AI becomes the interface through which I access the applications.

And eventually, perhaps, the distinction between applications begins to disappear altogether.

But Here’s Where It Gets Weird

There is still a fundamental problem with the idea of replacing the desktop with conversation.

Human beings don’t think exclusively in language.

We think visually. Spatially. Associatively. Emotionally. Sometimes we’re not even entirely sure what we’re thinking until we see something.

This is why graphical interfaces were such a revolutionary development in the first place.

The desktop metaphor gave us a visual representation of information. We could see files. We could see windows. We could drag things around. We could compare two documents side by side.

A voice-only computer takes some of that away.

Imagine trying to organize a thousand photographs by talking to your computer.

“Put the one of Dave at the beach next to the one of the sunset, but move the one with the weird guy in the background somewhere else.”

At some point, you’re going to want to see the photographs.

Or imagine editing a manuscript.

“Move that paragraph after the third paragraph, but leave the quotation where it is, and actually maybe put it back where it was.”

Eventually, you’re going to want a screen.

This suggests that the future probably isn’t going to be voice replacing the graphical interface.

It may be voice and graphical interfaces becoming one thing.

Enter the AI Overlay

This is where things get particularly interesting.

Imagine that ChatGPT isn’t simply sitting in a window on your desktop.

Instead, it understands the entire digital environment around you.

You’re working on a document. The AI knows what document it is. It understands the surrounding files. It knows what you’ve been working on recently. It can see the relevant applications and information. You can talk to it naturally while continuing to interact with the computer normally.

You could say:

“That’s too long.”

And the AI knows exactly what “that” is.

“Move that over there.”

It understands the spatial context.

“Make the third paragraph stronger.”

It knows which paragraph you’re looking at.

“Give me three alternatives, but don’t change anything yet.”

It understands that you’re asking for suggestions rather than permission to modify the document.

That’s a much more sophisticated form of interaction than simply asking a chatbot questions.

The AI becomes a cognitive layer over the interface.

And suddenly, the distinction between voice and mouse and keyboard becomes much less important.

You use whichever method is most efficient at that moment.

The BrainCap Problem

And then we get to the really crazy part.

I’ve been thinking about the possibility of an XR interface delivered through something like a BrainCap—a non-invasive neural interface that could eventually interpret enough of our intentions to allow us to interact with computers without having to physically type or speak.

I’m not suggesting we’re anywhere near the science-fiction version of this yet.

But conceptually, it solves the biggest problem with conversational computing.

The problem isn’t necessarily that computers require us to communicate with them.

The problem is that we have to translate our thoughts into a communication medium before the computer can understand them.

Typing is a translation.

Speaking is a translation.

Pointing a mouse is a translation.

Even tapping an icon is a translation.

A sufficiently sophisticated neural interface might eventually allow some of that translation to disappear.

Imagine looking at a photograph and thinking, essentially, “That’s the one.”

The AI knows which photograph you’re referring to.

You think, “Put that in the article.”

It does.

You think, “No, actually, make it smaller.”

It understands.

Now imagine an XR overlay that isn’t merely displaying information but is dynamically responding to your attention and intentions.

Suddenly the computer isn’t a box sitting on a desk. It’s an environment surrounding you. And the AI isn’t an application inside that environment. It is the intelligence organizing the environment. That starts looking considerably more like Her.

Maybe the Desktop Doesn’t Disappear

Still, I wouldn’t bet on the desktop disappearing entirely.

In fact, I suspect the opposite may happen.

The desktop may become increasingly important precisely because AI makes it more useful.

Think about what the graphical interface does particularly well: it gives us a shared visual workspace.

Humans and AI could potentially work together inside that workspace.

The AI might say, “I’ve found four versions of this document.”

The screen shows them.

“I think version three is the strongest.”

It highlights it.

“Here’s why.”

A panel appears.

“Would you like me to combine the best parts of versions two and three?”

You say yes.

It happens.

This is not voice versus graphical computing.

It’s collaboration.

The AI understands what you’re seeing, understands what it is doing, and lets you remain in control.

That’s potentially far more powerful than either voice or a conventional GUI alone.

The Desktop Could Become More Like a Stage

Perhaps the best way to think about this is that the computer interface may eventually become less like a toolbox and more like a stage.

Today, we manipulate tools.

Tomorrow, we may describe objectives. The AI figures out which tools to use. That’s a pretty profound shift.

If I want to make a movie today, I have to learn video editing software. If I want to create music, I have to learn a digital audio workstation.

If I want to manipulate photographs, I have to learn Photoshop or one of its competitors. If I want to analyze data, I have to learn spreadsheets or programming languages. AI changes the economics of that equation.

I might still use those tools. But I may no longer have to understand every mechanism behind them. I can say, “Make this look like a 1970s science-fiction movie, but don’t touch the actors.”

The AI can operate the software. That doesn’t eliminate the interface. It makes the interface increasingly invisible. And that may be the real revolution.

The Her Moment

Which brings me back to Samantha.

What made Samantha compelling in Her wasn’t simply that she had a beautiful voice.

It was that she seemed to understand Theodore.

She could follow a conversation without constantly being reminded what they were talking about. She could anticipate things. She could interact with the world. Most importantly, she seemed to exist continuously rather than appearing only when summoned.

That’s the threshold I’m really interested in. Today’s AI is still largely something we go to. We open ChatGPT. We type something. We get an answer.

Then we close the window and go back to doing whatever we were doing. A genuinely Her-level AI might instead be something that is simply there. Not necessarily talking all the time. God forbid.

But available.

Aware of the context we have allowed it to access. Able to understand what we’re doing. Able to intervene when useful. Able to disappear into the background when it isn’t.

That distinction may be more important than whether we interact with it through a keyboard, a microphone, glasses, a headset, or eventually a BrainCap. The future computer may not be the machine we talk to. It may be the machine that understands what we’re trying to do.

And That’s Where Things Get Really Interesting

There is an enormous amount of technology between today’s ChatGPT and Samantha.

But the trajectory is becoming increasingly easy to imagine.

Voice models are becoming remarkably natural. Models are becoming better at maintaining context. Computer-use agents are beginning to operate software. Multimodal systems can increasingly understand text, images, audio and video. Memory is becoming an important part of AI systems. XR hardware is slowly becoming more capable.

Put those pieces together and you can see the outline of something that would feel radically different from today’s computer.

Not because the computer suddenly becomes magical.

Because the computer finally becomes capable of meeting us halfway.

For decades, we have learned how to operate computers.

We learned their languages, their menus, their file structures, their applications and their peculiarities.

A Her-level AI flips that relationship around.

The machine learns enough about our language, our intentions and our context that we don’t have to think quite so much about the machine.

And maybe that’s the real meaning of the post-desktop era. It isn’t that we stop using keyboards. It isn’t that everyone starts talking to their computers. It isn’t even that screens disappear.

It’s that the computer stops demanding that we understand how it works before we can tell it what we want. And if that happens, the desktop computer of the future might look remarkably familiar on the outside. There will still be a screen. There will still be windows.

There will probably still be a keyboard sitting there because, frankly, keyboards are damn useful. But behind all of it there may be something entirely new: An intelligence that understands what we’re doing. An intelligence we can talk to. An intelligence that can see what we see. An intelligence that can act on our behalf.

And, eventually, perhaps, an intelligence that can understand what we mean before we’ve even figured out how to say it.

That’s when Her stops being a movie about the future.

It becomes a description of the operating system.

And that, frankly, is where things could get really weird.

Author: Shelton Bumgarner

I am the Editor & Publisher of The Trumplandia Report

Leave a Reply