OpenAI has expanded its ChatGPT Voice feature to the macOS and Windows desktop applications, powered by the GPT-Live model that operates in full duplex. This setup lets the system speak, hear, and process input simultaneously, aiming for conversations that feel less interrupted than earlier voice modes and follow instructions with greater consistency.
The update arrives after OpenAI first detailed GPT-Live earlier this month as a step beyond previous generations that often struggled with overlapping speech or rigid turn-taking. In practice, the desktop version now lets subscribers use voice not only for open-ended talk but also for directing Chat, Work, and Codex sessions. Users can launch new tasks, check the status of ongoing work, or adjust agents running in parallel threads. A developer might ask it to review a pull request, outline a codebase, draft a document, or coordinate several Codex jobs at once.
Because Work and Codex can draw on connected tools and files when permissions allow, the voice interface reaches into services such as Slack, GitHub, and Notion. On macOS an additional option called Screen context lets someone say “take a look at this,” after which the system receives a snapshot of the frontmost window for extra situational awareness. Behind the scenes the audio model hands off structured work to separate Codex threads that can call tools including appshots and computer-use agents.
This represents a clear shift from earlier AI voice experiences that stayed largely conversational. Systems like the first wave of smartphone assistants or even ChatGPT’s prior Advanced Voice Mode handled queries well enough for casual use but rarely helped complete multi-step professional work. The new desktop integration tries to close that gap by treating voice as another input method for productivity rather than a novelty. Whether the full-duplex design consistently reduces interruptions in noisy environments or complex coding sessions remains something users will test over time. Access is limited to Plus, Pro, Business, Edu, and Enterprise plans, and the same experience can also be reached remotely through the iOS app.
The rollout is global through the desktop clients. Early feedback already highlights the planning potential: one user described starting with a vague idea, refining it through spoken questions, mapping architecture in real time, then spinning up a Codex task to implement it. That workflow illustrates the intended direction—voice as a continuous steering mechanism rather than a one-shot command channel. Still, the dependence on subscription tiers, the privacy implications of screen capture, and the practical reliability of tool chaining will determine how widely the feature moves beyond early adopters. For now it simply extends an existing capability into the environments where many people already spend their working hours.
