OpenAI's GPT-Live Voice Mode Adds File Attachments, Projects, and Cross-Feature Integration
Key Takeaways
- •GPT-Live employs full-duplex voice architecture that enables simultaneous listening and speaking, supporting natural conversational behaviors such as mid-sentence interruptions and overlapping dialogue.
- •The flagship GPT-Live-1 model is available to paid subscribers with full capabilities, while free-tier users receive GPT-Live-1 mini with reduced processing headroom.
- •GPT-Live can route computationally demanding tasks to more powerful models such as GPT-5.5 while maintaining a responsive real-time voice layer.
- •Voice sessions now support file attachments, image processing, memory features, web search, and integration with the Projects feature across Work, Codex, and desktop environments.
- •The ChatGPT Library currently does not support direct file search or file addition during voice sessions, representing a remaining limitation in the updated system.

OpenAI launched GPT-Live on July 8, 2026, establishing it as the new default voice model family for ChatGPT. The update brings file attachments, project integration, memory, and search capabilities into a part of the product that previously functioned largely in isolation.
Users can now hold a voice conversation with ChatGPT while simultaneously sharing documents, eliminating the need to switch to text mode for content-based tasks. The change addresses a long-standing friction point in AI assistants, where voice interactions have typically been treated as a separate surface area rather than an integrated layer that can access the same tools and context available in text-based workflows.
Full-Duplex Voice Architecture
GPT-Live is a family of full-duplex voice models, meaning the system can listen and speak at the same time rather than alternating turns. This allows for natural conversational behaviors such as mid-sentence interruptions and overlapping dialogue, which traditional turn-based voice assistants could not handle without cutting off or restarting responses. The flagship model is GPT-Live-1, available to paid subscribers with the full feature set. A lighter variant, GPT-Live-1 mini, is offered to free-tier users and shares the same conversational architecture but with reduced capability headroom.
GPT-Live can delegate computationally intensive tasks to more powerful models, including GPT-5.5, while maintaining a responsive voice layer. This routing architecture mirrors a broader industry pattern in which lightweight, low-latency interfaces orchestrate work performed by heavier backend models, a design choice that balances real-time responsiveness with deep reasoning capability.
File Attachments and Projects Integration
The most operationally significant addition is file attachment support during live voice sessions. Users can attach files mid-conversation and have ChatGPT process that content in real time. The update also introduces image processing, memory features, and web search within voice sessions.
GPT-Live now connects to ChatGPT's Projects feature, which allows users to maintain persistent context across sessions through stored files, preferences, and custom instructions. This integration extends across Work, Codex, and desktop environments. For users who rely on ChatGPT across multiple surfaces—conversational, coding, and document-centric—the Projects integration effectively turns voice from a standalone interaction mode into a shared entry point that can pull from the same knowledge base as other interfaces.
One limitation remains: the ChatGPT Library, which stores documents separately from Projects, does not currently support direct file search or file addition during a voice session.
Evolution From Advanced Voice Mode
GPT-Live represents an incremental maturation of what OpenAI previously called Advanced Voice Mode rather than a fundamental reinvention. Advanced Voice Mode introduced more natural prosody and emotional range to ChatGPT's speech but still operated on a turn-based system. GPT-Live adds structural connectivity, linking voice to the broader product ecosystem including search, memory, file uploads, and Projects integration. The trajectory reflects a wider push across conversational AI providers to unify voice, text, and multimodal inputs under a single session model, rather than maintaining parallel capabilities with disjointed feature sets.