one year on
ChatGPT adds voice and image capabilities in multimodal expansion
OpenAI begins rolling out new voice conversation and image understanding features.
OpenAI has begun rolling out new voice and image capabilities in ChatGPT, allowing users to speak to the chatbot and share images for analysis.
The voice feature lets users speak to ChatGPT, and the image feature adds image capabilities in the app.
One commenter says Microsoft is planning a rollout of its own assistant as soon as tomorrow, though another replies that Windows 11 Copilot is not really the same thing.
The Hacker News community is divided on the news. While many welcome the update, a vocal contingent notes that the voice latency — several seconds per turn — is not yet compelling, pointing to faster local prototypes using Llama 2 and Whisper that achieve responses in hundreds of milliseconds. Others push back, arguing that latency is less important than accuracy and reliability at scale.
The record
Some commenters express disappointment about multi-second latency in the voice demo, while others share their own low-latency voice assistant prototypes using local models like Llama 2. One developer demonstrates a voice ordering system with negative latency, claiming it outperforms human drive-through attendants. A separate thread highlights the importance of interruption handling for natural conversation.
One year later — open only if you can handle spoilers
OpenAI's voice and image features have since become core to the ChatGPT experience, with over 100 million weekly active users. However, the initial voice latency complaints foreshadowed a year-long race to reduce delay, culminating in real-time streaming voice modes by late 2024.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy