Google launches Guided Vision in Gemini Live for accessibility

Google launches Guided Vision in Gemini Live for accessibility
News

Google launched Guided Vision in Gemini Live on October 1, 2026, bringing conversational visual assistance to compatible Android devices. Users can share their camera during a Gemini Live session and hear descriptions of what is around them, ask follow-up questions, and receive spoken prompts to adjust the framing. Google says the feature is available on Android 9 and later in regions and languages where Gemini Live is supported.

The interaction is designed for situations in which a still image description is not enough. Guided Vision can read fine print on a food label or menu, help locate an object such as an earbud or a spice jar, describe colours and patterns, and give an overview of a room or table. If the camera is pointed too high, too close, or to the side, Gemini can ask the user to pan, tilt, or step back so it gets more useful context. Those spoken reframing cues are important for people who cannot rely on a viewfinder, but they can also help anyone working in poor light or with small text.

Google says it developed Guided Vision with blind and low-vision communities and partnered with Aira. The company reports that Aira contributed tens of thousands of hours of visual-interpretation data and that more than 1,000 members of its Trusted Tester network helped refine the system across daily routines. These are Google’s and Aira’s descriptions of the development process, not an independent measurement of accuracy. Android Authority independently described the rollout as a tool for blind and low-vision users, while an accessible-technology help page documents the available shortcuts and TalkBack integration.

Google also sets clear limits. Guided Vision can make mistakes and is not a medical device, mobility aid, white-cane replacement, navigation system, or obstacle detector. Users should continue to rely on established mobility aids and safe-travel practices. That warning matters because a fluent spoken description can sound more certain than the camera and model actually are.

For users, the launch makes Gemini Live more useful as an interactive accessibility aid rather than a one-off image captioner. For makers, it shows the value of combining multimodal perception with audio guidance and user-controlled framing. The important test now is practical reliability: whether the feature recognises objects and text consistently across devices, languages, lighting conditions, and everyday environments without encouraging unsafe reliance.