The conversation
A conversation you can interrupt.
One session engine behind an embeddable widget and a real phone number. Voice is full duplex, so a visitor can cut in mid-sentence the way they would with a person.
What the capability actually does.
Full duplex, with barge-in
The agent stops when it is interrupted rather than talking over the caller. Interruption handling is implemented once in the session engine, so it behaves the same on the phone as in the browser.
Camera and screen share
A visitor can turn on their camera or share their screen partway through a voice conversation, and the agent describes what is actually in front of them. Useful for walking someone through a form they are stuck on.
Hand it a file mid-conversation
A photo, a receipt, a spreadsheet, a slide deck, a PDF. Images go down the live video channel; everything else is extracted to text or described by a vision call. It lands in the transcript as its own entry, so the conversation stays auditable.
Text mode asks for nothing
Typed chat with streamed replies and no microphone permission prompt. A visitor chooses at the start and can switch either way.
It cannot leak your styling
The widget renders inside a shadow root, so your page CSS cannot reach into it and its styles cannot escape. One script tag, and it can be restricted to domains you name.
The screens this happens on.

It can see what they are looking at
The visitor started this session by clicking a widget on a page, chose voice, and then shared their screen partway through. The agent is describing what is actually rendered in front of them rather than working from a description they had to type out. Audio runs full duplex throughout, so cutting in mid-sentence works the way it does with a person.
- Camera and screen share are voice-mode only
- The visitor starts the share, never the agent

The same agent, without a microphone
The same agent, reached the same way, with no microphone permission prompt. A visitor picks voice or text when the session opens and can switch afterwards. Replies stream in as they are generated rather than appearing all at once, which is what keeps a typed conversation from feeling like a form submission.
- Mounts inside a shadow root, so your CSS and its own cannot reach each other
What it deliberately does not do.
- The live voice engine runs on Gemini and only Gemini. No other provider currently offers an equivalent realtime bidirectional audio API, so this is a technical boundary rather than a preference.
- A session commits to voice or text when it starts, because the underlying model does.
- File attach is off by default and enabled per agent. Legacy binary .doc files are not among the accepted formats.