Claude Voice with Opus & Sonnet (July 2026): Reasoning-First Voice vs GPT-Live
- Understand what Anthropic actually shipped on July 23, 2026 — and what it deliberately did not change
- Know when to switch between Haiku, Sonnet, and Opus mid-conversation, and what changes when you do
- Use the connected-app connectors (Gmail, Calendar, Docs, Slack) from voice without hitting the free-tier ceiling
- See why Claude's turn-based voice is a different design bet from OpenAI's full-duplex GPT-Live — and pick the right tool for the job
On July 23, 2026, Anthropic opened Claude's voice mode to its heavier reasoning models — Opus and Sonnet, alongside Haiku — and wired connected apps (Gmail, Google Calendar, Google Docs, Slack, among others) into the voice experience. Two weeks earlier, on July 8, OpenAI had shipped GPT-Live, a speech-native full-duplex system that listens and talks simultaneously. Same week, two very different bets on what voice AI should be.
This page pulls the two apart, because the difference matters for what you should actually reach for.
What Anthropic shipped (and what it deliberately didn't)
Here's the non-obvious part: Anthropic did not upgrade the voice stack itself. A company spokesperson told TechCrunch and others that Claude keeps its turn-based design — "Claude listens, pauses to think, then answers" — and the underlying voice model was not changed with this release. What changed is the reasoning model behind the voice: you can now put Opus's thinking or Sonnet's balance behind the spoken back-and-forth, where before you got Haiku only.
That framing predicts everything else on this page:
- Expect deeper answers, not smoother conversation.
- Expect tool-heavy voice requests to actually work (draft an email, reschedule a meeting), because the model behind the mic is now capable enough to plan them.
- Do not expect GPT-Live-style barge-in, backchannels, or the model murmuring "mhmm" while you talk — that requires a full-duplex speech-native model, which Claude voice is not (yet).
If you last used Opus in text chat, voice starts with Opus — but it picks the fastest generation of that model family to keep the conversation responsive. Your text-mode habits set your voice defaults. Switching model in text is the easiest way to change what voice hands you next time you tap the mic.
The turn-based bet vs full-duplex GPT-Live
Both announcements landed in the same 15 days. They're solving different problems.
| Claude Voice (July 23, 2026) | GPT-Live (July 8, 2026) | |
|---|---|---|
| Architecture | Turn-based: STT → LLM → TTS pipeline (see full-duplex voice AI) | Speech-native full-duplex — one model that consumes and produces audio directly |
| Interrupts / backchannels | No; listens, then answers | Yes; can barge in, say "mhmm", stay quiet |
| Model tiers in voice | Haiku, Sonnet, Opus (paid); Haiku only (free) | GPT-Live-1 (paid), GPT-Live-1 mini (default, incl. free) |
| Reasoning depth | Inherits your text model — Opus can now think in voice | Fast conversational front-end; delegates heavy reasoning to a frontier model in the background |
| Connected tools in voice | Gmail, Google Calendar, Google Docs, Slack (paid = multiple; free = one) | Voice-native tool use during conversation |
| Latency profile | Depth-first: heavier model = deeper answer, longer pause before it speaks | Latency-first: designed for ~human-like turn-taking |
| Developer API | No developer voice API at time of writing | GPT-Live not yet in API; gpt-realtime remains the current developer product |
Pick Claude voice when you want to think out loud with a strong reasoning model and get real work done in your connected apps by voice. Pick GPT-Live when the conversation itself matters more than depth — interviews, tutoring, language practice, anything where interruption and rapid exchange are the point. Both can be right in the same week.
Anthropic's own framing is a give-away: "focused on intelligence and tool access." That's the honest way to describe what changed. If someone tells you "Opus voice mode is more natural" — it isn't. It's smarter and can now touch your inbox and calendar, but the conversational feel is the same turn-based experience as before.
Getting the update working
- The upgrade is rolling out in beta across Anthropic's mobile apps, Claude Desktop, and the web app. If you don't see the model picker yet, check for an app update and re-open the voice sheet.
- Voice defaults to the last model you picked in text chat, and runs the fastest version of that family. If you were on Opus in text yesterday, voice starts on Opus today. To change the default, switch model in a text chat first.
- Inside a live voice call, the model picker lets you jump between Haiku, Sonnet, and Opus without ending the call. Use it when the task changes — capture on Haiku, deepen on Sonnet or Opus.
- Add connectors (Gmail, Google Calendar, Google Docs, Slack) under Settings → Connectors. Paid plans (Pro, Max, Team, Enterprise) allow multiple; free is limited to Haiku and one connected tool.
- Anthropic warns that using several tools at once can add a short delay, and tool results may not fully display in voice. A common pattern: initiate by voice, then switch to the text view in the same chat to review the drafted email or the returned document.
Voice request that actually uses the new tool access
"Reschedule my 3 pm with Sara to Thursday morning, draft a short apology email in Gmail, and add the new slot to the shared team calendar — read the draft back to me before sending."
That single utterance now works because (a) the model behind the mic is capable enough to plan the sequence, (b) Google Calendar and Gmail connectors are live in voice, and (c) "read it back before sending" is exactly the checkpoint you want when you can't see the screen.
The model picker: when to switch to what
The picker is the single most useful lever in the update. Rule of thumb:
- Haiku — capture, quick answers, walking-and-talking, and everything on the free tier. Cheapest and fastest.
- Sonnet — the default balance. Long enough back-and-forth, feedback on a pitch, planning a small task, drafting messages that use one or two tools.
- Opus — deep reasoning by voice: strategy, code review by ear, multi-step tool sequences ("look at these three docs, summarize the differences, then draft an email"). Expect visibly longer pauses before it starts speaking — that's the depth you asked for.
A power move: start a call on Haiku to capture context fast, then switch to Opus for the "now think about all of that" moment. You get the responsiveness of Haiku during the messy setup and the depth of Opus only where it matters — without paying Opus's latency on the whole conversation.
Connected apps in voice — what actually works
At launch, the connectors most consistently reported as working in voice are Gmail, Google Calendar, Google Docs, and Slack. Several outlets also list Canva and Notion — Anthropic ships new connectors regularly, so treat any specific list as a moving target and check your Settings → Connectors sheet for what's actually available to you.
Free tier gets one connected app; paid plans get multiple. That single-app cap is the point where the free tier starts to feel obviously narrower — voice is where "connect all my tools" pays for itself.
A hands-free morning triage
"Read me the subject lines and first sentences of unread emails from the last 12 hours, group them by whether they need a reply today, and add three follow-up reminders to my calendar for tomorrow morning."
Practical patterns and gotchas
A few things worth internalizing:
- Language must be set manually. Voice supports English, French, German, Hindi, Indonesian, Italian, Japanese, Korean, Portuguese (Brazilian), and Spanish. You have to pick your language — the model won't auto-detect.
- Multi-tool calls add latency. If the response is slow, it's usually because the model is calling two or three connectors in sequence. Ask for fewer things per turn.
- Tool results may not fully display in voice. Use voice to initiate; switch to the text view of the same chat to inspect what Claude actually did. The chat is one artifact — the mic and the text field are just two ways into it.
- Voice is downstream of your text-chat memory and preferences. If you use Memory, voice sees the same personal context. If you use Reflect, voice sessions count towards your monthly recap.
- This is not a full-duplex system. If you catch yourself wishing you could interrupt cleanly or hear a "mhmm" while you think — that's a signal your task actually wants GPT-Live or a research demo like Kyutai Moshi, not Claude voice. Different tool.
- The July 23, 2026 update changed the model behind Claude's voice — not the voice stack. Same turn-based architecture, deeper reasoning and real tool access on top.
- Voice inherits your last text-chat model and runs its fastest version. Change model in text to change your voice default.
- The model picker inside a live call lets you jump Haiku → Sonnet → Opus mid-conversation — use it to keep Haiku's responsiveness for capture and Opus's depth for the payoff.
- Free tier: Haiku only + one connected app. Paid tier is where voice starts to feel like a full assistant, because multi-tool access is where the model earns its keep.
- Claude's bet is intelligence + tools on a conventional pipeline; OpenAI's GPT-Live bets on full-duplex conversation. Both are correct answers to different problems.
Check yourself
Check yourself
0/4Flashcards
Sources & further reading
- Anthropic updates Claude voice mode with more capable models — TechCrunch (July 23, 2026)
- Anthropic brings Opus and Sonnet to Claude Voice Mode — Unite.AI
- Anthropic adds model choice to Claude Voice Mode for all users — SQ Magazine
- Claude Voice Mode adds Sonnet, Opus, and app connectors — TechMyMoney (July 23, 2026)
- OpenAI releases new voice models for more natural live conversations — TechCrunch (July 8, 2026)
- Related pages on AILmanac: Full-duplex voice AI (GPT-Live, Moshi, and the 200 ms problem) · Talking to Claude (Voice Mode) · Connectors · Choosing a Model