Перейти к основному содержимому

Claude Voice with Opus & Sonnet (July 2026): Reasoning-First Voice vs GPT-Live

Начальный
What you'll learn
  • Understand what Anthropic actually shipped on July 23, 2026 — and what it deliberately did not change
  • Know when to switch between Haiku, Sonnet, and Opus mid-conversation, and what changes when you do
  • Use the connected-app connectors (Gmail, Calendar, Docs, Slack) from voice without hitting the free-tier ceiling
  • See why Claude's turn-based voice is a different design bet from OpenAI's full-duplex GPT-Live — and pick the right tool for the job

On July 23, 2026, Anthropic opened Claude's voice mode to its heavier reasoning models — Opus and Sonnet, alongside Haiku — and wired connected apps (Gmail, Google Calendar, Google Docs, Slack, among others) into the voice experience. Two weeks earlier, on July 8, OpenAI had shipped GPT-Live, a speech-native full-duplex system that listens and talks simultaneously. Same week, two very different bets on what voice AI should be.

This page pulls the two apart, because the difference matters for what you should actually reach for.

What Anthropic shipped (and what it deliberately didn't)

Here's the non-obvious part: Anthropic did not upgrade the voice stack itself. A company spokesperson told TechCrunch and others that Claude keeps its turn-based design — "Claude listens, pauses to think, then answers" — and the underlying voice model was not changed with this release. What changed is the reasoning model behind the voice: you can now put Opus's thinking or Sonnet's balance behind the spoken back-and-forth, where before you got Haiku only.

That framing predicts everything else on this page:

  • Expect deeper answers, not smoother conversation.
  • Expect tool-heavy voice requests to actually work (draft an email, reschedule a meeting), because the model behind the mic is now capable enough to plan them.
  • Do not expect GPT-Live-style barge-in, backchannels, or the model murmuring "mhmm" while you talk — that requires a full-duplex speech-native model, which Claude voice is not (yet).
Pro tip

If you last used Opus in text chat, voice starts with Opus — but it picks the fastest generation of that model family to keep the conversation responsive. Your text-mode habits set your voice defaults. Switching model in text is the easiest way to change what voice hands you next time you tap the mic.

The turn-based bet vs full-duplex GPT-Live

Both announcements landed in the same 15 days. They're solving different problems.

Claude Voice (July 23, 2026)GPT-Live (July 8, 2026)
ArchitectureTurn-based: STT → LLM → TTS pipeline (see full-duplex voice AI)Speech-native full-duplex — one model that consumes and produces audio directly
Interrupts / backchannelsNo; listens, then answersYes; can barge in, say "mhmm", stay quiet
Model tiers in voiceHaiku, Sonnet, Opus (paid); Haiku only (free)GPT-Live-1 (paid), GPT-Live-1 mini (default, incl. free)
Reasoning depthInherits your text model — Opus can now think in voiceFast conversational front-end; delegates heavy reasoning to a frontier model in the background
Connected tools in voiceGmail, Google Calendar, Google Docs, Slack (paid = multiple; free = one)Voice-native tool use during conversation
Latency profileDepth-first: heavier model = deeper answer, longer pause before it speaksLatency-first: designed for ~human-like turn-taking
Developer APINo developer voice API at time of writingGPT-Live not yet in API; gpt-realtime remains the current developer product

Pick Claude voice when you want to think out loud with a strong reasoning model and get real work done in your connected apps by voice. Pick GPT-Live when the conversation itself matters more than depth — interviews, tutoring, language practice, anything where interruption and rapid exchange are the point. Both can be right in the same week.

Watch out

Anthropic's own framing is a give-away: "focused on intelligence and tool access." That's the honest way to describe what changed. If someone tells you "Opus voice mode is more natural" — it isn't. It's smarter and can now touch your inbox and calendar, but the conversational feel is the same turn-based experience as before.

Getting the update working

Guided walkthrough1 of 5
  1. The upgrade is rolling out in beta across Anthropic's mobile apps, Claude Desktop, and the web app. If you don't see the model picker yet, check for an app update and re-open the voice sheet.

Voice request that actually uses the new tool access

"Reschedule my 3 pm with Sara to Thursday morning, draft a short apology email in Gmail, and add the new slot to the shared team calendar — read the draft back to me before sending."

That single utterance now works because (a) the model behind the mic is capable enough to plan the sequence, (b) Google Calendar and Gmail connectors are live in voice, and (c) "read it back before sending" is exactly the checkpoint you want when you can't see the screen.

The model picker: when to switch to what

The picker is the single most useful lever in the update. Rule of thumb:

  • Haiku — capture, quick answers, walking-and-talking, and everything on the free tier. Cheapest and fastest.
  • Sonnet — the default balance. Long enough back-and-forth, feedback on a pitch, planning a small task, drafting messages that use one or two tools.
  • Opus — deep reasoning by voice: strategy, code review by ear, multi-step tool sequences ("look at these three docs, summarize the differences, then draft an email"). Expect visibly longer pauses before it starts speaking — that's the depth you asked for.
Pro tip

A power move: start a call on Haiku to capture context fast, then switch to Opus for the "now think about all of that" moment. You get the responsiveness of Haiku during the messy setup and the depth of Opus only where it matters — without paying Opus's latency on the whole conversation.

Connected apps in voice — what actually works

At launch, the connectors most consistently reported as working in voice are Gmail, Google Calendar, Google Docs, and Slack. Several outlets also list Canva and Notion — Anthropic ships new connectors regularly, so treat any specific list as a moving target and check your Settings → Connectors sheet for what's actually available to you.

Free tier gets one connected app; paid plans get multiple. That single-app cap is the point where the free tier starts to feel obviously narrower — voice is where "connect all my tools" pays for itself.

A hands-free morning triage

"Read me the subject lines and first sentences of unread emails from the last 12 hours, group them by whether they need a reply today, and add three follow-up reminders to my calendar for tomorrow morning."

Practical patterns and gotchas

A few things worth internalizing:

  • Language must be set manually. Voice supports English, French, German, Hindi, Indonesian, Italian, Japanese, Korean, Portuguese (Brazilian), and Spanish. You have to pick your language — the model won't auto-detect.
  • Multi-tool calls add latency. If the response is slow, it's usually because the model is calling two or three connectors in sequence. Ask for fewer things per turn.
  • Tool results may not fully display in voice. Use voice to initiate; switch to the text view of the same chat to inspect what Claude actually did. The chat is one artifact — the mic and the text field are just two ways into it.
  • Voice is downstream of your text-chat memory and preferences. If you use Memory, voice sees the same personal context. If you use Reflect, voice sessions count towards your monthly recap.
  • This is not a full-duplex system. If you catch yourself wishing you could interrupt cleanly or hear a "mhmm" while you think — that's a signal your task actually wants GPT-Live or a research demo like Kyutai Moshi, not Claude voice. Different tool.
Key takeaways
  • The July 23, 2026 update changed the model behind Claude's voice — not the voice stack. Same turn-based architecture, deeper reasoning and real tool access on top.
  • Voice inherits your last text-chat model and runs its fastest version. Change model in text to change your voice default.
  • The model picker inside a live call lets you jump Haiku → Sonnet → Opus mid-conversation — use it to keep Haiku's responsiveness for capture and Opus's depth for the payoff.
  • Free tier: Haiku only + one connected app. Paid tier is where voice starts to feel like a full assistant, because multi-tool access is where the model earns its keep.
  • Claude's bet is intelligence + tools on a conventional pipeline; OpenAI's GPT-Live bets on full-duplex conversation. Both are correct answers to different problems.

Check yourself

Check yourself

0/4
  1. The July 23, 2026 Claude voice update — what actually changed under the hood?
  2. You want fast capture at the start of a call and deep synthesis at the end. Best move?
  3. You're on the free tier and you connected Gmail. You now try to add Google Calendar. What happens?
  4. A user says 'Claude voice on Opus feels more natural than before.' Is that a claim you should trust?

Flashcards

Нажмите Enter или пробел, чтобы перевернуть карточку. Используйте стрелки влево и вправо для перехода между карточками.Показан термин.
1 / 6

Sources & further reading