Gemini Speak to Window on macOS: the two voice modes and where the line sits

Gemini for macOS puts a system-wide voice feature on the Fn key. It runs in two distinct modes: plain dictation, which is on by default and only processes your voice, and reasoning with screen context, which is off until you opt in and which sends what Gemini can see on screen to Google. Everything worth understanding about this feature is the boundary between those two modes.

What the Fn key does

Google calls the feature Speak to Window. It binds to the Fn key (the globe key, ๐ŸŒ) and has three gestures:

  • Press and hold Fn, speak, release to submit. Gemini writes the result at your cursor in whatever app is frontmost.
  • Double-tap Fn to start hands-free recording, tap again to stop. This is the one to use for anything longer than a sentence or two.
  • Click the Speak to Window icon in the Gemini prompt bar, if you would rather not use the keyboard at all.

The shortcut is configurable under Settings > Speak to Window, which matters more than it sounds โ€” see the conflict note further down.

Mode 1: intelligent dictation (default, no opt-in)

Out of the box, Speak to Window is a transcription tool. It takes your audio and produces cleaned-up text: filler words removed, mid-sentence self-corrections resolved rather than transcribed literally, and structure applied โ€” bullet points and paragraph breaks where your speech implies them.

In this mode Gemini has no access to your screen. It receives audio and returns text. That is the entire transaction. The feature works in any text field in any app because it types the result rather than integrating with the app.

Mode 2: reasoning with screen context (opt-in)

Turning on reasoning changes what the same keypress does. Open Gemini Settings from the menu bar, go to the Speak to Window tab, and switch on Use reasoning.

Once enabled, Gemini attempts to classify each utterance: is this text to be typed, or an instruction to be executed? Say “the quarterly numbers came in under plan” and it dictates. Say “rewrite this in the past tense” with a paragraph selected and it rewrites. The classification is inferred, not signalled by a separate gesture, which is the feature’s main design risk.

With reasoning on, three categories of task open up:

  • Editing selected text in place โ€” tone changes, rewrites, reformatting.
  • Extracting from selected files โ€” highlight documents, PDFs, or images on the desktop and ask for a summary or specific details.
  • Generating images by voice description, inserted into the active window.

Google documents one non-obvious sequencing rule: to give Gemini context from a selection, highlight it first, then hold Fn and speak, then deselect after Speak to Window has started โ€” otherwise Gemini will overwrite the selection with its output instead of using it as input.

The two modes side by side

 Intelligent dictationReasoning
Enabled byDefaultSettings > Speak to Window > Use reasoning
Sent to GoogleAudio onlyAudio plus screen or selection content
OutputTranscribed text at cursorText, edits, or generated images
Reads selectionsNoYes
Failure modeBad transcriptionInstruction typed as text, or text executed as instruction

What the permissions actually grant

Reasoning mode is gated by macOS privacy permissions, not just the in-app toggle. The frontmost window is the default unit of context. To let Gemini read full browser pages beyond the visible viewport, or files across a connected local folder, you grant Accessibility to Gemini in System Settings > Privacy & Security. Screen Recording may also be requested depending on the type of context being shared.

These are broad permissions. Accessibility in particular is the permission that lets an application read and control the contents of other applications โ€” it is the same grant a password manager or a window manager asks for. The practical consequence is that any window you have in the foreground when you trigger reasoning is eligible to be sent to Google’s servers for processing. Nothing here runs locally.

If you handle client or regulated data

Leave reasoning off and use dictation only. Dictation gives you the majority of the day-to-day value โ€” fast, clean transcription anywhere โ€” without granting screen access. Turn reasoning on deliberately for a session where you need it, then turn it back off. There is no per-app allowlist.

Three separate ways Gemini gets your screen

Speak to Window is one of three screen-context paths in the Mac app, and they are easy to confuse because they overlap:

TriggerWhat it does
Option + SpaceOpens the mini chat window. Option + Shift + Space opens the full chat.
Command + CommandPulls the frontmost window into the chat as context, then you type.
Fn (long-press or double-tap)Speak to Window. Voice in, result dropped at your cursor without visiting the chat.

The distinction that matters: Command + Command brings your screen to Gemini’s window. Fn brings Gemini’s output to your window. If you want to read and iterate on a response, use the chat. If you want a result deposited where you are working, use Fn.

Against macOS built-in dictation

Apple’s dictation is on-device for many languages, works offline, has no account requirement, and transcribes literally โ€” including your “ums” and your abandoned half-sentences. Gemini’s dictation is cloud-processed, requires a signed-in Google account and a connection, and edits as it transcribes.

That editing is the actual differentiator, and it cuts both ways. For drafting prose from spoken thought, cleaned-up output saves a pass. For dictating anything where exact wording matters โ€” quotes, code, names, legal text โ€” a model that silently “fixes” your speech is a liability, and Apple’s literal transcription is the safer tool.

Shortcut conflict: both features live on the same physical key. macOS assigns a Fn/globe-key shortcut to its own dictation, configurable under System Settings > Keyboard > Dictation. If Speak to Window fires inconsistently, this is the first thing to check โ€” change one of the two bindings rather than trying to make them coexist.

Against ChatGPT Voice on desktop

OpenAI shipped ChatGPT Voice to the desktop app in late July 2026, days before Google’s announcement. The two features look adjacent and are built on opposite premises.

ChatGPT Voice is a conversation. Powered by GPT-Live, it speaks and listens simultaneously, and its stated purpose is directing work โ€” controlling the computer and coordinating agents running in ChatGPT Work or Codex. Its macOS Screen Context feature captures an appshot of the frontmost window, including text outside the visible scroll area, when you opt in and ask it to look.

Speak to Window is a transaction. Press, speak, release, receive. There is no back-and-forth, no spoken reply, and no agent to direct โ€” Gemini’s autonomous work on the Mac lives in a separate feature, Gemini Spark, which is gated to Google AI Ultra subscribers in supported countries.

 Gemini Speak to WindowChatGPT Voice on desktop
InteractionOne-shot: speak, release, get outputContinuous spoken conversation
Spoken repliesNoYes
Where output landsAt your cursor, in the active appIn the ChatGPT app
Screen contextOpt-in; frontmost window, wider with AccessibilityOpt-in; appshot of frontmost window including off-screen text
Plan requirementFree tier includedPlus, Pro, Business, Edu, Enterprise

They are not really competing for the same slot. If you want to talk through a problem, ChatGPT’s model fits. If you want text to appear in the document you already have open, Gemini’s does.

Requirements and limits

  • Apple Silicon only. Intel Macs cannot run the app at all.
  • macOS Sequoia (15.0) or later, 8 GB RAM minimum, 200 MB disk.
  • English only at launch, rolling out to all users of the Mac app. Google says more languages are coming without committing to a date.
  • Always online. Both modes are cloud-processed; there is no offline fallback.
  • Personal Google account, or a work or school account where an administrator has enabled Gemini Apps.

The Apple Silicon requirement is the one that quietly excludes people. This is not a feature you can evaluate on an older Mac.

Where it goes wrong

Either reasoning is off, or the classifier read your utterance as dictation. Check Settings > Speak to Window first. If reasoning is on, phrase instructions imperatively and reference the selection explicitly โ€” “rewrite this paragraph as bullet points” classifies more reliably than “maybe bullet points would be better here.”

The selection was still active when the output landed. Highlight, start Speak to Window, then deselect while it is listening. Gemini keeps the context and no longer has a selection to overwrite.

Without Accessibility permission, context is limited to what is visible in the frontmost window. Full-page browser reading and shared local folder access require enabling Accessibility for Gemini in System Settings > Privacy & Security. Decide whether the task justifies the grant.

Shortcut collision. Check System Settings > Keyboard > Dictation for Apple’s binding and Settings > Speak to Window for Gemini’s, and change one. Also confirm Gemini is actually running โ€” the feature depends on the menu bar app being active, not just installed.

Hold-to-talk is designed for short bursts and ends the moment you release. For anything sustained, double-tap Fn for hands-free recording and tap again to finish.

A reasonable way to set it up

  1. Install and use dictation only for a week. It requires no permissions beyond the microphone and tells you whether the cleaned-up transcription suits how you actually speak.
  2. Resolve the Fn conflict with Apple dictation before deciding the feature is unreliable.
  3. Turn on reasoning only once you have a specific recurring task for it, and only if you are comfortable with the frontmost window being sent to Google.
  4. Grant Accessibility last, and only if window-scope context proves insufficient.