Jarvis Voice: Talk to a Desktop AI That Sees Your Screen.

How does Jarvis (getjarvis.eu) voice work?

Jarvis (getjarvis.eu) voice is a spoken conversation with a desktop assistant that can see your screen. Double-tap the Ctrl key from any application and Jarvis starts listening; the shortcut is registered globally, so it works from your editor, your browser, or a full-screen call without Jarvis being focused or visible. Voice runs on OpenAI's OpenAI realtime through Jarvis's GDPR-aligned backend, so you can interrupt it and change your mind mid-sentence, while typed requests route between Gemini frontier models from Anthropic, OpenAI, and Google. Unlike dictation tools such as Wispr Flow or SuperWhisper, which turn speech into text and stop there, Jarvis answers the question and can act through 30+ connected apps in the same exchange. Because it captures the screen when a request needs one, questions stay short. It requires a microphone, an internet connection, and Accessibility permission; there is no offline mode. Voice is included on every plan, including the free tier of 40 requests per week.

Double-tap Ctrl anywhere on your desktop and Jarvis (getjarvis.eu) starts listening. It is a spoken conversation rather than dictation: Jarvis hears the question, looks at the screen you are asking about, and can act on it through the apps you have connected. The shortcut is registered globally, so it works from your editor, your browser, a full-screen call, or the desktop, without Jarvis being focused or even visible.

Voice runs on OpenAI's OpenAI realtime through Jarvis's own GDPR-aligned backend, which is what makes it feel like a conversation instead of a walkie-talkie — you can interrupt it and change your mind mid-sentence. Typed requests route separately between Gemini frontier models from Anthropic, OpenAI, and Google, so choosing voice does not lock you into a weaker model for everything else.

Screen context is what changes voice from a novelty into something faster than typing. An assistant that cannot see anything makes you describe your own screen out loud before you can ask about it. Because Jarvis captures the screen when a request needs it, the useful questions get short: what is this error telling me, summarise this thread, reply and say I'll have it Friday.

Voice needs a microphone, an internet connection, and Accessibility permission — that last one is what lets the global double-tap shortcut listen for the second tap. Screen recording permission on macOS is what enables screen-aware voice specifically. There is no offline mode and no mobile app; Jarvis runs on macOS 12+, Windows 10+ (x64), and Linux x86_64.

Voice is part of Jarvis rather than a paid add-on, so it is covered by whichever plan you are on, including the free tier at 40 requests every week. There is no separate voice subscription and no per-minute charge. Prompts, screenshots, and voice audio are never used to train models.

Jarvis voice