Show it. Say it.
Your agent gets both.
Screenshot, recording or dictation — one key, and it lands where your coding agent reads it.
macOS 26 or later · Apple Silicon · Free during early access, no account
Dictation into the text field
Speak instead of type — right where the cursor is. In Claude Code, the terminal, VS Code, Mail. Cueshot strips the filler words and puts the punctuation back.
Where it helps
Explaining takes longer than showing.
You see it. Your agent doesn't. Three keys are all that stands between the two.
Say it
Hold one key and talk. The text appears where the cursor already is — in your agent, your terminal, your mail.
Show it
The screen freezes, you drag a box around what matters. One key later it is in the chat you are writing in.
Record it
Record, talk, and mark the spot in red while it happens. Your agent gets the picture, the voice and the mark.
What your agent gets
A model can't watch a video.
Forty seconds of screen recording are worthless to an agent as long as they stay a video. And even if you sent every frame, it would be thousands of them, most of which look alike. So Cueshot's real work happens after the recording.
| Without Cueshot | With Cueshot |
|---|---|
| 40 seconds, 1,200 frames — you find the moment yourself | 3 to 6 frames, exactly where something happened |
| Whatever you said is gone | Transcript with timestamps, recognized on device |
| Your box is a red rectangle in the picture | The box is in the picture and as coordinates in the manifest |
| Megabytes of video on disk | 2,000 to 17,000 tokens in the context window |
A frame is kept when it explains something: because you marked there, because the picture changed noticeably, or because it's the start or the end. Everything else is dropped. And then all you say is: "Take a look at the last capture."
The keys
Three keys, nothing else
Screenshot with a note
The screen freezes. Drag a box, draw on it, hit Enter. What you marked reaches your agent as a coordinate — not just as a red line.
Recording with your voice
Show it and talk while you do. Cueshot writes it down and remembers when you marked. The recording becomes the moments that mattered.
Dictation into the text field
Speak instead of type — right where the cursor is. In Claude Code, the terminal, VS Code, Mail. Cueshot strips the filler words and puts the punctuation back.
If you prefer key combinations, you can remap them in Settings → Shortcuts.
On your Mac
Your work stays where it is
There is no server behind Cueshot. What you record is written into your own user folder and read from there — by you, and by the agent on this machine.
Your recordings never leave your device.
Screen, audio and transcript are processed on your Mac and live in your own user folder, readable by your account only.
Speech recognition runs here too.
On your Mac, not in a cloud. No audio leaves the machine — not even to be recognized. The model is downloaded once and then works offline.
Password fields are exempt.
With secure input active, Cueshot inserts the text and stores nothing.
Exactly one thing can go out.
Anonymous timings — how long recognition took, never a word of what you said. Off until you say yes, and one click turns it off again.
That isn't a promise, it's how it's built: there is no server to send a recording to.
Show it. Say it.
Free during early access — every feature, no account, no card. A paid plan comes later; if you're here now, you'll hear first.
And it keeps growing: everything that makes working with an agent feel less like typing and more like pointing at something is on the list.
macOS 26 or later · Apple Silicon
Questions