Desktop and phone
Recording and transcription
Capture a thought out loud, keep the audio, and turn it into text on your own terms.
You can record straight into a note, or open the Record view for a longer session. Either way the audio is yours: it is saved to your device, it plays back inside the note, and transcription only happens if and where you tell it to.
Recording
- 1Use the microphone button in a note to record in place, or open Record for a standalone session.
- 2Speak. A long recording is split into chunks and written to disk as it goes, so nothing is lost if the app stops.
- 3Press Stop when you are done, or Discard to throw it away.
There is no pause on the recording itself
The controls are Stop and Discard. The Pause and Resume buttons you may see belong to the transcription job that runs afterwards, not to the microphone.Very long recordings become books
A recording of a couple of hours or more is turned into a Library item with chapters instead of one enormous note, so it stays navigable. See Library and reader.Turning speech into text
Transcription is a separate step and it is optional. There are two routes, and they differ by platform.
| Route | Desktop | Phone | What it needs |
|---|---|---|---|
| On device | Yes | No | A sherpa-onnx recogniser downloaded through Settings ▸ Speech to Text. Nothing leaves the machine. |
| Your own server | Yes | Yes | A Whisper-compatible endpoint you host. Audio is sent to that address and nowhere else. |
| Live streaming | No | Yes | Text appears as you speak, rather than after you stop. |
Only two on-device models are currently offered
The desktop model catalogue lists four, but only entries with a verified checksum are made available for download, and two are not yet pinned. This is deliberate: a download that cannot be verified is refused rather than shipped with checking switched off.Playback alongside the transcript
A recording embedded in a note plays inline, directly beside the text it produced. That means you can check a wording against what was actually said without leaving the note.
Yes. Recording and playback are entirely local. Only server-based transcription needs the network, and on-device transcription on the desktop does not.
