Skip to content

Recordings and AI

Speech becomes text locally, on your device. Audio does not go to the cloud.

Several recognition models are available with different speed/accuracy trade-offs; the model is chosen at import, and your own recordings use the configured default.

The live transcript during recording and the final one after stopping are different passes: the first is optimised for latency, the second for quality.

The summary is built from the transcript and answers “what was this about”: topics, decisions, commitments. It starts automatically or from a button on the recording.

The summary is saved as a separate note marked as automatically created. You can edit it — it’s an ordinary note.

As with notes, a short description and tags can be generated for a recording. Both need finished text: with no transcript yet, they say so.

When microphone and system audio are recorded at once, the speaker split comes from the channels rather than being guessed from voices. That’s more reliable than any clustering: a channel is a fact, not a hypothesis.

Speakers can be renamed — real names instead of “Microphone” and “System”. The labels apply across the whole transcript.

  • Find a recording by meaning (“the conversation about the March budget”).
  • Answer from its content, citing the fragment.
  • Work through a meeting — pull out decisions and open questions and propose tasks. Tasks are created only after you confirm.

Audio and recognition never leave the device. Text may go to the cloud — for summaries and agent answers, if cloud features are connected — and passes sensitive-data masking before it does. Without the cloud, summarisation runs on the local model.