Skip to content

Recording and transcription

Recording starts from the recording panel and can be paused and resumed. An indicator is visible throughout — the app does not record quietly.

A live transcript appears as people speak. It is a draft: the final text is computed after you stop, and it is more accurate.

If the microphone is connected but there is no sound, Yttri warns about “silence in the microphone”. That catches the most common failure — an hour of emptiness recorded from the wrong input device.

+ New recording at the bottom of the submenu opens a dialog with two tabs — URL and File — and a choice of speech recognition engine. Links to external sources are supported: YouTube, SoundCloud and others.

While a recording is processed, a “Processing (N)” block appears above the list with progress per job. The stages are named: queued, downloading, decoding, transcribing, summarising.

All of it runs locally on your machine.

A card shows the title, the description, a relative date and the duration. While speech recognition runs, a spinner labelled “transcribing” sits next to it. A single click opens a preview, a double click opens its own tab.

The context menu is shorter than a note’s: open in a new tab, move to a collection, rename, delete. Recordings cannot be duplicated.

Transcription starts automatically. If it failed or the quality disappointed, it can be run again — the process is resumable, so it does not start from scratch.

The engine is chosen when the audio is imported; for Russian and English there are several models with a different balance of speed and accuracy.

The transcript can be corrected by hand — a name, a term, a slip. Edits are saved and used downstream: the summary is built from the corrected text.

The transcript and the summary can be exported separately. The audio remains a file in the data folder and can be taken from there directly.