Speech becomes a place you can jump to

A recording used to be one result. It is now a set of moments. Audio tracks are transcribed automatically — from video files and from audio files alike — and the transcript keeps its timings. So a search that matches a spoken phrase does not hand you three hours of tape and wish you luck. It hands you the point in the recording where the phrase was said, and takes you there.
The transcript is visible where you already are. In the video player, captions are a track you switch on. In the audio player, the transcribed segments are drawn as timed regions on the waveform, so the shape of the recording and the shape of what was said line up. The full text can be read and copied from the item’s own panel.
And it is searchable like everything else. Transcribed text is indexed with the rest of the record: no separate search, no separate place to look. A query that matches caption text returns the item, and opening it highlights the matching phrases in the transcript. Where subtitles exist in more than one language, a match reports which language it matched in.
Captions and subtitles come out of the same pass, so a recording that has become searchable has also become watchable with its words on screen — and a person recognised in it is listed once per continuous appearance rather than once per frame, which is the difference between jumping and scrubbing.
The spoken word, working on a real clip, is on the discovery page.