VitaliWeb Tools

Audio & Video Transcription

Recognize English or German speech locally, review timestamps and export text or subtitles.

Processed on your deviceFree · No account needed

Preparing your tool…

How to use it

  1. Choose an audio or video file up to 50 MiB and 60 seconds. Review the detected tracks, select one decodable audio track, and choose English or German as the spoken language.
  2. Confirm the 66.2 MB download of verified engine and model files from this site, then choose Transcribe locally. Follow the actual download and recognition activity; Cancel stops the job, while Reset and release model clears the working session.
  3. Listen to the source from each segment’s start time. Correct the words, names, numbers and timing, and fill missing times yourself. Review repeated or partial phrases around overlapping recognition windows before confirming your review.
  4. Choose TXT, SRT or VTT, then use the separate Download link. Subtitle exports require valid times. Continue in Subtitle Repair opens the reviewed segments on the same page for shifts, drift correction and another subtitle export.

In a real German test, spoken ‘neun Uhr dreißig’ was recognized as ‘39 Uhr’. Correcting the segment to ‘9:30 Uhr’ changes the edited export. A missing end time stays empty until you enter a valid value from the source.

Details & limits

Up to 50 MiB / 60 seconds. Mono/stereo audio, 8–96 kHz; video up to 1,920 px per side / 2.07 MP. Explicit 66.2 MB engine/model download. Browser decoding and device memory required. Review names, numbers and missing timestamps; no Swiss German accuracy promise or cloud fallback.

Which files work, and how much can I trust the result?

Supported containers are WAV, MP3, Ogg, ADTS/AAC, MP4/MOV and WebM/Matroska; the browser must decode the actual audio codec. Inputs allow at most eight tracks and mono/stereo audio at 8–96 kHz. Video is limited to 1,920 px per side and 2,073,600 pixels. Recognition produces an editable English/German draft: names, numbers, Swiss German, music and silence can cause mistakes or invented text. Times are model estimates on the original timeline, not guaranteed alignment. Overlapping 30-second windows can retain duplicate or partial phrases; no text is automatically deleted. Limits are 200 segments, 20,000 characters overall and 2,000 per segment. Audio decoding stops after 60 seconds of work; model initialization and recognition have five-minute outer limits and may fail earlier. The model can need several hundred MB of RAM; mobile devices may be too limited. It is reused in the current worker session and must load again after reset or cancellation. TXT uses the edited words; SRT/VTT are escaped and parsed again, but that structural check does not verify what was spoken.