Tom: Transcribe Audio & Video
Tom: audio & video naar tekst
Tom: Transcribe Audio & Video is #38 in Productivity Paid in Netherlands and charting in 1 of 24 countries we track.
DeveloperMarco Tini
CategoryProductivity
Price9.99 EUR
Released
Rating★ 0.0 (0)
Version1.0.7
Age rating4+
First charted🇳🇱 25 Sep 2026
Where it ranks now
Every top chart it appears in across 24 countries.
| Country | Charts | Rank |
|---|---|---|
| 🇳🇱 Netherlands | ProductivityPaid | #38 |
Members see more
Sign inEstimated downloads and revenue, ratings momentum, rank history, recent reviews per country and breakout scores for Tom: Transcribe Audio & Video.
About Tom: Transcribe Audio & Video
Transcribe audio & video with the engine that fits the job — three options, all 100% on-device.
Tom turns sound into clean, editable text right on your iPhone. Lectures, WhatsApp voice notes, screen recordings, interviews, podcasts, long meetings — pick your engine, drop in the file, read the transcript. No account, no upload, no subscription.
Built originally for a friend who is deaf, Tom is for anyone who needs to read what was said.
CHOOSE YOUR ENGINE
• Apple Speech — fast, low memory, built into iOS. Great default for live recording and short files. 25 languages.
• Whisper — OpenAI's recogniser, running locally. Six sizes from Tiny (80 MB) to Large v3 (3.1 GB), including Turbo for the best speed/quality balance. Up to 99 languages.
• Parakeet v3 — high-accuracy European transcription from FluidAudio, on-device. 25 European languages. Apple Silicon, iOS 17+.
Not happy with the result? Re-transcribe the same file with a different engine in one tap — the previous version is kept in the version history, so you can compare and switch back.
KEY FEATURES
• Live recording — tap once and watch the transcript appear as you speak.
• File import — any video from your library, or audio files from Files (including WhatsApp voice notes saved to Files).
• Multiple engines + re-transcribe — pick Apple Speech, Whisper, or Parakeet per file, and re-run the same audio through a different engine without losing the previous transcript.
• Editable segments — long-press any segment to fix a word the recogniser got wrong. Your edit stays on the device.
• Subtitle export — save any transcript as plain text, SubRip (.srt) or WebVTT (.vtt) for accessible video.
• Player with smart controls — tap-to-seek, follow the highlighted segment as it plays, skip ±10 seconds, and switch speed to 1.5x or 2x.
• Honest hallucination flag — when Whisper invents text on silence or music, suspect segments are clearly marked so you can verify or remove them. We never silently strip your transcript.
• Long files — chunked processing for meetings, lectures, conferences and full episodes.
• Model management — downloaded Whisper and Parakeet models live in Documents, visible in the Files app. Delete a model from Settings or from Files to reclaim space anytime.
• Wi-Fi only downloads — optional, on by default, so large models never eat your mobile data.
• Smart storage — save transcriptions with custom titles, search by title or content.
• Activity log for transparency on every step.
WHO IS IT FOR
• Students extracting notes from long lectures, videos, or screen recordings.
• Professionals transcribing interviews, meetings, or lengthy audio files.
• WhatsApp users who want to read voice messages instead of listening.
• Content creators turning audio or video into text or subtitles for TikTok, Reels, Shorts, or blog posts.
• Podcasters making episodes accessible with .srt subtitles.
• Anyone who is deaf or hard of hearing and needs reliable, private transcription.
WHY CHOOSE TOM
Mos
Tom turns sound into clean, editable text right on your iPhone. Lectures, WhatsApp voice notes, screen recordings, interviews, podcasts, long meetings — pick your engine, drop in the file, read the transcript. No account, no upload, no subscription.
Built originally for a friend who is deaf, Tom is for anyone who needs to read what was said.
CHOOSE YOUR ENGINE
• Apple Speech — fast, low memory, built into iOS. Great default for live recording and short files. 25 languages.
• Whisper — OpenAI's recogniser, running locally. Six sizes from Tiny (80 MB) to Large v3 (3.1 GB), including Turbo for the best speed/quality balance. Up to 99 languages.
• Parakeet v3 — high-accuracy European transcription from FluidAudio, on-device. 25 European languages. Apple Silicon, iOS 17+.
Not happy with the result? Re-transcribe the same file with a different engine in one tap — the previous version is kept in the version history, so you can compare and switch back.
KEY FEATURES
• Live recording — tap once and watch the transcript appear as you speak.
• File import — any video from your library, or audio files from Files (including WhatsApp voice notes saved to Files).
• Multiple engines + re-transcribe — pick Apple Speech, Whisper, or Parakeet per file, and re-run the same audio through a different engine without losing the previous transcript.
• Editable segments — long-press any segment to fix a word the recogniser got wrong. Your edit stays on the device.
• Subtitle export — save any transcript as plain text, SubRip (.srt) or WebVTT (.vtt) for accessible video.
• Player with smart controls — tap-to-seek, follow the highlighted segment as it plays, skip ±10 seconds, and switch speed to 1.5x or 2x.
• Honest hallucination flag — when Whisper invents text on silence or music, suspect segments are clearly marked so you can verify or remove them. We never silently strip your transcript.
• Long files — chunked processing for meetings, lectures, conferences and full episodes.
• Model management — downloaded Whisper and Parakeet models live in Documents, visible in the Files app. Delete a model from Settings or from Files to reclaim space anytime.
• Wi-Fi only downloads — optional, on by default, so large models never eat your mobile data.
• Smart storage — save transcriptions with custom titles, search by title or content.
• Activity log for transparency on every step.
WHO IS IT FOR
• Students extracting notes from long lectures, videos, or screen recordings.
• Professionals transcribing interviews, meetings, or lengthy audio files.
• WhatsApp users who want to read voice messages instead of listening.
• Content creators turning audio or video into text or subtitles for TikTok, Reels, Shorts, or blog posts.
• Podcasters making episodes accessible with .srt subtitles.
• Anyone who is deaf or hard of hearing and needs reliable, private transcription.
WHY CHOOSE TOM
Mos
Latest updates
1.0.7
This update includes bug fixes and performance improvements.