I built a private, on-device meeting scribe for Mac. Optional cloud is opt-in and anonymized first
I'm the developer, and a laser physicist. Most of my meetings are under NDA, and I wasn't going to hand that audio or the transcripts to a server I don't control, so I built Thoth to keep the whole thing on my Mac. Posting for feedback, not a launch. Problem Most AI meeting tools do the transcription or the summary in the cloud. Even the bot-free ones that grab system audio locally still ship the transcript and notes off to their servers for the AI step, and store them there. For anything under NDA that's a non-starter. I wanted the whole thing, recording, transcription, summary, to run on the device, and only touch the network if I decide it should. What it does Records your mic and the system audio as two separate tracks. So it gets both sides of a call with no bot joining, and the you-versus-them split is deterministic because each track is one source. Speakers get labelled on-device with Pyannote (SpeakerKit). Transcribes on-device with WhisperKit, Parakeet, or Apple's SpeechAnalyzer, and summarizes with a local model. No account, no backend. The way I prove it's local is by turning Wi-Fi off. Thoth records; transcribes & summarizes anyway (see attached video). If you do want a stronger cloud summary, it's your own key (OpenAI, Anthropic, Google) or Tinfoil's verified enclave. Either way a local pass masks names, places, and organizations before anything leaves, so you decide what goes. With Tinfoil it even surfaces the enclave's cryptographic proof, so you can check it instead of trusting me. Comparison Closest well-known app is Granola, and it's genuinely good, more polished than mine on the notes and team side. If you want a cloud notepad with integrations, it's a strong pick. The real difference is architecture. Granola transcribes locally, then sends the transcript and notes to its cloud for the AI step and keeps them there. Thoth never does that off-device unless you tell it to. By default nothing leaves, and it'll run with Wi-Fi off. If you do reach for cloud to get a better summary, it's your key or a verifiable enclave, and the anonymization runs before the send. On speaker labels, the you-versus-them split is deterministic here because the two sides arrive as separate tracks. Sorting out several people is Pyannote, on-device, model-based like everyone else's, so I'm not pretending I'm ahead there. Still working on it. And nothing joins your call as a bot, same as Granola. That's the one thing Otter, Fireflies, and Fathom all make you accept. Pricing Free tier: recordings capped at 30 min (mic) and 15 min (system or mixed), plus 20 AI actions a month. Paid: $7.99/month, $44.99/year, or $59.99 once for lifetime. Monthly and yearly come with a 1-month free trial. Mac App Store: https://apps.apple.com/app/id6756965313 More detail and privacy policy: https://thoth-app.com Mostly I want feedback.
已收集讨论
It's looks cool
Thanks :-)
looks amazing did you make the crow-man animation for loading yourself?
Thanks ! Partially, i used NanoBanana for the first frame, edited it in inkscape to clean it up and convert to vector then made the following frames (4 total)
Big fan of local ai nicely done
Thank you :-)
Looks awesome, love the idea of private transcriptions! How does it compare to Voice Memos transcriptions?
Thanks! TBH, I haven't benchmarked Voice Memos against Thoth on the same recording, so I won't fake a number. I'll do it once I get through a backlog that's gotten a bit long :p and I'll add Apple Speech to my public benchmark while I'm at it. Quick answer tho: one of Thoth's engines is Apple's SpeechAnalyzer, the same one Voice Memos uses, so on your own voice it should be similar. You can also switch to Whisper or Parakeet. Either way, both tools are built for different jobs. Voice Memos records one source, your mic, so on a call it only catches your side. Thoth records your mic and the system audio separately, so it actually captures the other people too, works out who said what, summarizes it, and redacts sensitive stuff before any cloud step.
This is great. how is the accuracy on this compared to say whisperAI or using macwhipser? curious if theres a meaningful quality gap?
Thanks :) TBH, i think similar accuracy. MacWhisper & Thoth apps both run the same transcription models, so i don't think there's a meaningful quality gap on transcription itself. I also suppose they do audio processing like me, but i don't know the depth of it. Where Thoth differs isn't accuracy, it's what happens around it: it records your mic and system audio as separate tracks, so the you-versus-them split is deterministic instead of diarized from one mixed stream; and it has an anonymization pass that masks names, IDs, and sensitive terms on-device before any cloud step, with private-enclave (Tinfoil) BYOK if you want a stronger cloud model.
Nicely done
Thank you :-)
I work in healthcare and academia, so everything I record touches HIPAA and FERPA. My current workflow is Spokenly free tier for post-meeting transcription, then a local Presidio script to redact PII before anything goes near a cloud AI. Local-first is a hard requirement for me, which is why Thoth caught my attention. Two questions: 1.How does diarization hold up on real Zoom calls with 4-6 remote participants? Since all remote voices arrive as one compressed system-audio stream, I’m curious what speaker confusion looks like in practice versus your clean-audio benchmarks. Any plans for voice profiles so recurring speakers get labeled automatically across recordings? 2. Can I select and copy text from the live transcription panel while recording is still in progress, or do I need to stop first? I sometimes want to paste a section into another AI tool mid-meeting for quick analysis. I appreciate the transparency in your published benchmarks. The accented-speech numbers are the kind of honesty I don’t usually see from transcription apps.
Thanks for your comment and for reading my benchmarks :-) On your two questions: 1 Your mic and the system audio are captured as two separate tracks, so the you-versus-them split is deterministic and never confuses. If you need to separate speakers on each stream, Pyannote is what Thoth (and all the industry) uses so i won't expect a massive difference there. Audio capture is what actually matter. For the system stream, not much can be done and Zoom, teams etc... already a good job. On the mic stream, Thoth uses voice isolation and auto gain, to put all chances on user's side for a great audio for transcription and diarization. You can even check the Fourier Transform after the recording to check the quality. Overall, yes, 4-6 overlapping compressed voices on the system stream do make diarization harder, that part is physics, we can only hope Pyannote gets better as time goes by :-) 1-bis Voice profiles for recurring speakers: not yet. What I get now is diarization, anonymous Speaker A/B/C within one recording, that you can rename. Speaker identification, e.g. enrolling a voiceprint once and recognizing that person across future recordings, is on Argmax's roadmap, but it isn't something I ship today so I won't claim it. Definitely adding it once they ship it ! 2 mid-recording copy: yes. The live transcript panel is selectable while recording is running, so you can grab a section and paste it elsewhere without stopping the session. you can even pin the live transcript window so it stays on top of every window. On anonymization, the redaction that runs before a cloud send is four deterministic on-device passes: NER for names, places and orgs (Apple's NaturalLanguage), NSDataDetector for phones, emails, URLs and addresses, your accepted keywords, and renamed speaker labels. On top of that there are two on-device suggestion engines that scan the transcript and propose sensitive terms as chips you accept, one via Apple's Foundation Models framework, one via a verified zero-retention Tinfoil enclave with fail-closed attestation for higher recall. Honest limit for your case: those suggesters find terms, but an accepted term is still a literal match for that value, and there's no way yet to define a standing format rule, so a structured identifier like an MRN or patient ID isn't caught deterministically or carried to the next recording the way your Presidio patterns are. For anything where a single miss isn't acceptable I'd keep Presidio in front of the cloud path. That's the one place a user-definable pattern rule would close the gap, and it's on my list.
Thank you for taking the time to write such a detailed and honest answer. That alone tells me a lot about how you build this app. Downloading it now to test. I'm not a developer, so I appreciate that you were upfront about the limits instead of overselling. The one thing I'll add as a real-world use case: I work in healthcare, medical research and education, so I deal with structured identifiers like medical record numbers and student IDs. Right now I run everything through a separate local script that catches those by pattern before anything goes to a cloud AI. If Thoth ever lets users define their own patterns for redaction, like "always catch any 7-digit number," it would cover that gap and I could probably drop the extra step entirely. Sounds like it's already on your list, so consider this one vote from a HIPAA/FERPA user who would use it every day. Again, thank you. Apps like this are exactly why I keep looking for local-first tools. I'm definitely putting the word out to colleagues about how this could fit their workflows.
Congrats, this looks very polished. It's nice to see more apps use a BYOK + Local AI approach. That's what I use for my apps, too. Out of curiosity: For your local model, do you include and host this as part of the app itself, or do you integrate with a host app like LM Studio, Ollama, or Llama.cpp?
Thanks ! :-) Both, user can connect to LM Studio local host and download local models (Phi, Gemma ; Llama ; Qwen)
Nice, I like that approach. Gemma 4 is my go-to, but Qwen is solid as well.
the on-device-by-default architecture is the right call for NDA work, the cloud tools make you just trust them and that's not good enough. What STT are u using? Are you satisfied with the quality? I tried various on-device models, and I didn’t like the result tbh, hopefully I don’t need this level of privacy, so I’m good using cloud tools
It uses Whisper, parakeet and Apple Speech, user can choose which ones. They're honestly good, especially if end goal is to extract a summary and action items. I used them at work meetings (dense, technical R&D work, accented english and the results were good, not perfect WER but hnestly good enough for my use case) What really improves is audio capture, eg placing the device where it can capture everyone, or using bluetooth capture tools is what genuinely helps
Looks nice! Is there a difference in features between subscription and one-time purchase?
No difference, all unlocks the same features
on device by default is the right call, the meeting apps that upload everything are a privacy nightmare for the other people in the call that never accepted nothing. the anonymized cloud part, what exactly it sends?
it sends the transcript with tokens thet replace sensitive information, like [PERSON_1] instead of Jane Doe, [ORG 1] instead of Microsoft; [PLACE_1] instead of Cupertino... Her's a screenshot as example
Thoth - Your Private AI Scribe — by Matthieu Veinhard Ratings: 0 (no ratings yet) Average: n/a ★ Age: released ~3 months ago · rated 17+ Overall score: 0 / 100