个人效率Reddit 原帖

[OS] TypeWhisper 1.5.0 - open-source macOS dictation with local/cloud engines, Workflows, and app-aware insertion

完整上下文原始内容 · Reddit

Hi r/macapps, I build TypeWhisper, a macOS dictation and text automation app. I just shipped 1.5.0 as the new stable release. Problem I use dictation a lot, but the frustrating part is not just speech-to-text accuracy. It is everything around it: text insertion behaves differently in Terminal, Slack, Mail, browsers, writing apps, and Electron apps sometimes I want local-only transcription, sometimes a faster cloud/API engine is fine many dictation flows still leave me cleaning up punctuation, spacing, names, numbers, or formatting by hand advanced workflows often require copy/paste through several apps TypeWhisper is my attempt to make that workflow more native and controllable on macOS. What it does system-wide push-to-talk or toggle dictation local and cloud transcription engines file transcription for audio/video, including SRT/WebVTT export Workflows that can run by app, website/domain, hotkey, or fallback dictionary corrections, learned replacements, snippets, and searchable history plugin support, local HTTP API, CLI, widgets, and watch folders What changed in 1.5.0 The main 1.5 work was making day-to-day dictation less brittle: app-aware insertion with better spacing, sentence position, terminal paste behavior, and rich-text target handling improved local model handling, including MLX memory controls, idle auto-unload, and stalled model-download recovery better dictionary learning and number normalization broader provider and plugin coverage reliability fixes for hotkeys, indicators, recording, uploads, workflows, fullscreen apps, and plugin metadata One example setup: Terminal gets raw plain-text insertion, Slack uses a fast engine with auto-submit, and Mail can run a workflow that cleans up or rewrites the text before insertion. Comparison The closest apps people usually compare this with are Apple Dictation, Wispr Flow, Superwhisper, and MacWhisper. I think they are all useful, but they optimize for different people. Apple Dictation Apple Dictation is the best answer if you want something built in, free, and simple. For short messages it can be enough. TypeWhisper is for the cases where I want more control: choosing the transcription engine, keeping a searchable history, applying dictionary corrections, running Workflows, processing text before insertion, or changing behavior per app/site. Wispr Flow Wispr Flow feels very polished if your main goal is: speak naturally, have the app clean it up, and get finished writing everywhere across desktop and mobile. Their positioning is very consumer-product focused: fast dictation, AI commands, auto-edits, cross-device support, and a simple free-to-paid path. TypeWhisper is less about one managed voice-writing experience and more about control. It is open source, macOS-native, supports local and cloud engines, has Workflows, app/site routing, plugins, CLI/API automation, file transcription, dictionary learning, snippets, and local model controls. So I would frame it like this: if you want the smoothest "just talk and it writes nicely" experience across devices, try Wispr Flow. If you want a macOS dictation system you can inspect, route, extend, and automate, TypeWhisper is closer to that. Superwhisper Superwhisper is probably the closest philosophical comparison. It also focuses on voice-to-text in any app, offline and cloud recognition, custom AI modes, 100+ languages, and polished AI text transformation. It has a strong cross-platform story too. Where TypeWhisper tries to differ is the power-user surface: GPLv3 source, community plugins, bundled integrations, local HTTP API, CLI, watch folders, app-aware insertion rules, workflow triggers, and explicit control over local/cloud providers. I am trying to make it feel less like "a dictation box with AI modes" and more like a programmable text pipeline for macOS. Superwhisper is likely a better fit if you want a mature, polished, cross-platform dictation product with strong default AI modes. TypeWhisper is likely a better fit if you care about open source, local-first options, automation, and deeply configurable per-app behavior. MacWhisper MacWhisper is very strong if your main job is transcribing audio and video files on a Mac. It also has dictation, ChatGPT prompt processing, and app-specific prompts in the direct version. TypeWhisper overlaps with that, but its center of gravity is system-wide dictation plus automation: Workflows, app-aware insertion, provider choice, plugins, local APIs, dictionary learning, snippets, and live text insertion behavior. I still think MacWhisper is a great app if file transcription is the core use case. TypeWhisper is aimed more at people who dictate into many apps all day and want that behavior to be programmable. Short version Use Apple Dictation if you want built-in and simple. Use Wispr Flow if you want the most polished managed voice-writing experience across devices. Use Superwhisper if you want a mature AI dictation product with strong cross-platform support and AI modes. Use MacWhisper if file transcription is your main workflow. Use TypeWhisper if you want open-source macOS dictation with local/cloud engines, app-aware insertion, Workflows, plugins, CLI/API automation, and more control over what happens after you speak. The tradeoff is that TypeWhisper is more configurable. I am working on making the defaults feel simpler, but the app is intentionally built for people who want to tune their dictation workflow instead of treating it as a black box. Privacy Local transcription is available, and local prompt processing on Apple Silicon is available through verified Gemma 4 MLX models. Cloud transcription and LLM providers are optional and only used when configured. Pricing TypeWhisper is open source under GPLv3. Installing and running an unmodified GPL copy is free, including personal use and GPL-compliant internal/commercial use. Commercial licenses are for proprietary distribution, non-GPL use, procurement, invoicing, or support: Individual: 5 EUR/month or 99 EUR lifetime Team: 19 EUR/month or 299 EUR lifetime Enterprise: 99 EUR/month or 999 EUR lifetime Pricing page: https://www.typewhisper.com/pricing Links Website: https://www.typewhisper.com GitHub: https://github.com/TypeWhisper/typewhisper-mac Release: https://github.com/TypeWhisper/typewhisper-mac/releases/tag/v1.5.0 Homebrew: brew install --cask typewhisper/tap/typewhisper I would appreciate feedback from people who already use dictation heavily on macOS, especially around Workflows, local models, and app-specific insertion behavior.

01需求标签
个人效率其他桌面应用未解决

已收集讨论

25 条已收集

25条已收集43条 Reddit 标称评论
u/SeoFood楼主回复

FluidVoice looks great, especially if you want fast, free, local-first dictation with strong defaults. TypeWhisper is more about control: Workflows per app/site/hotkey, local or cloud engines, file transcription, dictionary/snippets/history, plugins, and API/CLI automation. I would not claim one is more accurate without a real benchmark. Short version: FluidVoice is the simpler local dictation app, TypeWhisper is the more configurable dictation/workflow system.

u/Brief-Mongoose-6256

I will try this app just because of the way you have presented it in your opening. Very balanced and very informative. Thank you for taking the time and good luck in your journey.

u/SeoFood楼主回复

Thank you, that means a lot. I’m trying to keep the project communication grounded: what works today, where the tradeoffs are, and what kind of setup the app is actually useful for. Voice tools can get overhyped quickly, so I’d rather be specific than make broad claims. If you try it and hit anything confusing, I’m happy to hear the feedback.

u/oddcouplings

I've started using MacOS dictation a lot recently, looking forward to trying this out

u/MaxGaav

I think TypeWhisper is a great app. Use it for a couple of months now with Parakeet. It's very fast and acively developed, with almost daily little updates available (if you want those). Btw, I wanted to install Nemotron Streaming 3.5 but I cannot find out how to do that. Or is this model not supported?

u/SeoFood楼主回复

Thanks! Glad Parakeet is working well for you. Nemotron Streaming 3.5 is not a built-in TypeWhisper model right now. The supported local NVIDIA path in TypeWhisper 1.5 is Parakeet TDT v2/v3. There is one possible workaround: if you have Nemotron exposed through an OpenAI-compatible speech-to-text endpoint, you can try it through the OpenAI Compatible plugin by adding the endpoint and model name there. But the raw Hugging Face / ONNX / CoreML Nemotron 3.5 model is not installable from the TypeWhisper UI yet. I am interested in adding it, though. If you can share which runtime you wanted to use, for example NVIDIA NIM, Together, ONNX, CoreML, or something else, that would help me pick the right integration path.

u/molx730

is this different from the seven same things i have seen this week?

u/Brief-Mongoose-6256

Thank you I will. Happy to support your work. Very few genuine developers left these days. I'm glad to find one.

u/SeoFood楼主回复

I’m still learning how to talk about the project properly. The building part comes much more naturally to me than the marketing part, and I’m trying to explain the tradeoffs clearly instead of making everything sound bigger than it is.

u/swiftidnc739

What does the animation look like if there is another app that takes control of the notch usually (e.g. Alcove)?

u/frskia

Nice release; app-aware insertion is the part most dictation tools get wrong Genuine question... do the local engines hold up on long recordings, or is it tuned for short dictation bursts? I work on the long-audio side (I'm the maker of Loreo, a different use case, 1 to 4 hour meetings and interviews with speaker labels), so I am curious where the local-vs-cloud line lands for you as length grows. For quick voice-to-text on the Mac this looks solid.

u/SeoFood楼主回复

For TypeWhisper, local engines are mainly tuned around day-to-day dictation and file transcription where the user wants control: local/offline when privacy or cost matters, cloud when they want faster throughput or do not mind using an API. For 1-4 hour meetings/interviews with speaker labels, I would not pitch TypeWhisper as the primary tool. That becomes a long-audio product problem: diarization, progress, retries, chunking, review UI, export formats, etc. TypeWhisper is more focused on “I’m typing into real apps all day and want voice input to behave correctly there.” So local can hold up technically, depending on model/hardware, but the product focus is different from something like Loreo.

u/mrterrycarson

Very Nice app, but LTD bit high for me to support. Any coupons would be great. Keep building my friend.

u/SeoFood楼主回复

Small licensing clarification: TypeWhisper is open source under GPLv3. You can use it for free for personal use or GPL-compatible open-source use. The paid license is only for proprietary / non-GPL commercial usage. So if you are unsure, just try it first. The LTD is there for people who need the commercial license or want to support the project long-term.

u/Brief-Mongoose-6256

I’ve tried way too many dictation apps over the past year. I even bought a lifetime license for Super Whisper, but it never really clicked for me despite all the hype. Then I went through Audio Pen, Whisper Flow, Typeless, and Fluid Voice. They are good apps, but something always felt off. I installed TypeWhisper yesterday, and I’m genuinely impressed. It just works. Fast, accurate, lightweight, and it stays out of your way. I’m not a power user or a techie—I just need reliable dictation, and this nails it. The developer also seems genuinely engaged and responsive, which is refreshing. The app is free unless you need a commercial license, and honestly it punches well above its weight. It’s only been a day, so I’m still exploring it, but my first impression is that this app deserves a lot more attention.

u/SeoFood楼主回复

Really appreciate this, and thanks again for taking the time to share thoughtful feedback. That “reliable dictation that stays out of the way” part is exactly what I’m aiming for. I’m still tightening the rough edges, so if anything feels confusing or breaks your flow as you keep using it, I’d genuinely like to hear about it.

u/AchoMatico

Wow that looks amazing! How does this compare to just using the built-in dictation feature?

u/SeoFood楼主回复

Apple’s built-in dictation is genuinely solid for short, simple dictation. TypeWhisper is different when you want more control: local engines like WhisperKit/Parakeet, optional cloud providers, history/recovery, dictionary/snippets, and Workflows that can behave differently in Mail, Slack, Terminal, the browser, etc. So it’s less “Apple Dictation, but slightly different” and more a configurable voice-input layer.

u/SeoFood楼主回复

Thank you, that means a lot. I think the main reason it feels useful is that I actually use TypeWhisper every day and need it myself. A lot of the work is just me running into the same rough edges in daily use, then fixing them.

u/RoderickHossack

I'm using this on windows, and saw a comment about being able to see what you say as you say it by enabling "streaming," but I can't find that option.

u/SeoFood楼主回复

On Windows, install and enable the Live Transcript extension first. Then go to Settings → Appearance and enable Live Transcript plugin. It shows a floating live preview while you speak; the final text is still inserted after you stop. Real-time updates also depend on whether the selected transcription engine supports streaming.