# Build brief — a focused alternative to Superwhisper

> **Verdict:** Yes, for personal use · **Buildability:** 90/100 · **Category:** Voice Dictation
> **Source:** https://www.canitbevibecoded.com/superwhisper
> Independent editorial assessment from Can It Be Vibe Coded? Not affiliated with, endorsed by, or derived from Superwhisper. Verify current pricing and capabilities before acting.

## Context

**Superwhisper** — AI dictation for Mac that turns speech into text anywhere. It currently costs $8.49/mo.

A push-to-talk recorder that transcribes and pastes text into the active app is very buildable for one person, especially on macOS.

This brief describes a focused, single-operator replacement for the part of Superwhisper that is genuinely reproducible. It is deliberately narrower than the product it replaces, and it says so in writing. Build the useful core; do not pretend to have rebuilt the rest.

## What you are building

Bind a hotkey, capture microphone audio, transcribe with Whisper or speech API, optionally rewrite with an LLM, and paste into the active field.

- Build a focused single-user workflow with real persistence, search, and export.
- A responsive interface with real empty, loading, success, and error states.

## Requirements

### Functional

- MacOS automation/accessibility permissions.

### Data and integrations

- Speech API or local Whisper.
- Optional LLM API.
- Hotkey library.

Each of these needs a real account, credential, or quota. Set them up before writing feature code.

### Non-functional

- Accessibility: semantic markup, labelled controls, visible focus, and reduced-motion support.
- Security: server-side secrets, validated input, and no credentials in the client bundle.
- Reliability: retries with backoff on external calls, and a clear failure state when a provider is down.
- Portability: the operator can export their data and leave without losing it.

## Implementation brief

Build me a push-to-talk dictation tool for macOS to replace Superwhisper. Requirements:

- A global hotkey (default: hold right Option) records my mic while held, stops on
  release. A small Swift menu bar app or a Hammerspoon script, pick the simpler to
  ship.
- Record with ffmpeg (avfoundation) to a temp wav, transcribe locally with whisper.cpp
  (small.en by default, model path in a config file). Works fully offline, no cloud
  speech APIs.
- Paste the result into whatever field has focus (simulate Cmd+V via CGEvent or
  osascript, then restore my previous clipboard).
- Optional cleanup mode on a second hotkey: send the transcript to an LLM (key in
  .env) to fix punctuation and drop filler words, then paste. If no key is set, this
  mode just does a plain paste.
- Menu bar icon shows idle/recording/transcribing; clicking it lists the last 10
  transcripts with copy buttons.
- Append every transcript to ~/Dictation/YYYY-MM.md with a timestamp, and delete the
  audio after transcription. No accounts, no telemetry.
- Out of scope: per-app presets and custom vocabulary tuning. One good general mode.
- README: mic + accessibility permissions to grant, how to download the whisper
  model, and a note that the first run will trigger macOS permission prompts.

## Delivery standard

- Inspect the repository first, then write a short implementation plan before writing code.
- Deliver the smallest complete end-to-end workflow first; every primary control must work against persisted data.
- Use real validation and storage; never substitute fake dashboards, decorative controls, hard-coded success states, or mock integrations.
- Include responsive layouts plus genuine empty, loading, success, validation, and failure states.
- Keep secrets server-side in environment variables, provide .env.example, and never commit credentials or user data.
- Add structured logs around every external call and return actionable errors without leaking sensitive details.
- Write unit tests for the core logic and one automated test of the main user journey.
- Finish with a README covering setup, architecture, data location, backups, tests, deployment, and known limitations.

## Acceptance criteria

- [ ] A clean install starts the app using only the README and .env.example.
- [ ] The primary journey works from first visit through saved result, reload, edit, export, and deletion where applicable.
- [ ] Invalid input, missing configuration, provider failure, and an empty database each have a usable state.
- [ ] The interface works at 390px and 1440px, is keyboard navigable, and shows visible focus on every control.
- [ ] Tests, type checking, linting, and a production build all pass with no ignored failures.
- [ ] No part of the interface implies a live integration, security guarantee, or scale capability that was not actually built and verified.

## Non-goals

Do not build these, and do not claim to have replaced them:

- Beautiful native UX.
- Presets.
- Vocabulary/profile tuning.
- App-wide polish.
- The last 20 percent is sync, migration fidelity, speed, and edge cases.

## What you still own after launch

- Secure credentials, rotate secrets, and handle provider rate limits.
- Run migrations, backups, restores, and dependency updates.
- Test the critical journey after every model, API, or hosting change.
- Monitor failures and fix the edge cases a first prompt will miss.

## Risk

**Manageable.** A personal version is realistic if you test the critical journey and keep reliable backups.

Editorial confidence in this assessment: high. No independent one-shot implementation is linked yet.

## Prior art

Working open-source software you can read, fork, or borrow from before starting:

- [whisper.cpp](https://github.com/ggml-org/whisper.cpp) — Local transcription engine that makes a simple dictation clone realistic

---

Generated by [Can It Be Vibe Coded?](https://www.canitbevibecoded.com) · Full report: https://www.canitbevibecoded.com/superwhisper
