Can Cluely be vibe coded?
An always-on-top desktop overlay that watches your screen and listens to your calls, then feeds you AI answers in real time.
The core loop is genuinely thin: capture a screenshot, transcribe the mic, feed both to a multimodal model, stream the answer into a transparent always-on-top window on a hotkey. An agent gets that working in a session, and it will feel uncomfortably close to the demo. The gaps are the unglamorous parts: capturing the other side's audio takes virtual audio devices, latency has to beat the conversation, and the headline trick of staying invisible in screen shares depends on fragile platform window flags that vary by OS and conferencing app. Your DIY version will work fine when you are alone with your own screen, and may quietly betray you the one time it matters.
Jump to the build brief ↓Legacy-calibrated assessment
Checked Aug 2026
What you pay today, before any DIY hosting
medium editorial confidence
Tracked separately from the pricing check
Buildability by layer
Screens, forms, and focused interactions
The repeatable job the product performs
Availability and legality of required data
Uptime, queues, support, and maintenance
Security, compliance, and user confidence
The achievable core
- A local Electron overlay bound to a hotkey that screenshots your active display, transcribes recent audio, sends both to a multimodal LLM, and streams a short answer into a translucent window nobody else is supposed to see.
- Answer from retrieval over your own notes and files, with sources visible.
- A responsive interface with real empty, loading, success, and error states.
The parts a prompt cannot buy
- Screen-share invisibility that has actually been tested against Zoom, Meet, Teams and OS screenshot tools
- Sub-second latency and the prompt tuning that makes answers short enough to read while someone is talking
- System audio capture that just works without you installing and routing a virtual audio device
- Mobile and browser-extension coverage, plus a hosted account so your setup follows you to another machine
- The last 20 percent is sync, migration fidelity, speed, and edge cases.
- Trust, audits, and counterparties matter more than feature parity.
Build, switch, or keep paying
Narrower, with trade-offs
A local Electron overlay bound to a hotkey that screenshots your active display, transcribes recent audio, sends both to a multimodal LLM, and streams a short answer into a translucent window nobody else is supposed to see.
Use the build brief ↓No checked option yet
Compare the prior art below or build only the workflow you need.
$19.99/mo
Because the hard part is not the LLM call, it is the twenty small platform details that make an overlay silent, invisible and fast under pressure, and because people who reach for this tool are, by definition, not in the mood to debug a virtual audio driver ten minutes before an interview. Paying converts a fragile personal hack into something that mostly behaves on a laptop you did not configure yourself.
Visit Cluely ↗Why people still pay
Because the hard part is not the LLM call, it is the twenty small platform details that make an overlay silent, invisible and fast under pressure, and because people who reach for this tool are, by definition, not in the mood to debug a virtual audio driver ten minutes before an interview. Paying converts a fragile personal hack into something that mostly behaves on a laptop you did not configure yourself.
The last 20 percent is sync, migration fidelity, speed, and edge cases.
Trust, audits, and counterparties matter more than feature parity.
The brief
Context, requirements, acceptance criteria, non-goals, and the full production standard — as Markdown, ready for any coding agent.
Build brief — a focused alternative to Cluely
Context
Cluely — An always-on-top desktop overlay that watches your screen and listens to your calls, then feeds you AI answers in real time. It currently costs $19.99/mo.
The core loop is genuinely thin: capture a screenshot, transcribe the mic, feed both to a multimodal model, stream the answer into a transparent always-on-top window on a hotkey. An agent gets that working in a session, and it will feel uncomfortably close to the demo. The gaps are the unglamorous parts: capturing the other side's audio takes virtual audio devices, latency has to beat the conversation, and the headline trick of staying invisible in screen shares depends on fragile platform window flags that vary by OS and conferencing app. Your DIY version will work fine when you are alone with your own screen, and may quietly betray you the one time it matters.
This brief describes a focused, single-operator replacement for the part of Cluely that is genuinely reproducible. It is deliberately narrower than the product it replaces, and it says so in writing. Build the useful core; do not pretend to have rebuilt the rest.
What you are building
A local Electron overlay bound to a hotkey that screenshots your active display, transcribes recent audio, sends both to a multimodal LLM, and streams a short answer into a translucent window nobody else is supposed to see.
Answer from retrieval over your own notes and files, with sources visible.
A responsive interface with real empty, loading, success, and error states.
Requirements
Functional
MacOS or Windows desktop, plus screen recording and microphone permissions.
Node 20 and willingness to sign or self-trust an unsigned Electron build.
Data and integrations
An OpenAI API key (or any multimodal chat endpoint) in .env.
A virtual audio device for other-party audio: BlackHole on macOS, VB-Cable or WASAPI loopback on Windows.
Each of these needs a real account, credential, or quota. Set them up before writing feature code.
Non-functional
Accessibility: semantic markup, labelled controls, visible focus, and reduced-motion support.
Security: server-side secrets, validated input, and no credentials in the client bundle.
Reliability: retries with backoff on external calls, and a clear failure state when a provider is down.
Portability: the operator can export their data and leave without losing it.
Implementation brief
Build a local desktop AI overlay assistant called "Peek". Stack: Electron 30 with TypeScript, Vite for the renderer, no framework beyond plain TS and CSS. Target macOS first, keep Windows paths behind a small platform module. Everything runs locally except LLM calls.
Behavior:
1. On launch, create a frameless, transparent, always-on-top BrowserWindow, 420x520, positioned top-right of the primary display. No dock icon, no menu bar item beyond a tray icon with Quit.
2. Set the window to ignore mouse events by default; hold a modifier hotkey to make it interactive.
3. Global hotkeys: Cmd+Shift+Space asks a question from the current context, Cmd+Shift+H toggles visibility, Cmd+Shift+Enter follows up in the same thread.
4. On ask: capture a PNG of the active display via Electron desktopCapturer at max 1600px wide, grab the last 30 seconds of rolling audio transcript, and send both to the model.
5. Audio: capture from a configurable input device using the renderer's MediaRecorder in 5 second chunks, transcribe each chunk with OpenAI whisper-1, keep a rolling 60 second transcript buffer in memory only. Never write audio to disk.
6. LLM: call the OpenAI chat completions endpoint with a multimodal message (screenshot plus transcript plus user intent), stream tokens into the overlay. System prompt: answer in under 60 words, lead with the answer, bullets only when listing, no preamble.
7. Attempt screen-share exclusion: call setContentProtection(true) on the window, and on Windows use the SetWindowDisplayAffinity equivalent flag. Print a startup warning in the console that exclusion is best effort and must be verified manually.
8. Settings: a small JSON config at userData/config.json for model name, input device id, hotkeys, and transcript window length. No settings UI beyond a tray menu item that opens the file.
In scope: the overlay, hotkeys, screenshot capture, rolling transcription, streaming answers, tray quit, README with macOS permission steps and BlackHole routing instructions for capturing other-party audio.
Out of scope: accounts, login, cloud sync, telemetry, analytics, auto-update, code signing, mobile, browser extension, meeting integrations, transcript persistence, any database.
Secrets: OPENAI_API_KEY in .env, loaded in the main process only. Never expose the key to the renderer; proxy all API calls through IPC. Commit a .env.example.
Deliver: working npm scripts dev and build, a README with a 60 second setup, and a HONESTY.md file listing what will break, starting with screen-share invisibility and audio device routing.
Delivery standard
Inspect the repository first, then write a short implementation plan before writing code.
Deliver the smallest complete end-to-end workflow first; every primary control must work against persisted data.
Use real validation and storage; never substitute fake dashboards, decorative controls, hard-coded success states, or mock integrations.
Include responsive layouts plus genuine empty, loading, success, validation, and failure states.
Keep secrets server-side in environment variables, provide .env.example, and never commit credentials or user data.
Add structured logs around every external call and return actionable errors without leaking sensitive details.
Write unit tests for the core logic and one automated test of the main user journey.
Finish with a README covering setup, architecture, data location, backups, tests, deployment, and known limitations.
Acceptance criteria
A clean install starts the app using only the README and .env.example.
The primary journey works from first visit through saved result, reload, edit, export, and deletion where applicable.
Invalid input, missing configuration, provider failure, and an empty database each have a usable state.
The interface works at 390px and 1440px, is keyboard navigable, and shows visible focus on every control.
Tests, type checking, linting, and a production build all pass with no ignored failures.
No part of the interface implies a live integration, security guarantee, or scale capability that was not actually built and verified.
Non-goals
Do not build these, and do not claim to have replaced them:
Screen-share invisibility that has actually been tested against Zoom, Meet, Teams and OS screenshot tools.
Sub-second latency and the prompt tuning that makes answers short enough to read while someone is talking.
System audio capture that just works without you installing and routing a virtual audio device.
Mobile and browser-extension coverage, plus a hosted account so your setup follows you to another machine.
Somebody else's legal and PR exposure for a product whose pitch is 'cheat on everything'.
What you still own after launch
Secure credentials, rotate secrets, and handle provider rate limits.
Run migrations, backups, restores, and dependency updates.
Test the critical journey after every model, API, or hosting change.
Monitor failures and fix the edge cases a first prompt will miss.
Risk
Manageable. A personal version is realistic if you test the critical journey and keep reliable backups.
Editorial confidence in this assessment: medium. No reviewed project implementation is linked yet.
Generated by Can It Be Vibe Coded? · Full report: https://www.canitbevibecoded.com/cluely
You still own the product
- Secure credentials, rotate secrets, and handle provider rate limits.
- Run migrations, backups, restores, and dependency updates.
- Test the critical journey after every model, API, or hosting change.
- Monitor failures and fix the edge cases a first prompt will miss.
Projects built from this idea
No reviewed implementation has been linked for Cluely yet. A submission is evidence for review, not automatic proof that the whole product was replaced.
Built a version of Cluely?Submit the project as evidence for this report.
Before you start
Can Cluely be vibe coded?
Partly, if you narrow it. The core loop is genuinely thin: capture a screenshot, transcribe the mic, feed both to a multimodal model, stream the answer into a transparent always-on-top window on a hotkey. An agent gets that working in a session, and it will feel uncomfortably close to the demo. The gaps are the unglamorous parts: capturing the other side's audio takes virtual audio devices, latency has to beat the conversation, and the headline trick of staying invisible in screen shares depends on fragile platform window flags that vary by OS and conferencing app. Your DIY version will work fine when you are alone with your own screen, and may quietly betray you the one time it matters.
What can an AI coding agent reproduce from Cluely?
A local Electron overlay bound to a hotkey that screenshots your active display, transcribes recent audio, sends both to a multimodal LLM, and streams a short answer into a translucent window nobody else is supposed to see. Answer from retrieval over your own notes and files, with sources visible. A responsive interface with real empty, loading, success, and error states.
What will a DIY Cluely replacement still be missing?
Screen-share invisibility that has actually been tested against Zoom, Meet, Teams and OS screenshot tools; Sub-second latency and the prompt tuning that makes answers short enough to read while someone is talking; System audio capture that just works without you installing and routing a virtual audio device; Mobile and browser-extension coverage, plus a hosted account so your setup follows you to another machine; The last 20 percent is sync, migration fidelity, speed, and edge cases.; Trust, audits, and counterparties matter more than feature parity.
What do I still own after building a Cluely alternative?
Secure credentials, rotate secrets, and handle provider rate limits. Run migrations, backups, restores, and dependency updates. Test the critical journey after every model, API, or hosting change. Monitor failures and fix the edge cases a first prompt will miss.