# Build brief — a focused alternative to diclip

> **Verdict:** Partly, if you narrow it · **Buildability:** 58/100 · **Category:** Audio Video
> **Source:** https://www.canitbevibecoded.com/diclip
> Independent editorial assessment from Can It Be Vibe Coded? Not affiliated with, endorsed by, or derived from diclip. Verify current pricing and capabilities before acting.

## Context

**diclip** — Turns long recordings into ranked, captioned vertical clips worth posting. It currently costs $6/mo.

The core loop is genuinely one-shottable: transcribe, rank the strongest moments, cut vertical clips with burned captions. The gaps are execution and infra. Face-aware reframing with seat tracking, a full in-browser timeline editor, and a container-scale render pipeline are not a one-session build. A competent agent can produce a usable personal clipper quickly, but matching diclip's reframe quality and editing depth is a substantial project.

This brief describes a focused, single-operator replacement for the part of diclip that is genuinely reproducible. It is deliberately narrower than the product it replaces, and it says so in writing. Build the useful core; do not pretend to have rebuilt the rest.

## What you are building

Transcribe a long video, rank the strongest moments with an LLM, and render vertical clips with burned captions using FFmpeg.

- Transcribe supplied recordings, cut them on a timeline, and export finished files.
- A responsive interface with real empty, loading, success, and error states.

## Requirements

### Functional

- FFmpeg.
- Faster-whisper.
- Desktop with enough storage for source media and renders.

### Data and integrations

- LLM API key for moment ranking (optional, falls back to heuristics).

Each of these needs a real account, credential, or quota. Set them up before writing feature code.

### Non-functional

- Accessibility: semantic markup, labelled controls, visible focus, and reduced-motion support.
- Security: server-side secrets, validated input, and no credentials in the client bundle.
- Reliability: retries with backoff on external calls, and a clear failure state when a provider is down.
- Portability: the operator can export their data and leave without losing it.

## Implementation brief

Build me a personal substitute for diclip, not a platform clone.

- Use Python 3.12 + FastAPI for a localhost web app, a plain JavaScript
  frontend, faster-whisper for word-level transcription, and FFmpeg for
  rendering.
- I import an MP4, MOV, WebM, or MP3 file and get a transcript synced to the
  video with speaker timing.
- A ranking step selects the strongest 30-90 second moments: call an LLM API
  key from .env with the transcript, or fall back to a local heuristic
  (keyword density, question marks, pauses) when no key is set.
- Show each candidate clip with its hook line, a confidence score, and a one
  sentence reason, marked ready or pending.
- Let me accept or reject each clip, edit in/out points, and choose 9:16, 1:1,
  or 16:9 output.
- Render clips with FFmpeg using center-crop or blur-pad, burning captions
  styled by speaker. Never overwrite the source. Show progress and a useful
  failure message.
- Store projects as JSON under ~/diclipDIY/projects and renders under
  ~/diclipDIY/exports, with a recent-projects page and a delete action.
- Bind to localhost only. No accounts, telemetry, or network calls after the
  model and transcript are local, except the optional LLM ranking call.
- Deliberately exclude face-aware reframing with seat tracking, the full
  in-browser timeline editor, cloud container rendering, yt-dlp downloads, and
  credit billing.
- Add unit tests for the ranking fallback and caption grouping, plus one smoke
  test that imports a short fixture and produces a playable MP4.
- Include a README with setup, data paths, and an honest note that CPU
  transcription and rendering are slow on long videos.

## Delivery standard

- Inspect the repository first, then write a short implementation plan before writing code.
- Deliver the smallest complete end-to-end workflow first; every primary control must work against persisted data.
- Use real validation and storage; never substitute fake dashboards, decorative controls, hard-coded success states, or mock integrations.
- Include responsive layouts plus genuine empty, loading, success, validation, and failure states.
- Keep secrets server-side in environment variables, provide .env.example, and never commit credentials or user data.
- Add structured logs around every external call and return actionable errors without leaking sensitive details.
- Write unit tests for the core logic and one automated test of the main user journey.
- Finish with a README covering setup, architecture, data location, backups, tests, deployment, and known limitations.

## Acceptance criteria

- [ ] A clean install starts the app using only the README and .env.example.
- [ ] The primary journey works from first visit through saved result, reload, edit, export, and deletion where applicable.
- [ ] Invalid input, missing configuration, provider failure, and an empty database each have a usable state.
- [ ] The interface works at 390px and 1440px, is keyboard navigable, and shows visible focus on every control.
- [ ] Tests, type checking, linting, and a production build all pass with no ignored failures.
- [ ] No part of the interface implies a live integration, security guarantee, or scale capability that was not actually built and verified.

## Non-goals

Do not build these, and do not claim to have replaced them:

- Face-aware vertical reframing with seat tracking.
- Full in-browser timeline editor with scenes, tracks, and effects.
- Container-scale render pipeline and yt-dlp import.
- Credit-based capacity model and priority processing.
- Ready-to-post hooks and confidence scoring polished for clips.

## What you still own after launch

- Secure credentials, rotate secrets, and handle provider rate limits.
- Run migrations, backups, restores, and dependency updates.
- Test the critical journey after every model, API, or hosting change.
- Monitor failures and fix the edge cases a first prompt will miss.

## Risk

**Operational risk.** The code is achievable; dependable data, integrations, and ongoing operations are the real cost.

Editorial confidence in this assessment: medium. No reviewed project implementation is linked yet.

---

Generated by [Can It Be Vibe Coded?](https://www.canitbevibecoded.com) · Full report: https://www.canitbevibecoded.com/diclip
