# Build brief — a focused alternative to Midjourney

> **Verdict:** Not faithfully · **Buildability:** 21/100 · **Category:** AI Image
> **Source:** https://www.canitbevibecoded.com/midjourney
> Independent editorial assessment from Can It Be Vibe Coded? Not affiliated with, endorsed by, or derived from Midjourney. Verify current pricing and capabilities before acting.

## Context

**Midjourney** — Text-to-image generation service known for high-quality aesthetic output. It currently costs $10/mo.

You can build a wrapper around open models, but you cannot solo-recreate Midjourney's proprietary model quality, style tuning, infrastructure, and community distribution.

This brief describes a focused, single-operator replacement for the part of Midjourney that is genuinely reproducible. It is deliberately narrower than the product it replaces, and it says so in writing. Build the useful core; do not pretend to have rebuilt the rest.

## What you are building

Use Stable Diffusion/Flux-style local or hosted models with a prompt UI and gallery.

- Build a focused single-user workflow with real persistence, search, and export.
- A responsive interface with real empty, loading, success, and error states.

## Requirements

### Functional

- Model weights/license.
- Storage/gallery.
- Prompt UI.

### Data and integrations

- GPU or hosted image API.

Each of these needs a real account, credential, or quota. Set them up before writing feature code.

### Non-functional

- Accessibility: semantic markup, labelled controls, visible focus, and reduced-motion support.
- Security: server-side secrets, validated input, and no credentials in the client bundle.
- Reliability: retries with backoff on external calls, and a clear failure state when a provider is down.
- Portability: the operator can export their data and leave without losing it.

## Implementation brief

Build me a local image-generation studio wrapping open models, in place of
Midjourney. Requirements:

- A local web app (Node + Express, plain HTML/JS) on localhost:4890: prompt box,
  a generate-4 button, and a gallery grid of results.
- Backend targets whichever engine I have: a local ComfyUI instance running Flux
  or SDXL if there is a GPU, otherwise a hosted API like Replicate or fal.ai
  (key in .env). Same UI either way.
- Save every image to ~/ImageGen/YYYY-MM/ with a sidecar JSON: prompt, seed,
  model, steps, so any result is reproducible later.
- Gallery search over past prompts via SQLite FTS5, favorites, and one-click
  re-run with the same seed or a new variation.
- Style presets in a styles.json of prompt prefixes and suffixes I can toggle,
  a poor man's style reference.
- No accounts, no telemetry, everything stays local except the hosted-API calls.
- Out of scope: training or fine-tuning models, and matching Midjourney's look.
  Build for reproducibility and control, not a taste layer.
- README: both engine setups, the rough cost per image on the hosted route, and
  a plain note that this is not Midjourney, the proprietary model and its years
  of aesthetic tuning cannot be rebuilt, this trades that for privacy and control.

## Delivery standard

- Inspect the repository first, then write a short implementation plan before writing code.
- Deliver the smallest complete end-to-end workflow first; every primary control must work against persisted data.
- Use real validation and storage; never substitute fake dashboards, decorative controls, hard-coded success states, or mock integrations.
- Include responsive layouts plus genuine empty, loading, success, validation, and failure states.
- Keep secrets server-side in environment variables, provide .env.example, and never commit credentials or user data.
- Add structured logs around every external call and return actionable errors without leaking sensitive details.
- Write unit tests for the core logic and one automated test of the main user journey.
- Finish with a README covering setup, architecture, data location, backups, tests, deployment, and known limitations.

## Acceptance criteria

- [ ] A clean install starts the app using only the README and .env.example.
- [ ] The primary journey works from first visit through saved result, reload, edit, export, and deletion where applicable.
- [ ] Invalid input, missing configuration, provider failure, and an empty database each have a usable state.
- [ ] The interface works at 390px and 1440px, is keyboard navigable, and shows visible focus on every control.
- [ ] Tests, type checking, linting, and a production build all pass with no ignored failures.
- [ ] No part of the interface implies a live integration, security guarantee, or scale capability that was not actually built and verified.

## Non-goals

Do not build these, and do not claim to have replaced them:

- Frontier/proprietary model quality.
- Style consistency.
- Moderation.
- Compute scale.
- Model quality and inference operations are part of the product.
- The value comes from the people already using it.

## What you still own after launch

- Secure credentials, rotate secrets, and handle provider rate limits.
- Run migrations, backups, restores, and dependency updates.
- Test the critical journey after every model, API, or hosting change.
- Monitor failures and fix the edge cases a first prompt will miss.

## Risk

**Manageable.** A personal version is realistic if you test the critical journey and keep reliable backups.

Editorial confidence in this assessment: high. No independent one-shot implementation is linked yet.

## Prior art

Working open-source software you can read, fork, or borrow from before starting:

- [ComfyUI](https://github.com/comfyanonymous/ComfyUI) — Open-source node-based UI for running image-generation workflows locally or on a GPU serve

---

Generated by [Can It Be Vibe Coded?](https://www.canitbevibecoded.com) · Full report: https://www.canitbevibecoded.com/midjourney
