creator-studio

byPriyanshu Mogra

Build me a personal, self-hosted AI video creation platform called “Creator Studio”. This is NOT a kids-only website. It should be a general-purpose video creation tool that I can use for YouTube, Shorts, Reels, TikTok, educational videos, faceless videos, storytelling, documentaries, gaming content, nursery rhymes, and any other type of video. The most important requirement is: NO PAID API DEPENDENCIES by default. Design the platform around free and open-source models and local processing. API integrations can be optional, but the application must work without paid APIs. MAIN WORKFLOW: SCRIPT → VIDEO I should be able to paste any script and automatically turn it into a complete video. Input: - Paste script - Upload TXT/DOCX/PDF - Or enter a simple topic and let AI create the script Then automatically: 1. Analyze the script 2. Divide it into scenes 3. Estimate duration 4. Generate scene descriptions 5. Generate visual prompts 6. Generate voiceover 7. Generate or select visuals 8. Add background music 9. Add sound effects 10. Generate subtitles 11. Automatically synchronize everything 12. Assemble the final video 13. Export MP4 VIDEO TYPES: Provide presets: - YouTube video - YouTube Shorts - Instagram Reels - TikTok - Educational video - Story video - Faceless video - Kids animation - Documentary - Motivational video - Podcast clips - Custom Allow custom: - Aspect ratio - Resolution - FPS - Duration - Style SUPPORTED FORMATS: - 16:9 - 9:16 - 1:1 - 4:5 AI SCRIPT GENERATOR: Create a powerful script generator where I enter: Topic: Audience: Video length: Language: Tone: Style: Generate: - Hook - Full script - Narration - Scene breakdown - Dialogue - Visual instructions - CTA Allow me to edit the generated script before creating the video. AI SCENE GENERATOR: Automatically divide scripts into scenes. Each scene should contain: - Scene number - Duration - Narration - Visual description - Image prompt - Video prompt - Camera movement - Transition - Sound effect - Background music - Subtitle text Allow: - Regenerate scene - Edit scene - Delete scene - Duplicate scene - Reorder scene - Regenerate visual - Change duration TEXT TO SPEECH: Integrate free/local TTS options such as Piper TTS, Coqui TTS or other suitable open-source TTS engines. Allow: - Male voices - Female voices - Different languages - Different accents where available - Speaking speed - Pitch - Voice preview Keep the system modular so additional TTS providers can be added later. AI IMAGE GENERATION: Support local/open-source image generation where possible. Allow: - Generate image from prompt - Generate multiple variations - Select best image - Regenerate - Upload my own image - Image-to-image - Maintain character/style consistency AI VIDEO GENERATION: Provide a video-generation module designed to support locally available/open-source models when the user's computer has sufficient GPU resources. The system should support pluggable video generation backends rather than being locked to one provider. Allow: - Text-to-video - Image-to-video - Animate image - Camera movement - Scene duration - Motion strength If local video generation is unavailable because of hardware limitations, gracefully fall back to image-based video creation using: - Generated images - Ken Burns effect - Zoom - Pan - Camera movement - Transitions VIDEO EDITOR: Create a proper timeline-based video editor. Timeline tracks: 1. Video 2. Images 3. Voiceover 4. Music 5. Sound effects 6. Subtitles 7. Text 8. Overlays Features: - Cut - Split - Trim - Crop - Resize - Rotate - Speed control - Volume control - Fade in/out - Transitions - Text overlays - Subtitles - Stickers - Images - Audio - Video layers Use FFmpeg for video processing and rendering. CLIPPING TOOL: This is VERY IMPORTANT. Create an automatic clipping system where I can upload a long video and generate short clips from it. Input: - MP4/MKV/MOV/WebM - YouTube video file - Podcast - Lecture - Interview - Long-form video Automatically: 1. Transcribe the video 2. Detect important/high-engagement moments 3. Find potential clips 4. Suggest clip start/end times 5. Generate titles 6. Generate captions 7. Convert clips to 9:16 8. Add animated subtitles 9. Reframe the speaker 10. Export multiple Shorts Show suggested clips like: CLIP 01 00:14:32 → 00:15:18 “Most interesting moment” Score: 92% CLIP 02 00:28:10 → 00:29:03 “Strong hook” Score: 88% Allow me to preview, edit and export each clip. AUTO CAPTIONS: Use free/open-source speech recognition such as Whisper or faster-whisper. Features: - Automatic transcription - Word-level timestamps - Subtitle generation - SRT/VTT export - Animated captions - Highlight important words - Multiple subtitle styles Allow caption presets: - Minimal - Bold - Shorts - Karaoke - Highlight - Professional AUDIO: Provide: - Background music - Upload music - Sound effects - Audio trimming - Volume adjustment - Noise reduction - Voice enhancement - Audio ducking Use royalty-free/open-source music by default and clearly indicate the source/license when applicable. THUMBNAIL GENERATOR: Generate YouTube thumbnails from: - Video - Screenshot - Prompt Allow: - Text - Images - Background - AI-generated elements - Multiple variations - 1280×720 export AI CONTENT PACKAGE: After creating a video automatically generate: YouTube title YouTube description SEO keywords Hashtags Tags Thumbnail text Short description Social media captions PROJECT MANAGEMENT: Create: - New Project - Save Project - Duplicate Project - Rename - Delete - Export - Import project Automatically save progress. MEDIA LIBRARY: Store: - Images - Videos - Audio - Voiceovers - Music - SFX - Thumbnails - Generated clips Allow search, filtering and folders. FREE/LOCAL-FIRST ARCHITECTURE: Prioritize free and open-source tools. Use where appropriate: - FFmpeg for video processing - Whisper/faster-whisper for transcription - Piper/other open-source TTS for voice - Open-source LLMs through Ollama - Stable Diffusion/Flux-compatible local image generation where hardware allows - Open-source video-generation models where hardware allows - OpenCV for video/image processing - Python for AI/video processing - FastAPI or Flask for backend - React/Next.js for frontend The application should detect which local AI tools are installed and show their status. Create a “System Status” page showing: ✓ FFmpeg ✓ Whisper ✓ Ollama ✓ TTS ✓ Image Generation ✓ Video Generation ✓ GPU ✓ Storage If a component isn't installed, show: - What it does - Whether it is optional - Installation instructions IMPORTANT: Do NOT pretend that unlimited AI video generation is free if the required model needs expensive GPU hardware. The software itself should be free/open-source and local-first wherever possible. For users without a powerful GPU, provide a lightweight mode using: - Local LLM - Local TTS - Whisper - User-provided images/videos - FFmpeg - Motion effects - Transitions - Captions This should still allow complete videos to be created without paid APIs. DASHBOARD: Create a beautiful professional dashboard with: “Create Video” “Generate Script” “Script → Video” “Video → Clips” “AI Voice” “Generate Images” “Generate Thumbnail” “Video Editor” Show: Recent Projects Storage Processing Jobs System Status DESIGN: Use a modern dark professional interface similar to professional creative software. Dark background. Clean cards. Subtle gradients. Smooth animations. Modern typography. Clear icons. Responsive design. Do NOT make it look like a children's website. Make it feel like a personal combination of: AI script generator + CapCut + Canva + AI video generator + Whisper transcription + FFmpeg video processor but completely focused on my own personal workflow. TECHNICAL REQUIREMENTS: Frontend: React or Next.js TypeScript Tailwind CSS Backend: Python FastAPI Video processing: FFmpeg AI: Modular provider architecture Database: SQLite for local installation, with an option to switch to PostgreSQL later. Storage: Local filesystem by default. Authentication: Optional for local personal use. All API keys must be stored in environment variables and NEVER exposed in frontend code. Create clean modular architecture so I can replace any AI model later. The application should run locally with: npm install npm run dev and the backend with: pip install -r requirements.txt python server.py Also provide a README containing complete installation instructions for Windows. The final result should be a REAL WORKING APPLICATION, not merely a UI mockup. Prioritize the following features in the first working version: 1. Script → Scene Breakdown 2. Script → Voiceover 3. Script → Images 4. Images + Voiceover → Video 5. Automatic subtitles 6. FFmpeg rendering 7. Video → Automatic Clips 8. Shorts/Reels 9:16 export 9. Thumbnail generation 10. YouTube metadata generation Build the project so additional AI video models can be plugged in later without rewriting the entire application.

No preview

Comments (0)

No comments yet. Be the first!

System Requirements

Page 1 of 40

System Requirements Document for creator-studio

1. Introduction

Creator Studio is a personal, self-hosted AI video creation platform. It is a general-purpose video creation tool for YouTube, Shorts, Reels, TikTok, educational videos, faceless videos, storytelling, documentaries, gaming content, nursery rhymes, and any other type of video. It is explicitly not a kids-only website.

The product's defining constraint is that it has no paid API dependencies by default. The platform is designed around free and open-source models and local processing. API integrations may be added optionally, but the application must work without paid APIs. The software itself is free/open-source and local-first wherever possible, and it does not pretend that unlimited AI video generation is free when the required model needs expensive GPU hardware.

The audience is a single power-user creator (Priyanshu Mogra, IN) who owns and drives the machine: a workshop of FFmpeg, Whisper, Piper and Ollama that the user runs themselves. The product should feel like a personal combination of an AI script generator, CapCut, Canva, an AI video generator, Whisper transcription, and an FFmpeg video processor — completely focused on the user's own personal workflow.

The final result must be a real working application, not merely a UI mockup.

Page 2 of 40

2. System Overview

Creator Studio runs locally as two processes: a React/Next.js + TypeScript + Tailwind CSS frontend started with npm install and npm run dev, and a Python FastAPI backend started with pip install -r requirements.txt and python server.py. Video processing is performed by FFmpeg. AI capabilities are provided through a modular provider architecture so any AI model can be replaced later, and additional AI video models can be plugged in later without rewriting the entire application.

The core current capability is the SCRIPT → VIDEO pipeline: accept a pasted script, an uploaded TXT/DOCX/PDF, or a simple topic that AI turns into a script; then automatically analyze the script, divide it into scenes, estimate duration, generate scene descriptions, generate visual prompts, generate voiceover, generate or select visuals, add background music, add sound effects, generate subtitles, synchronize everything, assemble the final video, and export MP4.

Supporting current capabilities include video-type presets and custom format controls, an AI script generator, an AI scene generator, local text-to-speech, local image generation, pluggable local video generation with a graceful image-based fallback, a timeline-based video editor, an automatic clipping tool for long-form video, auto captions, audio management, a thumbnail generator, an AI content package, project management, a media library, a System Status page, and a professional dashboard.

Storage is the local filesystem by default. SQLite is used for local installation, with an option to switch to PostgreSQL later. Authentication is optional for local personal use. All API keys must be stored in environment variables and never exposed in frontend code.

Narrow exclusions and boundaries:

  • The platform must not require paid APIs to function; optional API integrations are permitted but never mandatory.
  • The interface must not look like a children's website.
  • The product must not claim unlimited free AI video generation when the underlying model requires expensive GPU hardware.
  • Local video generation is conditional on sufficient GPU resources; when unavailable, the system falls back to image-based video creation rather than failing.
Page 3 of 40

2a. Product Interpretation and Delivery Boundary

Creator Studio is delivered as a local, self-hosted application owned and operated by a single creator on their own machine. There is no hosted multi-tenant service, no subscription, and no mandatory cloud dependency. The frontend and backend run locally; media is written to the local filesystem; project state is persisted in a local SQLite database (with a documented option to switch to PostgreSQL later).

Access ownership is deliberately simple. Because authentication is optional for local personal use, the application does not gate its working surfaces behind an account system. The public entry surface (Landing) is anonymously reachable and explains what the product is; the working surfaces (Dashboard, production workspaces, editor, library, status) are reachable in the same local session. No differentiated permissions, roles, or role-based visibility are established by the source, and none are introduced here.

Delivery ownership is split between first-party application surfaces and locally installed external tools. The application owns the user's interaction: choosing a video type, entering or uploading a script, reviewing and editing scenes, configuring voices, generating and selecting images, driving the timeline, reviewing suggested clips, choosing caption presets, exporting thumbnails, and reading the content package. The heavy lifting is performed by locally installed, user-owned tools — FFmpeg, Whisper/faster-whisper, Piper/Coqui TTS, Ollama, Stable Diffusion/Flux-compatible image generation, and open-source video-generation models — which the application detects, reports on, and invokes. These tools are not part of the application's own UI; the application surfaces their availability, optionality, and installation instructions.

Current vs. future boundary: everything described in this document is current. Optional paid API integrations are explicitly permitted but are not current requirements and are not designed as mandatory paths. The modular provider architecture is a current requirement precisely so that future models and providers can be added without rewriting the application.

Page 4 of 40

2b. Source Content Inventory

Not applicable — no reference directive in this request declares a content_source.

2c. Page Content and Component Coverage

Landing

  • Information/state: Anonymous first impression of Creator Studio as a personal, general-purpose, local-first video creation platform. States the no-paid-API-by-default promise, the local-first architecture, and the range of video types supported (YouTube, Shorts, Reels, TikTok, educational, faceless, storytelling, documentary, gaming, nursery rhymes, and any other type). States honestly that local AI video generation depends on GPU hardware and that a lightweight mode exists for machines without a powerful GPU.
  • Primary actions: Enter the studio (go to Dashboard). Read the local-first / no-paid-API explanation.
  • Supporting actions: Jump to System Status to see which local tools are installed.
  • Domain entities: Product description, capability summary, supported video types, local tool list (FFmpeg, Whisper, Ollama, TTS, Image Generation, Video Generation, GPU, Storage).
  • Component responsibilities: Full-bleed hero composition with oversized stacked headline and primary CTA; framed vertical video still with quoted aspect-ratio label; faint 12-column cutting-mat grid; capability/pipeline summary; honest hardware note.
  • States: Loading (static content, no blocking fetch); empty (n/a); success (rendered); error (if the local backend is unreachable, show a clear "backend not running" notice with the python server.py instruction); recovery (retry connection).

Dashboard

  • Information/state: Central hub. Shows Recent Projects, Storage, Processing Jobs, and System Status. Presents the creation routes: "Create Video", "Generate Script", "Script → Video", "Video → Clips", "AI Voice", "Generate Images", "Generate Thumbnail", "Video Editor".
  • Primary actions: Launch any of the eight creation routes.
  • Supporting actions: Open a recent project; open System Status; inspect a processing job.
  • Domain entities: Project, Processing Job, Storage usage, System component status.
  • Component responsibilities: Creation-route module row; recent-projects filmstrip of framed thumbnails; storage readout; processing-jobs list with live progress; right-edge system ledger with per-component status chips.
  • States: Loading (skeleton plates for projects/jobs/status); empty (no projects yet — prompt to create the first project; no jobs — quiet empty ledger); success (populated modules); error (a failed module shows an inline error and does not block the others); recovery (per-module retry).
Page 5 of 40

Create Video

  • Information/state: Entry workspace for selecting a video type, source, and production path. Shows the preset list (YouTube video, YouTube Shorts, Instagram Reels, TikTok, Educational video, Story video, Faceless video, Kids animation, Documentary, Motivational video, Podcast clips, Custom) and the custom controls (aspect ratio, resolution, FPS, duration, style). Shows supported formats 16:9, 9:16, 1:1, 4:5.
  • Primary actions: Choose a preset or Custom; set custom aspect ratio, resolution, FPS, duration, and style; choose the script source (paste script / upload TXT/DOCX/PDF / enter a topic); continue into production.
  • Supporting actions: Switch to lightweight mode when GPU is insufficient.
  • Domain entities: Video Type Preset, Format (aspect ratio), Resolution, FPS, Duration, Style, Script Source.
  • Component responsibilities: Preset plate grid; custom parameter panel; script-source selector; format chips stamped with quoted labels ("16:9", "9:16", "1:1", "4:5"); hardware-aware mode indicator.
  • States: Loading (preset list resolving); empty (no preset selected — Custom is the implicit default); success (selection confirmed, continue enabled); error (invalid custom combination, e.g. unsupported resolution for the chosen aspect ratio — inline validation); recovery (reset to a valid preset).

Generate Script

  • Information/state: Script generation and editing workspace. Inputs: Topic, Audience, Video length, Language, Tone, Style. Outputs: Hook, Full script, Narration, Scene breakdown, Dialogue, Visual instructions, CTA.
  • Primary actions: Generate the script from the inputs; edit the generated script before creating the video; send the edited script onward to production.
  • Supporting actions: Regenerate; copy a section; inspect the scene breakdown preview.
  • Domain entities: Script, Hook, Narration, Dialogue, Visual Instructions, CTA, Scene Breakdown, generation inputs.
  • Component responsibilities: Input form; generated-output sections as separate labeled blocks; editable script surface; provenance of the local LLM provider used.
  • States: Loading (generation in progress with a running indicator); empty (no inputs yet — form only); success (all output sections populated and editable); error (local LLM unavailable — explain what Ollama does, that it is required for this step, and how to install it); recovery (retry generation, or paste a script manually instead).

Script → Video

  • Information/state: The complete script-based production pipeline. Displays the 13 stages: analyze the script, divide it into scenes, estimate duration, generate scene descriptions, generate visual prompts, generate voiceover, generate or select visuals, add background music, add sound effects, generate subtitles, synchronize everything, assemble the final video, export MP4. Each stage shows DONE / RUNNING / QUEUED.
  • Primary actions: Start the pipeline; watch stage progress; intervene at a stage; export the final MP4.
  • Supporting actions: Open the produced scenes, voiceover, images, captions, or editor from the relevant stage; switch to lightweight mode.
  • Domain entities: Pipeline Run, Stage, Scene, Voiceover, Visual, Background Music, Sound Effect, Subtitle, Final Video, MP4 export.
  • Component responsibilities: Numbered stage ledger with 1px rules and a progress bar on the running row; per-stage status stamps; stage-level links into the owning workspace; export control.
  • States: Loading (pipeline initializing); empty (no script supplied — route back to source selection); success (all stages DONE, MP4 available); error (a stage fails — the failing stage is named, the reason is shown, and the pipeline halts at that stage rather than silently continuing); recovery (retry the failed stage, or fix the input in the owning workspace and resume).
Page 6 of 40

Scenes

  • Information/state: Scene breakdown workspace. Each scene contains scene number, duration, narration, visual description, image prompt, video prompt, camera movement, transition, sound effect, background music, and subtitle text.
  • Primary actions: Regenerate scene; edit scene; delete scene; duplicate scene; reorder scene; regenerate visual; change duration.
  • Supporting actions: Inspect a scene's generated visual; jump to Generate Images or Video Generator for a specific scene.
  • Domain entities: Scene (with all eleven fields), Scene Order, Visual.
  • Component responsibilities: Scene list with numbered tags; per-scene field inspector; reorder affordance with a 1px insertion rule; per-scene action row; duration control.
  • States: Loading (scenes being derived from the script); empty (no scenes yet — prompt to run the breakdown); success (ordered scene list with all fields populated); error (a scene fails to regenerate — the previous scene content is preserved and the failure is reported); recovery (retry the scene, or edit it manually).

AI Voice

  • Information/state: Local TTS configuration and voiceover production. Shows available voices with male voices, female voices, different languages, and different accents where available. Shows speaking speed and pitch controls. Shows the local TTS provider in use (Piper TTS, Coqui TTS, or another suitable open-source TTS engine).
  • Primary actions: Select a voice; set speaking speed; set pitch; preview the voice; generate the voiceover.
  • Supporting actions: Switch TTS provider; regenerate a voiceover segment.
  • Domain entities: Voice, Language, Accent, Speaking Speed, Pitch, Voice Preview, Voiceover asset.
  • Component responsibilities: Voice picker; speed and pitch controls; preview player; provider selector reflecting the modular TTS architecture; generation control.
  • States: Loading (voice list resolving from the local engine); empty (no TTS engine installed — explain what TTS does, that it is required for voiceover, and how to install Piper/Coqui); success (voice selected, preview plays, voiceover generated); error (preview or generation fails — report the engine error); recovery (retry, or switch provider).

Generate Images

  • Information/state: Local/open-source image generation workspace. Shows the prompt, generated variations, and the currently selected image.
  • Primary actions: Generate image from prompt; generate multiple variations; select best image; regenerate; upload my own image; image-to-image; maintain character/style consistency.
  • Supporting actions: Attach a generated image to a scene; send an image to the Video Generator.
  • Domain entities: Image Prompt, Image Variation, Selected Image, Uploaded Image, Reference Image (for image-to-image and consistency).
  • Component responsibilities: Prompt input; variation grid with framed thumbnails; selection control; upload control; image-to-image reference control; consistency controls; provider indicator (Stable Diffusion/Flux-compatible local generation where hardware allows).
  • States: Loading (generation in progress with a running indicator); empty (no prompt yet, or no local image model installed — explain what image generation does, that it is optional, and how to install it); success (variations rendered, one selectable); error (generation fails — report the model/hardware error); recovery (retry, reduce variation count, or upload an image instead).
Page 7 of 40

Video Generator

  • Information/state: Pluggable local video generation workspace. Shows the selected backend and the hardware assessment. Exposes text-to-video, image-to-video, animate image, camera movement, scene duration, and motion strength.
  • Primary actions: Choose a backend; run text-to-video; run image-to-video; animate an image; set camera movement; set scene duration; set motion strength; generate.
  • Supporting actions: Switch to the image-based fallback (generated images, Ken Burns effect, zoom, pan, camera movement, transitions) when local video generation is unavailable.
  • Domain entities: Video Generation Backend, Motion Strength, Camera Movement, Scene Duration, Generated Clip, Fallback Motion Preset.
  • Component responsibilities: Backend selector reflecting the pluggable architecture; hardware capability banner; generation parameter panel; fallback switch; result preview.
  • States: Loading (backend availability check running); empty (no local video model installed, or insufficient GPU — state plainly that local video generation needs sufficient GPU resources and offer the fallback); success (clip generated and previewable); error (generation fails mid-run — report the failure and preserve the fallback path); recovery (retry, lower motion strength/duration, or switch to the image-based fallback).

Video Editor

  • Information/state: Timeline-based editor with eight tracks: Video, Images, Voiceover, Music, Sound Effects, Subtitles, Text, Overlays. Shows the ruler in tabular numerals, the playhead, and the selected clip's trim handles.
  • Primary actions: Cut; split; trim; crop; resize; rotate; speed control; volume control; fade in/out; transitions; text overlays; subtitles; stickers; images; audio; video layers. Render through FFmpeg.
  • Supporting actions: Add media from the Media Library; open Captions or Audio for their respective tracks; export.
  • Domain entities: Timeline, Track, Clip, Transition, Text Overlay, Sticker, Subtitle, Audio Layer, Video Layer.
  • Component responsibilities: Eight stacked tracks with a fixed label gutter; ruler; orange playhead; hazard-stripe trim handles on the selected clip; per-clip inspector for the full feature set; render control.
  • States: Loading (project timeline loading); empty (no clips on the timeline — prompt to add media or run a pipeline); success (timeline populated, edits applied, render completes); error (render fails — report the FFmpeg error and the failing clip/track); recovery (fix the offending clip and re-render; the timeline state is preserved).

Video → Clips

  • Information/state: Long-form ingestion and clip generation entry. Accepts MP4/MKV/MOV/WebM, a YouTube video file, a podcast, a lecture, an interview, or a long-form video. Shows the automatic sequence: transcribe the video, detect important/high-engagement moments, find potential clips, suggest clip start/end times, generate titles, generate captions, convert clips to 9:16, add animated subtitles, reframe the speaker, export multiple Shorts.
  • Primary actions: Upload a long video; start the automatic clipping run.
  • Supporting actions: Choose the transcription engine (Whisper/faster-whisper); monitor progress.
  • Domain entities: Source Video, Transcript, Detected Moment, Suggested Clip, Clip Title, Caption, 9:16 Conversion, Reframe.
  • Component responsibilities: Upload control with accepted-format list; run control; progress ledger for the ten automatic steps; handoff to Clips when suggestions are ready.
  • States: Loading (upload and transcription in progress); empty (no file selected); success (suggested clips produced and handed to Clips); error (unsupported format, or transcription unavailable — explain what Whisper does, that it is required for this step, and how to install it); recovery (retry, or convert the source to a supported format).
Page 8 of 40

Clips

  • Information/state: Suggested-clip workspace. Each suggestion is presented as a call-sheet tag showing the clip number, the start→end timecode range, a quoted moment label, and a score percentage — for example: CLIP 01, 00:14:32 → 00:15:18, "Most interesting moment", Score: 92%; CLIP 02, 00:28:10 → 00:29:03, "Strong hook", Score: 88%.
  • Primary actions: Preview a clip; edit a clip; export a clip; export multiple Shorts.
  • Supporting actions: Adjust start/end times; edit the generated title; apply animated subtitles; reframe the speaker; convert to 9:16.
  • Domain entities: Suggested Clip, Clip Number, Start/End Timecode, Moment Label, Score, Clip Title, Caption, Reframe, Exported Short.
  • Component responsibilities: Tag-style clip cards with score bars; preview player; per-clip edit controls; batch export control.
  • States: Loading (suggestions being computed); empty (no suggestions found — report that no high-engagement moments were detected and allow manual range selection); success (clips previewable and exportable); error (preview or export fails — report the failure for that clip only); recovery (retry the individual clip export).

Captions

  • Information/state: Subtitle and caption workspace. Shows the transcript with word-level timestamps and the applied caption style. Presets: Minimal, Bold, Shorts, Karaoke, Highlight, Professional.
  • Primary actions: Run automatic transcription; generate subtitles; apply a caption preset; highlight important words; export SRT/VTT; produce animated captions.
  • Supporting actions: Edit subtitle text and timing; switch subtitle style.
  • Domain entities: Transcript, Word-Level Timestamp, Subtitle, Caption Preset, SRT export, VTT export, Animated Caption.
  • Component responsibilities: Transcript/timeline view; preset selector; style preview; export controls; transcription engine indicator (Whisper/faster-whisper).
  • States: Loading (transcription running); empty (no transcript yet — prompt to transcribe); success (subtitles generated, preset applied, export available); error (transcription fails — report the engine error); recovery (retry, or import an existing transcript).

Audio

  • Information/state: Audio workspace covering background music, uploaded music, sound effects, audio trimming, volume adjustment, noise reduction, voice enhancement, and audio ducking. Shows the source/license of royalty-free/open-source music when applicable.
  • Primary actions: Add background music; upload music; add sound effects; trim audio; adjust volume; apply noise reduction; apply voice enhancement; apply audio ducking.
  • Supporting actions: Browse royalty-free/open-source music; inspect license details.
  • Domain entities: Background Music, Uploaded Music, Sound Effect, Audio Clip, Volume, Ducking Setting, License/Source.
  • Component responsibilities: Music and SFX browsers; waveform display as chunky 2px bars; per-clip audio inspector; license/source readout.
  • States: Loading (audio assets resolving); empty (no audio added); success (audio applied and audible in preview); error (unsupported audio file or processing failure — report it); recovery (retry or replace the asset).
Page 9 of 40

Generate Thumbnail

  • Information/state: Thumbnail workspace. Sources: video, screenshot, or prompt. Elements: text, images, background, AI-generated elements. Output: multiple variations, 1280×720 export.
  • Primary actions: Generate thumbnails from a video frame, a screenshot, or a prompt; add text; add images; set a background; add AI-generated elements; generate multiple variations; export at 1280×720.
  • Supporting actions: Select a variation; regenerate.
  • Domain entities: Thumbnail, Source Frame, Screenshot, Prompt, Text Element, Background, AI-Generated Element, 1280×720 Export.
  • Component responsibilities: Source selector; composition canvas; variation grid; export control.
  • States: Loading (variations generating); empty (no source selected); success (variations rendered, one exportable at 1280×720); error (generation or export fails — report it); recovery (retry, or switch source type).

Content Package

  • Information/state: Publishing metadata workspace, generated automatically after a video is created. Contains YouTube title, YouTube description, SEO keywords, hashtags, tags, thumbnail text, short description, and social media captions.
  • Primary actions: Review the generated package; copy individual fields; edit fields.
  • Supporting actions: Regenerate the package; send thumbnail text to Generate Thumbnail.
  • Domain entities: YouTube Title, YouTube Description, SEO Keywords, Hashtags, Tags, Thumbnail Text, Short Description, Social Media Captions.
  • Component responsibilities: Field-by-field blocks with copy controls; regeneration control; linkage to the thumbnail workspace.
  • States: Loading (package generating); empty (no completed video yet — explain that the package is produced after a video is created); success (all fields populated); error (generation fails — report it); recovery (retry generation).

Projects

  • Information/state: Durable project management. Lists projects with their state and last-saved time.
  • Primary actions: New Project; Save Project; Duplicate Project; Rename; Delete; Export; Import project.
  • Supporting actions: Open a project into its workspace; rely on automatic progress saving.
  • Domain entities: Project, Project Name, Save State, Exported Project File, Imported Project File.
  • Component responsibilities: Project list; per-project action row; import/export controls; autosave indicator.
  • States: Loading (project list resolving); empty (no projects — prompt to create one); success (project created/saved/duplicated/renamed/exported/imported); error (save, export, or import fails — report the reason and preserve the existing project); recovery (retry the operation; autosave retries on the next change).
Page 10 of 40

Media Library

  • Information/state: Stores images, videos, audio, voiceovers, music, SFX, thumbnails, and generated clips. Supports search, filtering, and folders.
  • Primary actions: Search; filter; create and use folders; open an asset; send an asset to a workspace.
  • Supporting actions: Delete or rename an asset; inspect asset metadata.
  • Domain entities: Media Asset (image, video, audio, voiceover, music, SFX, thumbnail, generated clip), Folder, Filter, Search Query.
  • Component responsibilities: Asset grid with framed thumbnails stamped with quoted aspect-ratio labels; folder tree; search and filter bar; asset detail panel.
  • States: Loading (assets indexing); empty (no media yet — prompt to generate or upload); success (assets browsable, searchable, filterable, foldered); error (an asset fails to load — show a placeholder and report it); recovery (refresh the asset).

System Status

  • Information/state: Displays the status of FFmpeg, Whisper, Ollama, TTS, Image Generation, Video Generation, GPU, and Storage. For each component that is not installed, shows what it does, whether it is optional, and installation instructions.
  • Primary actions: Read component status; open installation instructions for a missing component; re-check status.
  • Supporting actions: Jump to the workspace that depends on a given component.
  • Domain entities: Component (FFmpeg, Whisper, Ollama, TTS, Image Generation, Video Generation, GPU, Storage), Status, Optionality, Installation Instructions.
  • Component responsibilities: Eight-row status ledger with status chips; per-component detail panel with description, optionality, and install instructions; re-check control.
  • States: Loading (detection running); empty (n/a — all eight rows always render); success (each row shows ready or missing with its detail); error (detection itself fails — report that detection could not run and allow a re-check); recovery (re-run detection).

3. Functional Requirements

Page 11 of 40

Product and Architecture

FR-01 — General-purpose, self-hosted platform (explicit) As the Creator Studio Owner, I should run a personal, self-hosted AI video creation platform named "Creator Studio" that is general-purpose — usable for YouTube, Shorts, Reels, TikTok, educational videos, faceless videos, storytelling, documentaries, gaming content, nursery rhymes, and any other type of video — and explicitly not a kids-only website.

  • Trigger/input: Local installation and launch of the application.
  • Observable result: The application runs locally and presents general-purpose creation routes with no kids-only framing.
  • Access state: Local personal use; authentication optional.
  • Failure/recovery: If the backend is not running, the frontend reports it and states how to start it.
  • Continuation: The user proceeds to the Dashboard or a creation route.

FR-02 — No paid API dependencies by default (explicit) As the Creator Studio Owner, I should be able to complete the full video creation workflow without any paid API, because the platform is designed around free and open-source models and local processing; API integrations may exist but are optional.

  • Trigger/input: Any generation or processing action.
  • Observable result: The action completes using local/open-source tooling with no paid API call required.
  • Access state: Local.
  • Failure/recovery: If a required local tool is missing, the application names it, states whether it is optional, and gives installation instructions rather than silently substituting a paid service.
  • Continuation: The user installs the tool or uses the lightweight path.

FR-03 — Free/open-source, local-first software (explicit) As the Creator Studio Owner, I should have the software itself be free/open-source and local-first wherever possible.

  • Observable result: The application's own code and processing run locally and depend on free/open-source tooling.
  • Failure/recovery: Missing optional components degrade gracefully rather than blocking the whole application.
  • Continuation: The user continues with the capabilities their machine supports.

FR-04 — Honest hardware messaging (explicit) As the Creator Studio Owner, I should not be told that unlimited AI video generation is free when the required model needs expensive GPU hardware.

  • Trigger/input: Attempting local video generation, or viewing System Status / Video Generator.
  • Observable result: The application plainly states that local video generation depends on sufficient GPU resources.
  • Failure/recovery: When hardware is insufficient, the application offers the image-based fallback instead of implying free unlimited generation.
  • Continuation: The user chooses the fallback or upgrades hardware.

FR-05 — Pluggable AI provider architecture (explicit) As the Creator Studio Owner, I should have a clean modular architecture so I can replace any AI model later, and additional AI video models can be plugged in later without rewriting the entire application.

  • Trigger/input: Selecting or swapping a provider for TTS, LLM, image generation, video generation, or transcription.
  • Observable result: Providers are selectable and swappable without changing the rest of the application.
  • Failure/recovery: If a provider is unavailable, the application reports it and allows selecting another.
  • Continuation: The user continues with the newly selected provider.

FR-06 — Local installation commands (explicit) As the Creator Studio Owner, I should run the frontend locally with npm install and npm run dev, and the backend with pip install -r requirements.txt and python server.py.

  • Observable result: Both processes start and the application is usable locally.
  • Failure/recovery: Startup errors are reported with actionable detail.
  • Continuation: The user opens the application.

FR-07 — Windows installation README (explicit) As the Creator Studio Owner, I should have a README containing complete installation instructions for Windows.

  • Observable result: A README documents the complete Windows installation procedure.
  • Continuation: The user follows the README to install and run the application.

FR-08 — Real working application (explicit) As the Creator Studio Owner, I should receive a real working application, not merely a UI mockup.

  • Observable result: The described workflows actually execute and produce real outputs (voiceovers, images, subtitles, rendered MP4s, exported clips, thumbnails, metadata).
  • Failure/recovery: Failures are real, reported failures with recovery paths — not simulated success.
  • Continuation: The user completes real work in the application.

FR-09 — Local tool detection (explicit) As the Creator Studio Owner, I should have the application detect which local AI tools are installed and show their status.

  • Trigger/input: Application start, or an explicit re-check.
  • Observable result: Each detected component is reported as present or missing.
  • Failure/recovery: If detection cannot run, the application says so and allows a re-check.
  • Continuation: The user acts on the reported status.

FR-10 — System Status page (explicit) As the Creator Studio Owner, I should see a "System Status" page showing FFmpeg, Whisper, Ollama, TTS, Image Generation, Video Generation, GPU, and Storage; and for any component that isn't installed, I should see what it does, whether it is optional, and installation instructions.

  • Trigger/input: Opening System Status.
  • Observable result: Eight status rows render with ready/missing state; missing components expand to description, optionality, and install instructions.
  • Access state: Local.
  • Failure/recovery: Detection failure is reported with a re-check option.
  • Continuation: The user installs a missing component or proceeds with what is available.

FR-11 — Lightweight mode (explicit) As the Creator Studio Owner without a powerful GPU, I should have a lightweight mode using local LLM, local TTS, Whisper, user-provided images/videos, FFmpeg, motion effects, transitions, and captions, which still allows complete videos to be created without paid APIs.

  • Trigger/input: Selecting lightweight mode, or the application detecting insufficient GPU.
  • Observable result: A complete video can be produced using only the lightweight toolset.
  • Failure/recovery: If a lightweight component is missing, it is named with install instructions.
  • Continuation: The user exports a complete video without paid APIs.
Page 12 of 40

Main Workflow: Script → Video

FR-12 — Script input methods (explicit) As the Creator Studio Owner, I should paste a script, upload a TXT/DOCX/PDF, or enter a simple topic and let AI create the script.

  • Trigger/input: Paste, file upload, or topic entry.
  • Observable result: The script content is accepted into the project as the production source.
  • Failure/recovery: An unreadable or unsupported file is rejected with a clear message.
  • Continuation: The script proceeds to analysis.

FR-13 — Automatic script-to-video pipeline (explicit) As the Creator Studio Owner, I should have the platform automatically analyze the script, divide it into scenes, estimate duration, generate scene descriptions, generate visual prompts, generate voiceover, generate or select visuals, add background music, add sound effects, generate subtitles, automatically synchronize everything, assemble the final video, and export MP4.

  • Trigger/input: A supplied script and a start command.
  • Observable result: Each of the thirteen stages runs and reports its state; the final video is assembled and exported as MP4.
  • Access state: Local.
  • Failure/recovery: A failing stage is named with its reason; the pipeline halts at that stage rather than silently continuing; the user can fix the input and resume.
  • Continuation: The user exports the MP4 or opens the editor for refinement.
Page 13 of 40

Video Types and Formats

FR-14 — Video-type presets (explicit) As the Creator Studio Owner, I should choose from presets: YouTube video, YouTube Shorts, Instagram Reels, TikTok, Educational video, Story video, Faceless video, Kids animation, Documentary, Motivational video, Podcast clips, and Custom.

  • Trigger/input: Selecting a preset.
  • Observable result: The preset's production parameters are applied to the project.
  • Failure/recovery: If a preset cannot be applied, the reason is shown and Custom remains available.
  • Continuation: The user proceeds to source selection.

FR-15 — Custom format controls (explicit) As the Creator Studio Owner, I should set custom aspect ratio, resolution, FPS, duration, and style.

  • Trigger/input: Editing custom parameters.
  • Observable result: The custom values are applied to the project.
  • Failure/recovery: An invalid combination is rejected with inline validation.
  • Continuation: The user continues with valid custom settings.

FR-16 — Supported formats (explicit) As the Creator Studio Owner, I should be able to use 16:9, 9:16, 1:1, and 4:5.

  • Observable result: Each of the four aspect ratios is selectable and applied to output.
  • Failure/recovery: An unsupported ratio is not offered.
  • Continuation: The user proceeds with a supported ratio.
Page 14 of 40

AI Script Generator

FR-17 — Script generator inputs (explicit) As the Creator Studio Owner, I should enter Topic, Audience, Video length, Language, Tone, and Style into the script generator.

  • Trigger/input: Filling the six inputs.
  • Observable result: The inputs are captured and used for generation.
  • Failure/recovery: Missing required inputs are flagged before generation.
  • Continuation: The user generates the script.

FR-18 — Script generator outputs (explicit) As the Creator Studio Owner, I should receive a Hook, Full script, Narration, Scene breakdown, Dialogue, Visual instructions, and CTA.

  • Trigger/input: Running generation.
  • Observable result: All seven output sections are produced.
  • Failure/recovery: If the local LLM is unavailable, the application explains what Ollama does, that it is required for this step, and how to install it.
  • Continuation: The user reviews the output.

FR-19 — Edit generated script before video creation (explicit) As the Creator Studio Owner, I should be able to edit the generated script before creating the video.

  • Trigger/input: Editing the generated script.
  • Observable result: Edits are saved and used as the production source.
  • Failure/recovery: Unsaved edits are preserved if generation is re-run, with a clear prompt.
  • Continuation: The user proceeds to video creation with the edited script.
Page 15 of 40

AI Scene Generator

FR-20 — Automatic scene division (explicit) As the Creator Studio Owner, I should have scripts automatically divided into scenes.

  • Trigger/input: A script in the project.
  • Observable result: An ordered scene list is produced.
  • Failure/recovery: If division fails, the reason is reported and the script is preserved.
  • Continuation: The user reviews the scenes.

FR-21 — Scene fields (explicit) As the Creator Studio Owner, each scene should contain scene number, duration, narration, visual description, image prompt, video prompt, camera movement, transition, sound effect, background music, and subtitle text.

  • Observable result: All eleven fields are present and editable per scene.
  • Failure/recovery: A field that fails to generate is marked and can be regenerated or edited manually.
  • Continuation: The user refines scenes.

FR-22 — Scene operations (explicit) As the Creator Studio Owner, I should be able to regenerate scene, edit scene, delete scene, duplicate scene, reorder scene, regenerate visual, and change duration.

  • Trigger/input: Invoking any scene operation.
  • Observable result: The scene list reflects the operation immediately.
  • Failure/recovery: A failed regeneration preserves the previous scene content and reports the failure.
  • Continuation: The user continues editing or proceeds to production.
Page 16 of 40

Text to Speech

FR-23 — Free/local TTS integration (explicit) As the Creator Studio Owner, I should have free/local TTS options integrated, such as Piper TTS, Coqui TTS, or other suitable open-source TTS engines.

  • Trigger/input: Opening AI Voice.
  • Observable result: Available local TTS engines and their voices are listed.
  • Failure/recovery: If no TTS engine is installed, the application explains what TTS does, that it is required for voiceover, and how to install it.
  • Continuation: The user selects a voice.

FR-24 — Voice options (explicit) As the Creator Studio Owner, I should have male voices, female voices, different languages, different accents where available, speaking speed, pitch, and voice preview.

  • Trigger/input: Selecting a voice and adjusting speed/pitch.
  • Observable result: The chosen voice, language, accent, speed, and pitch are applied; preview plays the configured voice.
  • Failure/recovery: A preview failure is reported with the engine error.
  • Continuation: The user generates the voiceover.

FR-25 — Modular TTS providers (explicit) As the Creator Studio Owner, I should have the TTS system kept modular so additional TTS providers can be added later.

  • Observable result: TTS providers are selectable and additional providers can be added without rewriting the application.
  • Failure/recovery: An unavailable provider is reported and another can be selected.
  • Continuation: The user continues with the selected provider.
Page 17 of 40

AI Image Generation

FR-26 — Local/open-source image generation (explicit) As the Creator Studio Owner, I should have local/open-source image generation supported where possible.

  • Trigger/input: Opening Generate Images.
  • Observable result: Local image generation is available when the hardware and model allow.
  • Failure/recovery: If no local image model is installed, the application explains what image generation does, that it is optional, and how to install it.
  • Continuation: The user generates or uploads images.

FR-27 — Image generation operations (explicit) As the Creator Studio Owner, I should be able to generate image from prompt, generate multiple variations, select best image, regenerate, upload my own image, image-to-image, and maintain character/style consistency.

  • Trigger/input: A prompt, a reference image, or an upload.
  • Observable result: Variations are produced, one is selectable, uploads are accepted, and consistency controls are applied.
  • Failure/recovery: A failed generation is reported and the previous selection is preserved.
  • Continuation: The user attaches the selected image to a scene or the timeline.
Page 18 of 40

AI Video Generation

FR-28 — Pluggable local video generation (explicit) As the Creator Studio Owner, I should have a video-generation module designed to support locally available/open-source models when my computer has sufficient GPU resources, with pluggable backends rather than being locked to one provider.

  • Trigger/input: Opening Video Generator.
  • Observable result: Available local backends are listed and selectable.
  • Failure/recovery: Insufficient GPU or a missing model is reported plainly, with the fallback offered.
  • Continuation: The user generates a clip or switches to the fallback.

FR-29 — Video generation operations (explicit) As the Creator Studio Owner, I should be able to use text-to-video, image-to-video, animate image, camera movement, scene duration, and motion strength.

  • Trigger/input: Selecting a mode and setting parameters.
  • Observable result: A generated clip reflects the chosen mode, camera movement, duration, and motion strength.
  • Failure/recovery: A failed generation is reported; parameters can be reduced and retried.
  • Continuation: The user places the clip on the timeline.

FR-30 — Graceful image-based fallback (explicit) As the Creator Studio Owner, when local video generation is unavailable because of hardware limitations, I should get a graceful fallback to image-based video creation using generated images, Ken Burns effect, zoom, pan, camera movement, and transitions.

  • Trigger/input: Local video generation unavailable.
  • Observable result: An image-based video is produced using the listed motion effects and transitions.
  • Failure/recovery: If the fallback also fails, the reason is reported and the images remain available.
  • Continuation: The user exports the image-based video.
Page 19 of 40

Video Editor

FR-31 — Timeline tracks (explicit) As the Timeline Editor, I should have a proper timeline-based video editor with tracks for Video, Images, Voiceover, Music, Sound Effects, Subtitles, Text, and Overlays.

  • Trigger/input: Opening Video Editor.
  • Observable result: Eight tracks render with a fixed label gutter and a ruler.
  • Failure/recovery: A track that fails to load is reported without blocking the others.
  • Continuation: The user places and edits clips.

FR-32 — Editing features (explicit) As the Timeline Editor, I should be able to cut, split, trim, crop, resize, rotate, control speed, control volume, apply fade in/out, apply transitions, add text overlays, add subtitles, add stickers, add images, add audio, and work with video layers.

  • Trigger/input: Selecting a clip and invoking a feature.
  • Observable result: The edit is applied to the timeline and reflected in the preview.
  • Failure/recovery: An invalid edit is rejected with a clear message and the timeline is unchanged.
  • Continuation: The user continues editing or renders.

FR-33 — FFmpeg processing and rendering (explicit) As the Timeline Editor, I should have FFmpeg used for video processing and rendering.

  • Trigger/input: Rendering the timeline.
  • Observable result: FFmpeg processes and renders the final output.
  • Failure/recovery: An FFmpeg failure is reported with the failing clip/track identified; the timeline state is preserved.
  • Continuation: The user fixes the issue and re-renders.
Page 20 of 40

Clipping Tool

FR-34 — Long-form input support (explicit) As the Clip Repurposer, I should upload MP4/MKV/MOV/WebM, a YouTube video file, a podcast, a lecture, an interview, or a long-form video into the clipping tool.

  • Trigger/input: Uploading a long-form file.
  • Observable result: The file is accepted and queued for processing.
  • Failure/recovery: An unsupported format is rejected with a clear message.
  • Continuation: The automatic clipping sequence begins.

FR-35 — Automatic clipping sequence (explicit) As the Clip Repurposer, I should have the system automatically transcribe the video, detect important/high-engagement moments, find potential clips, suggest clip start/end times, generate titles, generate captions, convert clips to 9:16, add animated subtitles, reframe the speaker, and export multiple Shorts.

  • Trigger/input: A queued long-form video.
  • Observable result: Each of the ten steps runs and reports progress; multiple Shorts are exportable.
  • Failure/recovery: A failing step is named with its reason; completed steps are preserved.
  • Continuation: The user reviews suggestions in Clips.

FR-36 — Suggested clip presentation (explicit) As the Clip Repurposer, I should see suggested clips presented with a clip number, a start→end timecode range, a moment label, and a score percentage — for example CLIP 01, 00:14:32 → 00:15:18, "Most interesting moment", Score: 92%; and CLIP 02, 00:28:10 → 00:29:03, "Strong hook", Score: 88%.

  • Observable result: Each suggestion renders with all four elements.
  • Failure/recovery: If no moments are detected, the application says so and allows manual range selection.
  • Continuation: The user previews a clip.

FR-37 — Clip preview, edit, and export (explicit) As the Clip Repurposer, I should be able to preview, edit, and export each clip.

  • Trigger/input: Selecting a suggested clip.
  • Observable result: The clip previews, edits are applied, and the clip exports.
  • Failure/recovery: A failed export is reported for that clip only; other clips remain exportable.
  • Continuation: The user exports the remaining clips.
Page 21 of 40

Auto Captions

FR-38 — Free/open-source speech recognition (explicit) As the Creator Studio Owner, I should have auto captions use free/open-source speech recognition such as Whisper or faster-whisper.

  • Trigger/input: Running transcription.
  • Observable result: A transcript is produced locally.
  • Failure/recovery: If Whisper is not installed, the application explains what it does, that it is required for this step, and how to install it.
  • Continuation: The user generates subtitles.

FR-39 — Caption features (explicit) As the Creator Studio Owner, I should have automatic transcription, word-level timestamps, subtitle generation, SRT/VTT export, animated captions, highlight important words, and multiple subtitle styles.

  • Trigger/input: A transcript and a caption configuration.
  • Observable result: Subtitles are generated with word-level timing, important words can be highlighted, animated captions render, and SRT/VTT files export.
  • Failure/recovery: An export or render failure is reported with the reason.
  • Continuation: The user applies the captions to the video.

FR-40 — Caption presets (explicit) As the Creator Studio Owner, I should be able to use caption presets: Minimal, Bold, Shorts, Karaoke, Highlight, and Professional.

  • Trigger/input: Selecting a preset.
  • Observable result: The selected style is applied to the captions.
  • Failure/recovery: A preset that fails to apply leaves the previous style intact and reports the failure.
  • Continuation: The user exports or continues editing.
Page 22 of 40

Audio

FR-41 — Audio capabilities (explicit) As the Creator Studio Owner, I should have background music, upload music, sound effects, audio trimming, volume adjustment, noise reduction, voice enhancement, and audio ducking.

  • Trigger/input: Adding or selecting an audio asset and invoking a capability.
  • Observable result: The audio is added or processed as requested and is audible in preview.
  • Failure/recovery: An unsupported file or processing failure is reported.
  • Continuation: The user continues editing or renders.

FR-42 — Royalty-free/open-source music with license indication (explicit) As the Creator Studio Owner, I should have royalty-free/open-source music used by default, with the source/license clearly indicated when applicable.

  • Trigger/input: Adding background music.
  • Observable result: The music's source and license are displayed alongside the asset.
  • Failure/recovery: If license information is unavailable for an asset, that is stated rather than omitted silently.
  • Continuation: The user keeps or replaces the track.
Page 23 of 40

Thumbnail Generator

FR-43 — Thumbnail sources (explicit) As the Creator Studio Owner, I should generate YouTube thumbnails from a video, a screenshot, or a prompt.

  • Trigger/input: Selecting a source.
  • Observable result: Thumbnail variations are generated from the chosen source.
  • Failure/recovery: A failed generation is reported and another source can be chosen.
  • Continuation: The user refines the thumbnail.

FR-44 — Thumbnail elements and export (explicit) As the Creator Studio Owner, I should be able to use text, images, background, and AI-generated elements, generate multiple variations, and export at 1280×720.

  • Trigger/input: Composing the thumbnail and exporting.
  • Observable result: Multiple variations are produced and a 1280×720 file is exported.
  • Failure/recovery: An export failure is reported with the reason.
  • Continuation: The user uses the exported thumbnail.

AI Content Package

FR-45 — Automatic content package (explicit) As the Creator Studio Owner, after creating a video I should automatically get a YouTube title, YouTube description, SEO keywords, hashtags, tags, thumbnail text, short description, and social media captions.

  • Trigger/input: Completing a video.
  • Observable result: All eight package fields are generated.
  • Failure/recovery: A generation failure is reported and can be retried.
  • Continuation: The user copies or edits the fields.
Page 24 of 40

Project Management

FR-46 — Project operations (explicit) As the Creator Studio Owner, I should be able to create a New Project, Save Project, Duplicate Project, Rename, Delete, Export, and Import project.

  • Trigger/input: Invoking a project operation.
  • Observable result: The project list and project state reflect the operation.
  • Failure/recovery: A failed save, export, or import is reported and the existing project is preserved.
  • Continuation: The user continues working in the project.

FR-47 — Automatic progress saving (explicit) As the Creator Studio Owner, my progress should be saved automatically.

  • Trigger/input: Any change to project state.
  • Observable result: Progress is persisted and can be resumed.
  • Failure/recovery: A failed autosave is surfaced and retried on the next change.
  • Continuation: The user resumes work later without losing progress.
Page 25 of 40

Media Library

FR-48 — Media storage (explicit) As the Creator Studio Owner, I should have a media library that stores images, videos, audio, voiceovers, music, SFX, thumbnails, and generated clips.

  • Observable result: All eight asset categories are stored and browsable.
  • Failure/recovery: An asset that fails to load shows a placeholder and is reported.
  • Continuation: The user reuses assets in projects.

FR-49 — Search, filtering, and folders (explicit) As the Creator Studio Owner, I should be able to search, filter, and organize media into folders.

  • Trigger/input: A search query, a filter, or a folder action.
  • Observable result: The library view reflects the query, filter, or folder organization.
  • Failure/recovery: A failed folder operation is reported and the library state is preserved.
  • Continuation: The user selects an asset.
Page 26 of 40

Dashboard

FR-50 — Dashboard creation routes (explicit) As the Creator Studio Owner, I should have a beautiful professional dashboard with "Create Video", "Generate Script", "Script → Video", "Video → Clips", "AI Voice", "Generate Images", "Generate Thumbnail", and "Video Editor".

  • Observable result: All eight routes are present and launch their respective workspaces.
  • Failure/recovery: A route that cannot open reports why.
  • Continuation: The user enters a workspace.

FR-51 — Dashboard information modules (explicit) As the Creator Studio Owner, the dashboard should show Recent Projects, Storage, Processing Jobs, and System Status.

  • Observable result: All four modules render with current data.
  • Failure/recovery: A failed module shows an inline error without blocking the others.
  • Continuation: The user opens a project, inspects storage, monitors a job, or opens System Status.

First Working Version Priorities

FR-52 — First working version priority order (explicit) As the Creator Studio Owner, the first working version should prioritize: 1) Script → Scene Breakdown, 2) Script → Voiceover, 3) Script → Images, 4) Images + Voiceover → Video, 5) Automatic subtitles, 6) FFmpeg rendering, 7) Video → Automatic Clips, 8) Shorts/Reels 9:16 export, 9) Thumbnail generation, 10) YouTube metadata generation.

  • Observable result: These ten capabilities are functional in the first working version.
  • Failure/recovery: A prioritized capability that is not yet functional is clearly marked rather than presented as working.
  • Continuation: The user completes the prioritized end-to-end workflow.
Page 27 of 40

Technical and Security

FR-53 — Frontend stack (explicit) As the Creator Studio Owner, the frontend should be React or Next.js with TypeScript and Tailwind CSS.

  • Observable result: The frontend is built with these technologies.
  • Continuation: The user runs the frontend locally.

FR-54 — Backend stack (explicit) As the Creator Studio Owner, the backend should be Python with FastAPI.

  • Observable result: The backend is a Python FastAPI service started with python server.py.
  • Continuation: The frontend communicates with the local backend.

FR-55 — Database and storage (explicit) As the Creator Studio Owner, I should have SQLite for local installation with an option to switch to PostgreSQL later, and local filesystem storage by default.

  • Observable result: Project and asset state persist in SQLite; media is written to the local filesystem; a documented path exists to switch to PostgreSQL.
  • Failure/recovery: A database or filesystem error is reported with the affected operation.
  • Continuation: The user continues working with persisted state.

FR-56 — Optional authentication (explicit) As the Creator Studio Owner, authentication should be optional for local personal use.

  • Observable result: The application is usable locally without requiring authentication.
  • Failure/recovery: If authentication is enabled, a failed sign-in is reported without blocking the local-only configuration.
  • Continuation: The user continues working.

FR-57 — API key handling (explicit) As the Creator Studio Owner, all API keys must be stored in environment variables and never exposed in frontend code.

  • Trigger/input: Configuring an optional API integration.
  • Observable result: Keys live in backend environment variables and never appear in frontend code or responses.
  • Failure/recovery: A missing key for an optional integration is reported without breaking local-only operation.
  • Continuation: The user continues with local processing or the configured optional integration.

FR-58 — Local AI toolchain (explicit) As the Creator Studio Owner, the platform should use, where appropriate: FFmpeg for video processing, Whisper/faster-whisper for transcription, Piper/other open-source TTS for voice, open-source LLMs through Ollama, Stable Diffusion/Flux-compatible local image generation where hardware allows, open-source video-generation models where hardware allows, OpenCV for video/image processing, Python for AI/video processing, FastAPI or Flask for backend, and React/Next.js for frontend.

  • Observable result: These tools are the platform's processing foundation, used where appropriate.
  • Failure/recovery: A missing tool is reported with its role, optionality, and installation instructions.
  • Continuation: The user installs the tool or uses an available alternative.
Page 28 of 40

Design

FR-59 — Modern dark professional interface (explicit) As the Creator Studio Owner, I should have a modern dark professional interface similar to professional creative software, with a dark background, clean cards, subtle gradients, smooth animations, modern typography, clear icons, and responsive design — and it must not look like a children's website.

  • Observable result: The interface renders dark, professional, responsive, and free of kids-site signalling.
  • Continuation: The user works comfortably across viewports.

FR-60 — Personal workflow feel (explicit) As the Creator Studio Owner, the product should feel like a personal combination of an AI script generator, CapCut, Canva, an AI video generator, Whisper transcription, and an FFmpeg video processor, completely focused on my own personal workflow.

  • Observable result: The interface and workflows read as one personal studio rather than a generic multi-tenant SaaS product.
  • Continuation: The user completes their own end-to-end workflow.

4. User Personas

Page 29 of 40

Creator Studio Owner (Solo Creator)

Product context: The single personal user of a self-hosted, local-first studio running on their own machine. They own the hardware, install the tools, and drive every step of production themselves. Their machine's capabilities — GPU, storage, installed models — directly shape what they can do, which is why System Status is part of their regular loop.

Primary goal: Produce a complete, publishable video and its metadata without paid APIs.

Distinct accepted responsibilities: They paste or upload a script (or enter a topic) and drive it through scene breakdown, voiceover, images, subtitles, and FFmpeg rendering to an exported MP4. They choose video-type presets or custom aspect ratio, resolution, FPS, duration, and style. They generate and edit scripts, refine scenes, configure local TTS voices, generate and select images, manage audio and music licensing, generate thumbnails, review the AI content package, manage projects, and organize the media library.

Relevant inputs or decisions: Script text or file; topic, audience, video length, language, tone, and style; preset vs. custom format; voice, language, accent, speed, and pitch; image prompts and reference images; caption preset; music and SFX choices; thumbnail source; whether to run the full pipeline or the lightweight path.

Interactions with other accepted participants: The Owner is the initiator of the script-to-video lifecycle. Their work hands off to the Clip Repurposer workflow when they feed long-form material in, and to the Timeline Editor workflow when they refine a rendered timeline. They also read the System Status page to see which local AI tools are installed and what to install.

Observable success: A complete, publishable video and its metadata, produced without paid APIs, with the local toolchain status understood and any missing components installed.

Constraints carried from source: No paid API dependencies by default; the software is free/open-source and local-first wherever possible; the interface must not look like a children's website; authentication is optional for local personal use.

Page 30 of 40

Clip Repurposer

Product context: The same personal workflow applied to long-form input rather than a script. This is a distinct working context: the source is an existing recording, the authoritative state is a transcript plus a set of scored moment suggestions, and the completion outcome is a batch of exported vertical clips rather than a single assembled video.

Primary goal: Turn one long recording into a set of exported vertical clips ready for Shorts, Reels, and TikTok.

Distinct accepted responsibilities: Uploading an MP4/MKV/MOV/WebM, a YouTube video file, a podcast, a lecture, an interview, or a long-form video; running the automatic sequence of transcription, high-engagement moment detection, potential clip discovery, start/end suggestion, title generation, caption generation, 9:16 conversion, animated subtitle addition, speaker reframing, and multiple Shorts export; reviewing scored suggestions; previewing, editing, and exporting each clip.

Relevant inputs or decisions: The long-form file; which suggested clips to keep; adjusted start/end times; edited titles; caption styling; reframing choices; which clips to export.

Interactions with other accepted participants: The Clip Repurposer consumes output produced by the Owner's long-form upload and hands refined clips onward to the Timeline Editor when a clip needs manual timeline work. They rely on the same local Whisper/faster-whisper transcription the Owner uses for captions.

Observable success: A set of exported vertical clips, each with a title and animated captions, ready to publish.

Constraints carried from source: The clipping tool is explicitly very important; transcription must use free/open-source speech recognition; local video generation is not required for this workflow.

Page 31 of 40

Timeline Editor

Product context: The editing-focused workflow in the timeline-based editor. This is a distinct working context from generation: the authoritative state is the timeline itself — eight tracks of clips, transitions, and layers — and the completion outcome is a correctly rendered timeline rather than a generated asset.

Primary goal: Manually refine a timeline so it renders correctly through FFmpeg.

Distinct accepted responsibilities: Working across video, image, voiceover, music, sound effect, subtitle, text, and overlay tracks; cutting, splitting, trimming, cropping, resizing, rotating, adjusting speed and volume, applying fades and transitions, and layering text, stickers, images, and audio.

Relevant inputs or decisions: Which clips to place on which track; trim and split points; crop, resize, and rotation values; speed and volume levels; fade and transition choices; text, sticker, image, and audio layer placement; when to render.

Interactions with other accepted participants: The Timeline Editor receives generated assets from the Owner's pipeline and refined clips from the Clip Repurposer, and returns a rendered output. They depend on FFmpeg being installed and on the Captions and Audio workspaces for their respective tracks.

Observable success: A manually refined timeline that renders correctly through FFmpeg.

Constraints carried from source: FFmpeg must be used for video processing and rendering; the editor must be timeline-based with the eight specified tracks.

Page 32 of 40

5. Core User Flows

Flow 1 — Owner: Script → Video end-to-end

  1. The Owner starts the backend with pip install -r requirements.txt and python server.py, and the frontend with npm install and npm run dev, then opens the application.
  2. On Landing, the Owner reads what Creator Studio is, confirms the no-paid-API-by-default promise, and enters the studio.
  3. On Dashboard, the Owner sees Recent Projects, Storage, Processing Jobs, and System Status, and chooses "Script → Video".
  4. On Create Video, the Owner selects a preset (for example YouTube video) or Custom, and sets aspect ratio, resolution, FPS, duration, and style — choosing from 16:9, 9:16, 1:1, or 4:5.
  5. The Owner supplies the script source: pastes a script, uploads a TXT/DOCX/PDF, or enters a simple topic and lets AI create the script.
  6. If a topic was entered, the Owner goes to Generate Script, fills Topic, Audience, Video length, Language, Tone, and Style, and generates. The application produces the Hook, Full script, Narration, Scene breakdown, Dialogue, Visual instructions, and CTA. The Owner edits the generated script before creating the video.
  7. The Owner starts the pipeline on Script → Video. The thirteen stages run in order: analyze the script, divide it into scenes, estimate duration, generate scene descriptions, generate visual prompts, generate voiceover, generate or select visuals, add background music, add sound effects, generate subtitles, synchronize everything, assemble the final video, export MP4. Each stage shows DONE / RUNNING / QUEUED.
  8. On Scenes, the Owner reviews the breakdown. Each scene shows scene number, duration, narration, visual description, image prompt, video prompt, camera movement, transition, sound effect, background music, and subtitle text. The Owner regenerates, edits, deletes, duplicates, reorders, regenerates visuals, or changes durations as needed.
  9. On AI Voice, the Owner picks a male or female voice, a language, and an accent where available, sets speaking speed and pitch, previews the voice, and generates the voiceover.
  10. On Generate Images, the Owner generates images from prompts, produces multiple variations, selects the best, regenerates, uploads their own images, uses image-to-image, and maintains character/style consistency.
  11. On Video Generator, the Owner either generates video with a local backend (text-to-video, image-to-video, animate image, camera movement, scene duration, motion strength) or, when the GPU is insufficient, accepts the graceful fallback to image-based video creation using generated images, Ken Burns effect, zoom, pan, camera movement, and transitions.
  12. On Captions, the Owner runs automatic transcription with Whisper/faster-whisper, gets word-level timestamps, generates subtitles, applies a preset (Minimal, Bold, Shorts, Karaoke, Highlight, or Professional), highlights important words, and exports SRT/VTT.
  13. On Audio, the Owner adds background music (royalty-free/open-source by default, with source/license clearly indicated), uploads music, adds sound effects, trims audio, adjusts volume, applies noise reduction, voice enhancement, and audio ducking.
  14. The pipeline assembles the final video and exports MP4. The Owner opens Video Editor to refine the timeline if needed.
  15. On Content Package, the Owner reviews the automatically generated YouTube title, YouTube description, SEO keywords, hashtags, tags, thumbnail text, short description, and social media captions.
  16. On Generate Thumbnail, the Owner generates thumbnails from the video, a screenshot, or a prompt, adds text, images, background, and AI-generated elements, produces multiple variations, and exports at 1280×720.
  17. Failure/recovery: If a stage fails, the failing stage is named with its reason and the pipeline halts there. The Owner fixes the input in the owning workspace (for example installing a missing local tool reported on System Status) and resumes. If the local LLM, TTS, or Whisper is missing, the application explains what it does, whether it is optional, and how to install it.
  18. Continuation: The Owner's project is saved automatically and appears in Recent Projects on the Dashboard and in Projects and the Media Library.
Page 33 of 40

Flow 2 — Owner: Lightweight mode on a machine without a powerful GPU

  1. On System Status, the Owner sees that GPU is insufficient for local video generation, and that FFmpeg, Whisper, Ollama, and TTS are present.
  2. The Owner reads the honest explanation that local video generation depends on sufficient GPU resources, and that the software itself remains free/open-source and local-first.
  3. The Owner selects lightweight mode, which uses local LLM, local TTS, Whisper, user-provided images/videos, FFmpeg, motion effects, transitions, and captions.
  4. The Owner supplies their own images or videos, generates a script and voiceover locally, and lets FFmpeg apply motion effects, transitions, and captions.
  5. Observable result: A complete video is produced and exported without any paid API.
  6. Continuation: The Owner proceeds to captions, thumbnail, and the content package exactly as in Flow 1.

Flow 3 — Owner: Generate a script only

  1. On Dashboard, the Owner chooses "Generate Script".
  2. On Generate Script, the Owner enters Topic, Audience, Video length, Language, Tone, and Style.
  3. The application generates the Hook, Full script, Narration, Scene breakdown, Dialogue, Visual instructions, and CTA.
  4. The Owner edits the generated script.
  5. Failure/recovery: If Ollama is not installed, the application explains what it does, that it is required for this step, and how to install it; the Owner can instead paste a script manually.
  6. Continuation: The Owner sends the edited script into the Script → Video pipeline.

Flow 4 — Clip Repurposer: Long video → multiple Shorts

  1. On Dashboard, the Clip Repurposer chooses "Video → Clips".
  2. On Video → Clips, they upload an MP4/MKV/MOV/WebM, a YouTube video file, a podcast, a lecture, an interview, or a long-form video.
  3. The automatic sequence runs: transcribe the video, detect important/high-engagement moments, find potential clips, suggest clip start/end times, generate titles, generate captions, convert clips to 9:16, add animated subtitles, reframe the speaker, and prepare multiple Shorts for export.
  4. On Clips, the suggestions appear as call-sheet tags. For example: CLIP 01, 00:14:32 → 00:15:18, "Most interesting moment", Score: 92%; and CLIP 02, 00:28:10 → 00:29:03, "Strong hook", Score: 88%.
  5. The Clip Repurposer previews a clip, edits its start/end times and title, applies animated subtitles, reframes the speaker, and confirms the 9:16 conversion.
  6. The Clip Repurposer exports the clip, then exports the remaining clips as multiple Shorts.
  7. Failure/recovery: If transcription is unavailable, the application explains what Whisper does, that it is required for this step, and how to install it. If a single clip export fails, only that clip is affected and can be retried.
  8. Continuation: Exported clips land in the Media Library under generated clips, and any clip needing manual work is opened in Video Editor.
Page 34 of 40

Flow 5 — Timeline Editor: Refine and render a timeline

  1. From a completed pipeline run or an exported clip, the Timeline Editor opens Video Editor.
  2. The timeline shows eight tracks: Video, Images, Voiceover, Music, Sound Effects, Subtitles, Text, and Overlays, with a ruler in tabular numerals and an orange playhead.
  3. The Timeline Editor cuts, splits, trims, crops, resizes, rotates, adjusts speed and volume, applies fade in/out and transitions, and adds text overlays, subtitles, stickers, images, audio, and video layers.
  4. For subtitle work they open Captions; for music, SFX, trimming, volume, noise reduction, voice enhancement, and ducking they open Audio.
  5. The Timeline Editor renders the timeline; FFmpeg performs the processing and rendering.
  6. Failure/recovery: If rendering fails, the FFmpeg error is reported with the failing clip or track identified, and the timeline state is preserved so the Editor can fix the offending clip and re-render.
  7. Continuation: The rendered output is exported and the project is saved automatically.

Flow 6 — Owner: Project and media management

  1. On Projects, the Owner creates a New Project, saves, duplicates, renames, deletes, exports, or imports a project.
  2. Progress is saved automatically as the Owner works, so production and editing can be resumed later.
  3. On Media Library, the Owner browses stored images, videos, audio, voiceovers, music, SFX, thumbnails, and generated clips, and uses search, filtering, and folders to organize them.
  4. Failure/recovery: A failed save, export, or import is reported and the existing project is preserved; a failed autosave is surfaced and retried on the next change.
  5. Continuation: The Owner reuses assets across projects and resumes work from where they left off.

Flow 7 — Owner: Check and install local tools

  1. On Dashboard, the Owner opens System Status (also reachable from the persistent right-edge system ledger).
  2. The page shows FFmpeg, Whisper, Ollama, TTS, Image Generation, Video Generation, GPU, and Storage, each with a ready or missing state.
  3. For any component that isn't installed, the Owner sees what it does, whether it is optional, and installation instructions.
  4. The Owner installs the missing component and re-checks status.
  5. Failure/recovery: If detection itself cannot run, the application says so and allows a re-check.
  6. Continuation: The Owner returns to the workflow that needed the component.
Page 35 of 40

6. Visuals, Colors and Theme

Muse and headline: Virgil Abloh — Industrial remix: a creator's workbench tagged in safety orange. The emotional register is workshop, not showroom: a machine room of FFmpeg, Whisper, Piper and Ollama that the user owns and drives themselves. Professional creative software (dark, dense, capable) with the swagger of a maker's own tool — never a kids' site, never a generic SaaS dashboard.

Color tokens (dark mode):

RoleHexUsage
Background#0B0B0CConcrete floor; ~60% of pixels
Surface#151517Panel/steel plate for cards, timeline tracks, inspector rails; ~25%
Text#F2F2EFOff-white paint, never pure white; ~15:1 on background
Primary#FF5A00Safety orange — the single hot accent: active action, playhead, render-progress fill, "REC" state, primary CTA; under 8% of pixels
Accent#FFE600Hazard yellow — construction-tape stripes, quote marks around labels, highlight color inside animated captions; never a button fill
Muted#8A8A85Metadata, timestamps, inactive track labels, disabled states; 4.6:1 on #0B0B0C
Hairline border#2A2A2C1px borders on every surface
Grid line#1E1E20Faint 12-column cutting-mat grid

No blue anywhere in the interface chrome; the only blue on screen is the user's own media.

Typography:

  • Headings: Archivo Black, all-caps, 0.94 line-height, -0.01em tracking. Set enormous for page titles and section openers, small-and-loud for panel headers.
  • Functional labels: Archivo at 11–12px, uppercase, 0.18em letterspacing, in muted grey, wrapped in literal quotation marks — "SCENE 04", "TRACK 06 / SFX", "PROVIDER: PIPER", "STATUS: OK".
  • Numerals: Tabular, Archivo SemiBold, so timecodes (00:14:32 → 00:15:18) align in a column like a call sheet.
  • Body: Archivo.
  • Type scale: 1.25 modular, mobile-first — display 40/56/88/128 (clamp), section 28/34, panel header 18, body 15/16, label 11/12, timecode 13 tabular. Display sizes: 40px @375, 56px @768, 88–128px @1280.
  • Monospace: Reserved strictly for code and terminal output in the install instructions.

Shape language: Hard industrial geometry. 0–2px radii on panels and buttons (no pills, no blobs). 1px hairline borders in #2A2A2C on every surface. 4px exposed grid lines drawn on the workspace background like a cutting mat. Chunky 2px black outlines around media thumbnails and clip cards so they read as physical objects on a table. Hazard stripes (repeating 45° bands of #FFE600 and #0B0B0C, 10px pitch) used as 8px dividers, as the top edge of the render queue, and as the trim handles on selected timeline clips. Tag/label motif: a small rotated rectangle with a punched hole and a 1px rule, used for scene numbers, provider chips and clip scores.

Spacing rhythm: Exposed 12-column grid with visible 1px column rules on the dashboard and workspace, like a spec sheet. Left rail 72px collapsed / 248px expanded as a vertical label stack of quoted uppercase entries with thin rules between them. Timeline uses a fixed 44px label gutter with 8 stacked tracks.

Imagery style: No stock photography, no 3D blobs, no gradients-as-decoration. Imagery is the product's own output and its own machinery: user-uploaded frames and generated stills presented as physical prints pinned inside 2px black frames on the concrete ground; macro shots of hardware (a GPU fan, an SSD, an XLR jack, a tape reel) rendered desaturated with a single orange gel; halftone-dot textures at 6% opacity over panel headers; diagrammatic pipeline art — arrows, numbered nodes, bracket marks — drawn in 1px #8A8A85. Waveforms are drawn as chunky 2px bars, never soft gradients. Thumbnails are always shown cropped to their true aspect ratio with the ratio stamped in the corner as a quoted label: "16:9", "9:16".

Page 36 of 40

7. Signature Design Concept

"The Cutting Mat" — the public entry (Landing) is a full-bleed dark concrete workspace, not a centred SaaS hero.

  • Left two-thirds: an oversized Archivo Black all-caps headline stacked in three lines — SCRIPT / IN, / VIDEO / OUT. at clamp(40px, 9vw, 128px), flush-left, breaking the grid so OUT. overhangs into the right column. Beneath its baseline sits the primary CTA as a hard-edged #FF5A00 rectangle with black uppercase text, pinned flush to the headline's left edge.
  • Right third: a solid #FF5A00 block bleeding off the right viewport edge. On top of it, a single framed 9:16 vertical video still (the user's own recent render) rotated -2deg, with hazard-stripe tape across its top corner and a quoted label "9:16 · RENDER 04" beneath it.
  • Behind everything: a faint 12-column cutting-mat grid in #1E1E20 with column numbers.
  • Below the hero: the pipeline as a numbered spec sheet — the 13 stages of SCRIPT → VIDEO rendered as a vertical ledger of numbered rows with 1px rules, each row stamped DONE / RUNNING / QUEUED in tabular caps, with the running row's left edge marked by a 4px #FF5A00 bar.
  • Honesty element: a plainly worded note that local AI video generation depends on sufficient GPU hardware, with the lightweight mode named as the alternative — recomposing accepted content only.
  • Responsive behavior: at 375px the orange block drops below the headline as a full-width band, the framed still scales to 60vw and stays whole inside the viewport, and the headline wraps to four lines at 40px with no clipping. All readable text and controls stay whole at 375px, 768px and 1280px.

8. Interaction Model & Motion Direction

Interaction Model: Animated Motion Tempo: expressive Hero Dimensionality: flat

Landing Hero Motion Brief

  • Focal subject: The framed 9:16 vertical video still — the user's own recent render — pinned on the safety-orange block, with hazard-stripe tape across its top corner and the quoted label "9:16 · RENDER 04" beneath it.
  • Input → transformation → outcome thesis: The user's own media (input) is pinned, taped and labeled onto the concrete workspace (transformation), becoming a finished, tagged artifact on the cutting mat (outcome). The motion shows the studio's own output being handled like a physical object — it never invents a capability.
  • Motion vocabulary: Snappy and mechanical, never bouncy. 120–180ms cubic-bezier(0.2, 0, 0, 1) on all state changes. Marquee tickers run for the render queue and the "recent activity" strip at ~40px/s. Clip suggestions stamp in with a 90ms scale-from-0.98 plus a 1px outline flash in #FFE600. Playhead and progress fill move linearly with real job progress, with no easing decoration. Hover on a project thumbnail swaps to its poster frame with a hard cut, no crossfade. Scene reorder is a direct drag with a 1px orange insertion rule, no spring physics.
  • Composed first frame: The headline fully set at its clamp size, flush-left, with OUT. overhanging the right column; the orange block bleeding off the right edge; the framed 9:16 still already pinned and taped at -2deg with its quoted label; the cutting-mat grid faintly visible; the primary CTA sitting on the headline's baseline. Nothing is mid-animation — the first frame is a complete, readable composition.
  • Reduced-motion state: Under prefers-reduced-motion, all marquees stop and wrap into readable rows; the ticker items sit in a horizontally scrollable row (overflow-x: auto) whose further items are reached by scrolling; stamp-in and progress animations resolve immediately to their final state; the framed still, headline, labels and CTA remain whole and fully readable.
Page 37 of 40

9. Non-Functional Requirements

NFR-01 — No paid API dependency (explicit) The platform must be designed around free and open-source models and local processing. API integrations may be optional, but the application must work without paid APIs. Rationale: the user's most important requirement.

NFR-02 — Honest capability claims (explicit) The application must not pretend that unlimited AI video generation is free if the required model needs expensive GPU hardware. Rationale: explicit user constraint.

NFR-03 — Hardware-conditional local video generation with graceful fallback (explicit) Local video generation is conditional on the user's computer having sufficient GPU resources; when it is unavailable because of hardware limitations, the system must gracefully fall back to image-based video creation. Rationale: explicit user constraint.

NFR-04 — Hardware-conditional local image generation (explicit) Local/open-source image generation is supported where possible; Stable Diffusion/Flux-compatible local image generation and open-source video-generation models are used where hardware allows. Rationale: explicit user constraint.

NFR-05 — Free/open-source, local-first software (explicit) The software itself should be free/open-source and local-first wherever possible. Rationale: explicit user constraint.

NFR-06 — API key secrecy (explicit) All API keys must be stored in environment variables and NEVER exposed in frontend code. Rationale: explicit user constraint.

NFR-07 — Optional authentication (explicit) Authentication is optional for local personal use. Rationale: explicit user constraint.

NFR-08 — Storage and database defaults (explicit) Storage is the local filesystem by default; SQLite is used for local installation with an option to switch to PostgreSQL later. Rationale: explicit user constraint.

NFR-09 — Music licensing transparency (explicit) Use royalty-free/open-source music by default and clearly indicate the source/license when applicable. Rationale: explicit user constraint.

NFR-10 — Dark professional design, not a kids' site (explicit) The design must be a modern dark professional interface and must NOT look like a children's website. Rationale: explicit user constraint.

NFR-11 — Real working application (explicit) The final result must be a real working application, not merely a UI mockup. Rationale: explicit user constraint.

NFR-12 — Extensibility without rewrite (explicit) The application must be built so additional AI video models can be plugged in later without rewriting the entire application. Rationale: explicit user constraint.

NFR-13 — Local tool detection before generation (required_inference) Local or optional provider capabilities must be detected before generation jobs run, so that a missing tool is reported with its role, optionality, and installation instructions rather than failing opaquely. Rationale: required to make the accepted detection and System Status behavior usable.

NFR-14 — GPU assessment before model selection (required_inference) GPU availability must be assessed before selecting local video or image generation, with the lightweight image-based fallback remaining available. Rationale: required to make the accepted fallback behavior usable.

NFR-15 — Local TTS and Whisper availability (required_inference) Local TTS and Whisper or faster-whisper are required for the corresponding voiceover and caption workflows. Rationale: required to make the accepted voiceover and caption workflows executable.

NFR-16 — Progress persistence (required_inference) Project progress must be persisted so production and editing can be resumed. Rationale: required to make the accepted automatic-save and resume behavior usable.

NFR-17 — Responsive, readable interface (explicit + direction) The interface must be responsive, with headlines, wordmarks, labels, numbers, card text and controls staying entirely inside the viewport and their container at 375px, 768px and 1280px, wrapping or scaling to fit. Rationale: explicit responsive-design requirement plus the authoritative creative direction.

Page 38 of 40

10. Tech Stack

Frontend (explicit)

  • React or Next.js
  • TypeScript
  • Tailwind CSS
  • Started locally with npm install and npm run dev

Backend (explicit)

  • Python
  • FastAPI (the source permits FastAPI or Flask; FastAPI is the stated primary)
  • Started locally with pip install -r requirements.txt and python server.py

Video processing (explicit)

  • FFmpeg for video processing and rendering

AI toolchain (explicit)

  • Whisper / faster-whisper for transcription
  • Piper TTS / Coqui TTS / other open-source TTS engines for voice
  • Open-source LLMs through Ollama
  • Stable Diffusion / Flux-compatible local image generation where hardware allows
  • Open-source video-generation models where hardware allows
  • OpenCV for video/image processing
  • Python for AI/video processing
  • Modular provider architecture so any AI model can be replaced later

Database (explicit)

  • SQLite for local installation, with an option to switch to PostgreSQL later

Storage (explicit)

  • Local filesystem by default

Authentication (explicit)

  • Optional for local personal use

Secrets (explicit)

  • All API keys stored in environment variables and never exposed in frontend code

Documentation (explicit)

  • README containing complete installation instructions for Windows

Design tokens (from authoritative creative direction)

  • Fonts: Archivo Black (headings), Archivo (body and labels); monospace reserved for code/terminal output
  • Palette: #0B0B0C, #151517, #F2F2EF, #FF5A00, #FFE600, #8A8A85, #2A2A2C, #1E1E20
Page 39 of 40

11. Assumptions and Constraints

Assumptions

  • A-01 (required_inference) The frontend and Python backend are installed and run locally on the user's own machine, as the source's run commands imply.
  • A-02 (required_inference) SQLite and the local filesystem are available by default for project state and media.
  • A-03 (required_inference) FFmpeg is available for rendering; if it is missing, the application detects this and explains what it does, whether it is optional, and how to install it.
  • A-04 (required_inference) Local or optional provider capabilities are detected before generation jobs run.
  • A-05 (required_inference) GPU availability is assessed before selecting local video or image generation; the lightweight image-based fallback remains available regardless.
  • A-06 (required_inference) Local TTS and Whisper/faster-whisper are required for the voiceover and caption workflows respectively.
  • A-07 (required_inference) Project progress is persisted so production and editing can be resumed.
  • A-08 (required_inference) If optional API integrations are configured, their keys remain in backend environment variables and are never exposed to the frontend.
  • A-09 (Default — not specified by user) The application is used by a single local user; no multi-user collaboration, sharing, or role model is assumed.
  • A-10 (Default — not specified by user) The README's Windows installation instructions are the primary documented setup path, matching the user's stated environment.

Constraints

  • C-01 (explicit) No paid API dependencies by default: the platform must be designed around free and open-source models and local processing; API integrations may be optional, but the application must work without paid APIs.
  • C-02 (explicit) Do not pretend that unlimited AI video generation is free if the required model needs expensive GPU hardware.
  • C-03 (explicit) Local video generation is conditional on the user's computer having sufficient GPU resources; when it is unavailable because of hardware limitations, the system must gracefully fall back to image-based video creation.
  • C-04 (explicit) Local/open-source image generation is supported where possible; Stable Diffusion/Flux-compatible local image generation and open-source video-generation models are used where hardware allows.
  • C-05 (explicit) The software itself should be free/open-source and local-first wherever possible.
  • C-06 (explicit) All API keys must be stored in environment variables and NEVER exposed in frontend code.
  • C-07 (explicit) Authentication is optional for local personal use.
  • C-08 (explicit) Storage is the local filesystem by default; SQLite is used for local installation with an option to switch to PostgreSQL later.
  • C-09 (explicit) Use royalty-free/open-source music by default and clearly indicate the source/license when applicable.
  • C-10 (explicit) The design must be a modern dark professional interface and must NOT look like a children's website.
  • C-11 (explicit) The final result must be a real working application, not merely a UI mockup.
  • C-12 (explicit) The application must be built so additional AI video models can be plugged in later without rewriting the entire application.
  • C-13 (explicit) The clipping tool is explicitly very important and must be treated as a first-class capability.
  • C-14 (explicit) The product is general-purpose and explicitly not a kids-only website, even though "Kids animation" and "nursery rhymes" are among the supported video types.
  • C-15 (creative direction, authoritative) No blue or indigo in interface chrome; no rounded pills or 16px+ card radii; no gradient-blob heroes, glassmorphism, soft drop shadows, or floating frosted cards; no Inter/Roboto/Arial/Helvetica/Poppins/Lato/Open Sans/system-ui as the interface typeface; no kid-friendly signalling; no stock photography of smiling creators; hazard yellow is never a button fill or body text; the generic indigo/blue-on-white SaaS template is forbidden.
Page 40 of 40

12. Glossary

  • Creator Studio — The personal, self-hosted AI video creation platform specified in this document.
  • SCRIPT → VIDEO — The main workflow that turns a pasted script, an uploaded TXT/DOCX/PDF, or an AI-generated script into a complete exported MP4 through thirteen automatic stages.
  • Scene — A unit of the script breakdown containing scene number, duration, narration, visual description, image prompt, video prompt, camera movement, transition, sound effect, background music, and subtitle text.
  • Preset — A named video-type configuration: YouTube video, YouTube Shorts, Instagram Reels, TikTok, Educational video, Story video, Faceless video, Kids animation, Documentary, Motivational video, Podcast clips, or Custom.
  • Aspect ratio / Format — One of the supported output shapes: 16:9, 9:16, 1:1, or 4:5.
  • TTS — Text-to-speech; the local voice synthesis used for voiceover, provided by Piper TTS, Coqui TTS, or another suitable open-source engine.
  • Whisper / faster-whisper — Free/open-source speech recognition used for automatic transcription and word-level timestamps.
  • Ollama — The local runtime for open-source LLMs used for script generation and other language tasks.
  • FFmpeg — The video processing and rendering engine used throughout the platform.
  • Ken Burns effect — The image-based motion technique (zoom and pan) used in the graceful fallback when local video generation is unavailable.
  • Lightweight mode — The complete-video path for machines without a powerful GPU, using local LLM, local TTS, Whisper, user-provided images/videos, FFmpeg, motion effects, transitions, and captions.
  • Clip / Short — A short vertical (9:16) excerpt produced by the clipping tool from a long-form video, with a title, captions, animated subtitles, and a reframed speaker.
  • Score — The percentage assigned to a suggested clip indicating its detected engagement level.
  • Caption preset — One of Minimal, Bold, Shorts, Karaoke, Highlight, or Professional.
  • Audio ducking — Automatically lowering music volume beneath voiceover.
  • Content package — The automatically generated publishing metadata: YouTube title, YouTube description, SEO keywords, hashtags, tags, thumbnail text, short description, and social media captions.
  • Media Library — The store for images, videos, audio, voiceovers, music, SFX, thumbnails, and generated clips, with search, filtering, and folders.
  • System Status — The page showing FFmpeg, Whisper, Ollama, TTS, Image Generation, Video Generation, GPU, and Storage, with what each does, whether it is optional, and installation instructions when missing.
  • Provider — A pluggable AI backend (TTS, LLM, image generation, video generation, transcription) that can be replaced without rewriting the application.
  • Lightweight fallback — See Lightweight mode.
  • Hazard tape — The repeating 45° #FFE600/#0B0B0C stripe motif used as dividers, render-queue top edges, and timeline trim handles.

No completed page designs yet.

Completed design pages will appear here when they are ready to preview.

Landing: Read platform capabilities
Dashboard: Launch Video → Clips
Video → Clips: Upload long-form video
Video → Clips: Start automatic clipping
Video → Clips: Choose transcription engine
Clips: Review scored suggestions
Clips: Select manual range
Clips: Preview a clip
Clips: Edit times and title
Clips: Apply animated subtitles
Clips: Reframe speaker to 9:16
Clips: 1. Export clip as Short
Clips: 2. Retry failed clip export
Captions: Style clip captions
Media Library: Open exported clips

No completed page designs yet.

Completed design pages will appear here when they are ready to preview.

Landing: Read platform capabilities
Dashboard: Launch Video → Clips
Video → Clips: Upload long-form video
Video → Clips: Start automatic clipping
Video → Clips: Choose transcription engine
Clips: Review scored suggestions
Clips: Select manual range
Clips: Preview a clip
Clips: Edit times and title
Clips: Apply animated subtitles
Clips: Reframe speaker to 9:16
Clips: 1. Export clip as Short
Clips: 2. Retry failed clip export
Captions: Style clip captions
Media Library: Open exported clips